Image encoding / decoding method and recording medium for storing bitstream

By constructing a candidate list for inter prediction using serially derived motion information, the method enhances prediction accuracy in video encoding and decoding, addressing the challenge of high data volumes in high-resolution image transmission and storage.

WO2025147175A1PCT designated stage expired Publication Date: 2025-07-10KT CORP

Patent Information

Application Number
PCT/KR2025/000247
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-03
Filing Date
2025-01-06
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images, including stereoscopic content, leads to higher data volumes, resulting in increased transmission and storage costs, necessitating more efficient image compression technologies.

Method used

A method for constructing a candidate list for inter prediction or intra block copying, involving deriving new candidates based on serially searched motion information, and a device for implementing this method, which enhances prediction accuracy by inserting chained inter prediction candidates into the inter prediction list.

Benefits of technology

Improves prediction accuracy in video encoding and decoding processes, thereby reducing data volume and costs associated with transmitting and storing high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025000247_10072025_PF_FP_ABST
    Figure KR2025000247_10072025_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding method according to the present disclosure may comprise the steps of: configuring an inter-prediction list of a current block; deriving a motion vector of the current block from the inter-prediction list; and acquiring a prediction block for the current block on the basis of the motion vector. At this time, a chained inter-prediction candidate derived on the basis of initial motion information of the current block and at least one piece of motion information searched in a chain on the basis of the initial motion information can be inserted into the inter-prediction list.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and recording medium for storing bitstream

[0001] The present disclosure relates to a video signal processing method and device.

[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, is increasing across various application fields. As image data becomes higher in resolution and quality, the relative amount of data increases compared to conventional image data. Therefore, transmitting image data using existing media such as wired and wireless broadband lines or storing it using existing storage media leads to increased transmission and storage costs. To address these issues arising from the increasing resolution and quality of image data, high-efficiency image compression technologies can be utilized.

[0003] There are various technologies for image compression, such as inter-picture prediction technology that predicts pixel values ​​included in the current picture from pictures before or after the current picture, intra-picture prediction technology that predicts pixel values ​​included in the current picture using pixel information in the current picture, and entropy encoding technology that assigns short codes to values ​​with high frequency of appearance and long codes to values ​​with low frequency of appearance. Using these image compression technologies, image data can be effectively compressed and transmitted or stored.

[0004] Meanwhile, as demand for high-resolution video grows, so does the demand for stereoscopic video content as a new video service. Discussions are underway on video compression technologies to effectively deliver high-resolution and ultra-high-resolution stereoscopic video content.

[0005] The present disclosure aims to provide a method for constructing a candidate list for inter prediction or intra block copying and a device therefor.

[0006] The present disclosure aims to provide a method and a device for deriving a new candidate based on initial motion information and sequentially searched motion information.

[0007] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0008] A video decoding method according to the present disclosure may include the steps of: configuring an inter prediction list of a current block; deriving a motion vector of the current block from the inter prediction list; and obtaining a prediction block for the current block based on the motion vector. At this time, a chained inter prediction candidate derived based on initial motion information of the current block and at least one piece of motion information chainedly searched based on the initial motion information may be inserted into the inter prediction list.

[0009] A video encoding method according to the present disclosure may include the steps of: configuring an inter prediction list of a current block; deriving a motion vector of the current block from the inter prediction list; and obtaining a prediction block for the current block based on the motion vector. At this time, a chained inter prediction candidate derived based on initial motion information of the current block and at least one piece of motion information chainedly searched based on the initial motion information may be inserted into the inter prediction list.

[0010] In the video encoding / decoding method according to the present disclosure, the initial motion information may be derived from an inter prediction candidate having the smallest index in the inter prediction list.

[0011] In the video encoding / decoding method according to the present disclosure, the initial motion information may be derived from a spatial neighboring block or a temporal neighboring block at a predefined position of the current block.

[0012] In the video encoding / decoding method according to the present disclosure, when at least one motion information is searched based on the initial motion information, the reference picture index of the chain inter prediction candidate is set to the reference picture index of the motion information searched most recently, and the motion vector of the chain inter prediction candidate can be derived by combining the motion vectors of the initial motion information and the searched motion information, respectively.

[0013] In the video encoding / decoding method according to the present disclosure, the search for motion information for deriving the chain inter prediction candidate may be repeatedly performed until there is no motion information available for the reference block specified by the last searched motion information.

[0014] In the video encoding / decoding method according to the present disclosure, the motion information of the reference block can be obtained from a predefined position within the reference block.

[0015] In the video encoding / decoding method according to the present disclosure, the search for motion information to derive the chain inter prediction candidate can be performed repeatedly until the number of searches reaches a threshold.

[0016] In the video encoding / decoding method according to the present disclosure, if there is a block vector available for a reference block indicated by the initial motion information or the searched motion information, the block vector may be additionally used when deriving the chain inter prediction candidate.

[0017] In the video encoding / decoding method according to the present disclosure, when a plurality of pieces of motion information are searched based on the initial motion information, a plurality of chain inter prediction candidates can be derived based on the plurality of pieces of motion information.

[0018] In the video encoding / decoding method according to the present disclosure, a first chain inter prediction candidate among the plurality of chain inter prediction candidates may be derived using the plurality of motion information, and a second inter prediction candidate among the plurality of chain inter prediction candidates may be derived by excluding the last motion information among the plurality of motion information.

[0019] In the video encoding / decoding method according to the present disclosure, when a plurality of pieces of motion information are searched based on the initial motion information, the motion vector of the chain inter prediction candidate can be derived by scaling the sum of the motion vectors of the initial motion information and each of the plurality of pieces of motion information.

[0020] In the video encoding / decoding method according to the present disclosure, the scaling may be performed based on the distance between the current picture and the reference picture of the current block and the distance between the current picture and the reference picture indicated by the last searched motion information.

[0021] In the video encoding / decoding method according to the present disclosure, the inter prediction list may be a motion information merge list or a motion vector prediction list.

[0022] According to the present disclosure, a computer-readable recording medium for storing a bitstream generated by an image encoding method can be provided.

[0023] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0024] According to the present disclosure, a method for constructing an improved candidate list for inter prediction or intra block copying and a device therefor are provided, thereby improving prediction accuracy.

[0025] According to the present disclosure, prediction accuracy can be improved by providing a method for deriving new candidates based on initial motion information and sequentially searched motion information.

[0026] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0027] FIG. 1 is a block diagram illustrating an image encoding device according to an embodiment of the present disclosure.

[0028] FIG. 2 is a block diagram showing an image decoding device according to an embodiment of the present disclosure.

[0029] Figure 3 is a diagram schematically illustrating the process of performing inter prediction in an encoder and decoder.

[0030] Figure 4 shows an example in which motion estimation is performed.

[0031] Figures 5 and 6 illustrate examples in which a prediction block of a current block is generated based on motion information generated through motion estimation.

[0032] Figure 7 shows the locations referenced to derive motion vector prediction values.

[0033] Figure 8 is a diagram for explaining a template-based motion estimation method.

[0034] Figure 9 shows examples of template configurations.

[0035] Figure 10 is a diagram for explaining a motion estimation method based on a bilateral matching method.

[0036] Figure 11 is a diagram for explaining a motion estimation method based on a one-way matching method.

[0037] Figures 12 and 13 illustrate examples in which prediction blocks are generated according to the precision of a motion vector.

[0038] Figure 14 shows an example in which motion compensation based on a translational model and a zooming model is performed for the current block.

[0039] Figure 15 shows an example in which motion compensation based on a translational model and a rotational model is performed for the current block.

[0040] Figures 16 and 17 illustrate examples of generating a prediction block for a current block using control point motion vectors.

[0041] Figure 18 shows an example of generating a prediction block for the current block using three control point motion vectors.

[0042] Figure 19 shows an example in which a motion vector is derived in sub-block units.

[0043] Figures 20 and 21 illustrate examples in which motion vectors are derived in units of sub-blocks within the current block when SbTMVP is applied.

[0044] Figures 22 and 23 are diagrams showing examples in which prediction blocks are derived according to motion vector precision.

[0045] Figures 24 and 25 are diagrams for explaining the process of encoding and decoding a motion vector difference value, respectively, when the AMVR method is applied.

[0046] Figure 26 is an example diagram to explain the configuration of the BV area.

[0047] FIG. 27 is a flowchart of a method for configuring a motion information merge list according to one embodiment of the present disclosure.

[0048] Figure 28 shows an example of deriving a chain merge candidate.

[0049] Figure 29 shows an example of deriving a chain merging candidate using a block vector.

[0050] Figure 30 shows the location from which the motion information of the reference block is derived.

[0051] Figure 31 shows an example of deriving a chain block vector candidate.

[0052] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.

[0053] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.

[0054] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0055] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0056] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereinafter, identical components in the drawings will be designated by the same reference numerals, and redundant descriptions of identical components will be omitted.

[0057] FIG. 1 is a block diagram illustrating an image encoding device according to an embodiment of the present disclosure.

[0058] Referring to FIG. 1, a video encoding device (100) may include a picture segmentation unit (110), a prediction unit (120, 125), a transformation unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse transformation unit (145), a filter unit (150), and a memory (155).

[0059] Each component shown in Fig. 1 is independently depicted to represent different characteristic functions in the video encoding device, and does not mean that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or one component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present disclosure as long as they do not deviate from the essence of the present disclosure.

[0060] Additionally, some components may not be essential components that perform the essential functions of the present disclosure, but may be optional components merely used to enhance performance. The present disclosure may be implemented by including only components essential to implementing the essence of the present disclosure, excluding components used solely for performance enhancement. A structure that includes only essential components, excluding optional components used solely for performance enhancement, is also within the scope of the present disclosure.

[0061] The picture splitting unit (110) can split the input picture into at least one processing unit. At this time, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The picture splitting unit (110) can split one picture into a combination of multiple coding units, prediction units, and transform units, and select one combination of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function) to encode the picture.

[0062] For example, a picture can be split into multiple coding units. A recursive tree structure such as a quad tree, a ternary tree, or a binary tree can be used to split a coding unit in a picture. A coding unit that is split into other coding units starting from an image or the largest coding unit as the root can be split into as many child nodes as the number of split coding units. A coding unit that cannot be split any further according to a certain restriction becomes a leaf node. For example, assuming that a quad tree split is applied to a coding unit, a coding unit can be split into at most four different coding units.

[0063] Hereinafter, in the embodiments of the present disclosure, the encoding unit may be used to mean a unit that performs encoding or may be used to mean a unit that performs decoding.

[0064] A prediction unit may be divided into at least one square or rectangular shape of the same size within a single coding unit, or may be divided such that one prediction unit among the divided prediction units within a single coding unit has a different shape and / or size from another prediction unit.

[0065] When predicting within a screen, the transformation unit and the prediction unit can be set to be the same. In this case, the encoding unit can be divided into multiple transformation units, and then intra-screen prediction can be performed for each transformation unit. The encoding unit can be divided in the horizontal direction or the vertical direction. The number of transformation units generated by dividing the encoding unit can be 2 or 4, depending on the size of the encoding unit. Alternatively, when the size of the transformation unit is small, multiple transformation units can be set as a single prediction unit.

[0066] The prediction unit (120, 125) may include an inter-prediction unit (120) that performs inter-prediction and an intra-prediction unit (125) that performs intra-prediction. It may be determined whether to use inter-prediction or intra-prediction for an encoding unit, and specific information (e.g., reference sample line, intra-prediction mode, motion vector, reference picture, etc.) according to each prediction method may be determined. At this time, the processing unit where prediction is performed and the processing unit where the prediction method and specific contents are determined may be different. For example, the prediction method and prediction mode, etc. are determined in the encoding unit, and the prediction may be performed in the prediction unit or the transformation unit. The residual value (residual block) between the generated prediction block and the original block may be input to the transformation unit (130). In addition, the prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value in the entropy encoding unit (165) and transmitted to the decoding device. When using a specific encoding mode, it is also possible to encode the original block as is and transmit it to the decoding unit without generating a prediction block through the prediction unit (120, 125).

[0067] The inter-screen prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous or subsequent pictures of the current picture, and in some cases, may predict a prediction unit based on information of a portion of an encoded region within the current picture. The inter-screen prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.

[0068] The reference picture interpolation unit can receive reference picture information from the memory (155) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.

[0069] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods can be used to derive a motion vector, such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm). The motion vector can have a motion vector value of 1 / 2 or 1 / 4 pixel unit based on the interpolated pixel. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various methods can be used as motion prediction methods, such as the Skip method, the Merge method, the AMVP (Advanced Motion Vector Prediction) method, and the Intra Block Copy method.

[0070] The on-screen prediction unit (125) can generate a prediction block based on reference pixel information, which is pixel information within the current picture. The reference pixel information can be derived from one selected from among a plurality of reference pixel lines. The Nth reference pixel line among the plurality of reference pixel lines can include left pixels having an x-axis difference of N from the upper left pixel within the current block and upper pixels having a y-axis difference of N from the upper left pixel. The number of reference pixel lines that the current block can select can be 1, 2, 3, or 4.

[0071] If the neighboring blocks of the current prediction unit are blocks that have performed inter-screen prediction and the reference pixel is a pixel that has performed inter-screen prediction, the reference pixel included in the block that has performed inter-screen prediction can be replaced with the reference pixel information of the neighboring block that has performed intra-screen prediction. That is, if the reference pixel is unavailable, the unavailable reference pixel information can be replaced with information from at least one of the available reference pixels.

[0072] In intra-screen prediction, the prediction mode can have a directional prediction mode that uses reference pixel information according to the prediction direction, and a non-directional mode that does not use directional information when performing prediction. The mode for predicting luminance information and the mode for predicting chrominance information can be different, and the intra-screen prediction mode information used to predict luminance information or the predicted luminance signal information can be utilized to predict chrominance information.

[0073] When performing intra-screen prediction, if the size of the prediction unit and the size of the transformation unit are the same, intra-screen prediction for the prediction unit can be performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side.

[0074] The on-screen prediction method can generate prediction blocks by applying a smoothing filter to reference pixels according to the prediction mode. Depending on the selected reference pixel line, whether or not the smoothing filter is applied can be determined.

[0075] In order to perform an intra-screen prediction method, the intra-screen prediction mode of the current prediction unit can be predicted from the intra-screen prediction modes of prediction units existing around the current prediction unit. When the prediction mode of the current prediction unit is predicted using mode information predicted from the surrounding prediction units, if the intra-screen prediction modes of the current prediction unit and the surrounding prediction units are the same, information indicating that the prediction modes of the current prediction unit and the surrounding prediction units are the same can be transmitted using predetermined flag information, and if the prediction modes of the current prediction unit and the surrounding prediction units are different, entropy encoding can be performed to encode the prediction mode information of the current block.

[0076] Additionally, a residual block containing residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130).

[0077] In the transformation unit (130), the residual block including the residual value information of the prediction unit generated through the original block and the prediction unit (120, 125) can be transformed using a transformation method such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on at least one of the size of the transformation unit, the shape of the transformation unit, the prediction mode of the prediction unit, or the prediction mode information within the screen of the prediction unit. Meanwhile, the transformation can be performed by separating the horizontal direction and the vertical direction.

[0078] After performing transformations in the horizontal and vertical directions, a secondary transformation can be performed. The secondary transformation may be in a form in which the horizontal and vertical directions are not separated. The secondary transformation can be performed on the transformation coefficients obtained by the primary transformation to generate final transformation coefficients. Meanwhile, the number of final transformation coefficients output by the secondary transformation may be smaller than the number of transformation coefficients input for the secondary transformation. Specifically, the secondary transformation can be performed using a reduced transformation matrix having different numbers of columns and rows.

[0079] The quantization unit (135) can quantize values ​​converted to the frequency domain by the transformation unit (130). The quantization coefficients can vary depending on the block or the importance of the image. The values ​​produced by the quantization unit (135) can be provided to the dequantization unit (140) and the reordering unit (160).

[0080] The rearrangement unit (160) can perform rearrangement of coefficient values ​​for quantized residual values.

[0081] The reordering unit (160) can change a two-dimensional block-shaped coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the reordering unit (160) can change the two-dimensional block-shaped coefficient into a one-dimensional vector form by scanning from the DC coefficient to the coefficient of the high-frequency region using a zig-zag scan method. Depending on the size of the conversion unit and the intra-screen prediction mode, a vertical scan that scans the two-dimensional block-shaped coefficient in the column direction, a horizontal scan that scans the two-dimensional block-shaped coefficient in the row direction, or a diagonal scan that scans the two-dimensional block-shaped coefficient in the diagonal direction may be used instead of the zig-zag scan. That is, depending on the size of the conversion unit and the intra-screen prediction mode, it is possible to determine which scan method among the zig-zag scan, the vertical scan, the horizontal scan, or the diagonal scan is to be used.

[0082] The entropy encoding unit (165) can perform entropy encoding based on the values ​​produced by the rearrangement unit (160). Entropy encoding can use various encoding methods such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).

[0083] The entropy encoding unit (165) can encode various information such as residual value coefficient information of the encoding unit, block type information, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information from the rearrangement unit (160) and the prediction unit (120, 125).

[0084] The entropy encoding unit (165) can entropy encode the coefficient values ​​of the encoding unit input from the rearrangement unit (160).

[0085] The inverse quantization unit (140) and the inverse transformation unit (145) inversely quantize the values ​​quantized in the quantization unit (135) and inversely transform the values ​​transformed in the transformation unit (130). The residual values ​​generated in the inverse quantization unit (140) and the inverse transformation unit (145) can be combined with the predicted prediction units predicted through the motion estimation unit, motion compensation unit, and intra-screen prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.

[0086] The filter unit (150) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter).

[0087] A deblocking filter can remove block distortion caused by boundaries between blocks in a reconstructed picture. To determine whether to perform deblocking, a deblocking filter can be applied to the current block based on the pixels contained in several columns or rows within the block. When applying a deblocking filter to a block, a strong filter or a weak filter can be applied depending on the required deblocking filtering strength. Furthermore, when applying a deblocking filter, horizontal and vertical filtering can be processed in parallel when performing vertical and horizontal filtering.

[0088] The offset correction unit can correct the offset from the original image on a pixel-by-pixel basis for an image that has undergone deblocking. To perform offset correction for a specific picture, the pixels contained in the image can be divided into a certain number of regions, the regions to be offset can be determined, and the offset can be applied to those regions. Alternatively, the offset can be applied by considering the edge information of each pixel.

[0089] Adaptive Loop Filtering (ALF) can be performed based on the comparison of the filtered restored image with the original image. After dividing the pixels included in the image into predetermined groups, a filter to be applied to each group can be determined, and filtering can be performed differentially for each group. Information regarding whether to apply ALF can be transmitted by luminance signal for each coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied can vary depending on each block. Furthermore, an ALF filter of the same form (fixed form) can be applied regardless of the characteristics of the target block.

[0090] The memory (155) can store a restoration block or picture produced through the filter unit (150), and the stored restoration block or picture can be provided to the prediction unit (120, 125) when performing inter-screen prediction.

[0091] FIG. 2 is a block diagram showing an image decoding device according to an embodiment of the present disclosure.

[0092] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (210), a rearrangement unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).

[0093] When a video bitstream is input to a video encoding device, the input bitstream can be decoded in the opposite procedure to that of the video encoding device.

[0094] The entropy decoding unit (210) can perform entropy decoding in a procedure opposite to that of the entropy encoding unit of the video encoding device. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied in response to the method performed in the video encoding device.

[0095] The entropy decoding unit (210) can decode information related to intra-screen prediction and inter-screen prediction performed in the encoding device.

[0096] The reordering unit (215) can perform reordering based on the method in which the bitstream entropy-decoded by the entropy decoding unit (210) is reordered by the encoding unit. The coefficients expressed in the form of a one-dimensional vector can be reordered by restoring them back to coefficients in the form of a two-dimensional block. The reordering unit (215) can perform reordering by receiving information related to the coefficient scanning performed by the encoding unit and performing reverse scanning based on the scanning order performed by the corresponding encoding unit.

[0097] The dequantization unit (220) can perform dequantization based on the quantization parameters provided from the encoding device and the coefficient values ​​of the rearranged block.

[0098] The inverse transform unit (225) can perform an inverse transform of the transform performed by the transform unit on the quantization result performed by the image encoding device. That is, at least one of an inverse transform of a secondary transform (secondary inverse transform) or an inverse transform for DCT, DST, and KLT (i.e., first inverse transform) can be performed. The inverse transform can be performed based on a transmission unit determined by the image encoding device. The inverse transform unit (225) of the image decoding device can determine a transform matrix for the second inverse transform or a transform technique (e.g., DCT, DST, KLT) for the first inverse transform according to a plurality of pieces of information such as a prediction method, the size and shape of the current block, the prediction mode, and the prediction direction within the screen. Alternatively, information for determining the transform matrix or the transform technique may be explicitly encoded and signaled.

[0099] The prediction unit (230, 235) can generate a prediction block based on the prediction block generation related information provided by the entropy decoding unit (210) and the previously decoded block or picture information provided by the memory (245).

[0100] As described above, when performing intra-screen prediction in the same manner as the operation in the video encoding device, if the size of the prediction unit and the size of the transformation unit are the same, intra-screen prediction for the prediction unit is performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side. However, when performing intra-screen prediction, if the size of the prediction unit and the size of the transformation unit are different, intra-screen prediction can be performed using reference pixels based on the transformation unit. In addition, intra-screen prediction using NxN division only for the minimum coding unit can be used.

[0101] The prediction unit (230, 235) may include a prediction unit determination unit, an inter-screen prediction unit, and an intra-screen prediction unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (210), prediction mode information of an intra-screen prediction method, and motion prediction-related information of an inter-screen prediction method, and may distinguish a prediction unit from a current encoding unit and determine whether the prediction unit performs inter-screen prediction or intra-screen prediction. The inter-screen prediction unit (230) may perform inter-screen prediction on the current prediction unit based on information included in at least one of a previous picture or a subsequent picture of the current picture including the current prediction unit, using information necessary for inter-screen prediction of the current prediction unit provided from the video encoding device. Alternatively, inter-screen prediction may be performed based on information on a pre-restored portion of the current picture including the current prediction unit.

[0102] In order to perform inter-screen prediction, it is possible to determine whether the motion prediction method of the prediction unit included in the encoding unit is Skip Mode, Merge Mode, AMVP Mode, or Intra-screen Block Copy Mode based on the encoding unit.

[0103] The intra-screen prediction unit (235) can generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit that has performed intra-screen prediction, intra-screen prediction can be performed based on intra-screen prediction mode information of the prediction unit provided by the video encoding device. The intra-screen prediction unit (235) can include an AIS (Adaptive Intra Smoothing) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a part that performs filtering on the reference pixels of the current block, and can determine and apply whether to apply the filter according to the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the video encoding device. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.

[0104] The reference pixel interpolation unit can generate a reference pixel of a pixel unit less than an integer value by interpolating the reference pixel when the prediction mode of the prediction unit is a prediction unit that performs intra-screen prediction based on the pixel value interpolated from the reference pixel. If the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.

[0105] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.

[0106] Information regarding whether a deblocking filter has been applied to a corresponding block or picture may be received from a video encoding device, and if a deblocking filter has been applied, information regarding whether a strong or weak filter has been applied. The deblocking filter of the video decoding device may receive information related to the deblocking filter provided by the video encoding device, and the video decoding device may perform deblocking filtering on the corresponding block.

[0107] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and offset value information.

[0108] ALF can be applied to an encoding unit based on information such as whether ALF is applied and ALF coefficient information provided from an encoding device. This ALF information can be provided by being included in a specific parameter set.

[0109] The memory (245) can store a restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to an output unit.

[0110] As described above, in the following embodiments of the present disclosure, for convenience of explanation, the term coding unit is used as an encoding unit, but it may also be a unit that performs not only encoding but also decoding.

[0111] In addition, the current block represents a block to be encoded / decoded, and may represent a coding tree block (or coding tree unit), an encoding block (or encoding unit), a transform block (or transform unit), a prediction block (or prediction unit), or a block to which an in-loop filter is applied, depending on the encoding / decoding step. In this specification, a 'unit' represents a basic unit for performing a specific encoding / decoding process, and a 'block' may represent a pixel array of a predetermined size. Unless otherwise distinguished, 'block' and 'unit' may be used with the same meaning. For example, in the embodiment described below, an encoding block (coding block) and an encoding unit (coding unit) may be understood to have the same meaning.

[0112] Furthermore, we will refer to the picture that contains the current block as the current picture.

[0113] When encoding the current picture, redundant data between pictures can be removed through inter-prediction. Inter-prediction can be performed on a block-by-block basis. Specifically, a prediction block of the current block can be generated from a reference picture using motion information of the current block. Here, the motion information can include at least one of a motion vector, a reference picture index, and a prediction direction.

[0114] Figure 3 is a diagram schematically illustrating the process of performing inter prediction in an encoder and decoder.

[0115] As in the example illustrated in FIG. 3, to perform inter prediction, motion information for the current block can be acquired (S310). Here, the motion information can include at least one of a motion vector, a reference picture index, or a weight applied to the prediction block. For the current block, motion information for at least one of the L0 direction or the L1 direction can be acquired.

[0116] In the encoder, motion information of the current block can be derived through motion estimation, and the derived motion information can be encoded and signaled to the decoder. Meanwhile, the encoding / decoding of motion information can be based on a motion information merging mode, a motion vector prediction mode, a template-based motion estimation method, or a bilateral matching method, which will be described later.

[0117] In the decoder, motion information of the current block can be derived based on the information transmitted from the encoder.

[0118] Alternatively, the motion information of the current block can be derived from the decoder in the same manner as in the encoder. This method can be referred to as decoder-side motion estimation.

[0119] Once motion information for the current block is derived, a prediction block for the current block can be obtained based on the derived motion information (S320). For example, a reference block spaced apart by a motion vector from the current block's position within the reference picture can be set as the prediction block for the current block.

[0120] Below, we will explain in more detail the process of calculating inter predictions.

[0121] The motion information of the current block can be generated through motion estimation.

[0122] Figure 4 shows an example in which motion estimation is performed.

[0123] In Fig. 4, it is assumed that the POC (Picture Order Count) of the current picture is T, and the POC of the reference picture is (T-1).

[0124] A search range for motion estimation can be set from the same location as the reference point of the current block within the reference picture. Here, the reference point may be the location of the upper left sample of the current block.

[0125] For example, in FIG. 4, a rectangle of size (w0+w01) and (h0+h1) is set as a search range centered on a reference point. In the above example, w0, w1, h0, and h1 may have the same value. Alternatively, at least one of w0, w1, h0, and h1 may be set to have a different value from the other. Alternatively, the sizes of w0, w1, h0, and h1 may be determined so as not to exceed a Coding Tree Unit (CTU) boundary, a slice boundary, a tile boundary, or a picture boundary.

[0126] Within the search range, reference blocks of the same size as the current block can be set, and the cost of each reference block relative to the current block can be measured. The cost can be calculated using the similarity between the two blocks.

[0127] For example, the cost can be calculated based on the absolute sum of the differences between the original samples in the current block and the original samples (or reconstructed samples) in the reference block. A smaller absolute sum can reduce the cost.

[0128] Afterwards, the cost of each reference block is compared, and the reference block with the optimal cost can be set as the prediction block of the current block.

[0129] Additionally, the distance between the current block and the reference block can be set as a motion vector. Specifically, the x-coordinate difference and the y-coordinate difference between the current block and the reference block can be set as the motion vector.

[0130] Furthermore, the index of the picture containing the reference block identified through motion estimation is set as the reference picture index.

[0131] Additionally, the prediction direction can be set based on whether the reference picture belongs to the L0 reference picture list or the L1 reference picture list.

[0132] Additionally, motion estimation can be performed for each of the L0 direction and the L1 direction. If prediction is performed for both the L0 direction and the L1 direction, motion information in the L0 direction and motion information in the L1 direction can be generated, respectively.

[0133] Figures 5 and 6 illustrate examples in which a prediction block of a current block is generated based on motion information generated through motion estimation.

[0134] Figure 5 shows an example of generating a prediction block by unidirectional (i.e., L0 direction) prediction, and Figure 6 shows an example of generating a prediction block by bidirectional (i.e., L0 and L1 direction) prediction.

[0135] In the case of unidirectional prediction, a prediction block of the current block is generated using a single motion information. For example, the motion information may include an L0 motion vector, an L0 reference picture index, and prediction direction information indicating the L0 direction.

[0136] In the case of bidirectional prediction, a prediction block is generated using two pieces of motion information. For example, a reference block in the L0 direction, determined based on motion information about the L0 direction (L0 motion information), can be set as an L0 prediction block, and an L1 prediction block can be generated based on a reference block in the L1 direction, determined based on motion information about the L1 direction (L1 motion information). Thereafter, the L0 prediction block and the L1 prediction block can be weighted and combined to generate a prediction block of the current block.

[0137] In the examples shown in FIGS. 4 to 6, the L0 reference picture is illustrated as existing in the previous direction of the current picture (i.e., having a POC value smaller than that of the current picture), and the L1 reference picture is illustrated as existing in the subsequent direction of the current picture (i.e., having a POC value larger than that of the current picture).

[0138] However, unlike the illustrated example, the L0 reference picture may exist in the subsequent direction of the current picture, or the L1 reference picture may exist in the previous direction of the current picture. For example, both the L0 reference picture and the L1 reference picture may exist in the previous direction of the current picture, or both may exist in the subsequent direction of the current picture. Alternatively, bidirectional prediction may be performed using the L0 reference picture existing in the subsequent direction of the current picture and the L1 reference picture existing in the previous direction of the current picture.

[0139] Motion information for blocks for which inter prediction has been performed can be stored in memory. At this time, the motion information can be stored on a sample-by-sample basis. Specifically, the motion information for a block to which a specific sample belongs can be stored as motion information for that specific sample. The stored motion information can be used to derive motion information for neighboring blocks to be encoded / decoded in the future.

[0140] In the encoder, information encoding residual samples corresponding to the difference between the sample of the current block (i.e., the original sample) and the predicted sample, and motion information required to generate a predicted block can be signaled to the decoder. The decoder can decode information about the signaled difference value to derive a difference sample, and add a prediction sample within the predicted block generated using the motion information to the difference sample to generate a restored sample.

[0141] At this time, in order to effectively compress the motion information signaled to the decoder, one of a plurality of inter prediction modes may be selected. Here, the plurality of inter prediction modes may include a motion information merging mode and a motion vector prediction mode.

[0142] The motion vector prediction mode is a mode that signals by encoding the difference between a motion vector and a motion vector prediction value. Here, the motion vector prediction value can be derived based on motion information of neighboring blocks or neighboring samples adjacent to the current block.

[0143] Figure 7 shows the locations referenced to derive motion vector prediction values.

[0144] For convenience of explanation, the current block is assumed to have a size of 4x4.

[0145] In the illustrated example, 'LB' represents a sample contained in the leftmost column and bottommost row within the current block. 'RT' represents a sample contained in the rightmost column and topmost row within the current block. A0 to A4 represent samples neighboring to the left of the current block, and B0 to B5 represent samples neighboring to the top of the current block. For example, A1 represents a sample neighboring to the left of LB, and B1 represents a sample neighboring to the top of RT.

[0146] Col indicates the location of a sample neighboring the lower right of the current block within a co-located picture. A co-located picture is a picture different from the current picture, and information for specifying the co-located picture (e.g., a co-located picture index) can be explicitly encoded and signaled in the bitstream. Alternatively, a reference picture having a predefined reference picture index can be set as the co-located picture.

[0147] The motion vector prediction value of the current block can be derived from at least one motion vector prediction candidate included in a motion vector prediction list.

[0148] The number of motion vector prediction candidates that can be inserted into the motion vector prediction list (i.e., the size of the list) may be predefined in the encoder and decoder. For example, the maximum number of motion vector prediction candidates may be 2.

[0149] A motion vector stored at the location of a neighboring sample adjacent to the current block or a scaled motion vector derived by scaling the motion vector can be inserted into the motion vector prediction list as a motion vector prediction candidate. At this time, the motion vector prediction candidates can be derived by scanning the neighboring samples adjacent to the current block in a predefined order.

[0150] For example, it is possible to check whether a motion vector is stored at each location in the order of A0 to A4. Then, according to the above scanning order, the first available motion vector found can be inserted into the motion vector prediction list as a motion vector prediction candidate.

[0151] As another example, in the order of A0 to A4, it is checked whether a motion vector is stored at each position, and the motion vector of the position that is found first and has the same reference picture as the current block can be inserted into the motion vector prediction list as a motion vector prediction candidate. If there is no neighboring sample that has the same reference picture as the current block, a motion vector prediction candidate can be derived based on the first found available vector. Specifically, the first found available motion vector can be scaled, and then the scaled motion vector can be inserted into the motion vector prediction list as a motion vector prediction candidate. At this time, the scaling can be performed based on the output order difference between the current picture and the reference picture (i.e., the POC difference) and the output order difference between the current picture and the reference picture of the neighboring sample (i.e., the POC difference).

[0152] Furthermore, it is possible to check whether a motion vector is stored at each location in the order of B0 to B5. Then, according to the above scanning order, the first available motion vector found can be inserted into the motion vector prediction list as a motion vector prediction candidate.

[0153] As another example, in the order of B0 to B5, it is checked whether a motion vector is stored at each position, and the motion vector of the position that has the same reference picture as the current block that is found first can be inserted into the motion vector prediction list as a motion vector prediction candidate. If there is no neighboring sample that has the same reference picture as the current block, a motion vector prediction candidate can be derived based on the first found available vector. Specifically, the first found available motion vector can be scaled, and then the scaled motion vector can be inserted into the motion vector prediction list as a motion vector prediction candidate. At this time, the scaling can be performed based on the output order difference between the current picture and the reference picture (i.e., the POC difference) and the output order difference between the current picture and the reference picture of the neighboring sample (i.e., the POC difference).

[0154] As in the example described above, a motion vector prediction candidate can be derived from a sample adjacent to the left of the current block, and a motion vector prediction candidate can be derived from a sample adjacent to the top of the current block.

[0155] At this time, the motion vector prediction candidate derived from the left sample may be inserted into the motion vector prediction list before the motion vector prediction candidate derived from the upper sample. In this case, the index assigned to the motion vector prediction candidate derived from the left sample may have a smaller value than the motion vector prediction candidate derived from the upper sample.

[0156] Conversely, the motion vector prediction candidate derived from the top sample may be inserted into the motion vector prediction list before the motion vector prediction candidate derived from the left sample.

[0157] Among the motion vector prediction candidates included in the above motion vector prediction list, the motion vector prediction candidate with the highest encoding efficiency can be set as the motion vector predictor (MVP) of the current block. In addition, index information indicating the motion vector prediction candidate set as the motion vector predictor of the current block among the plurality of motion vector prediction candidates can be encoded and signaled to a decoder. When the number of motion vector prediction candidates is two, the index information can be a 1-bit flag (e.g., an MVP flag). In addition, a motion vector difference (MVD), which is the difference between the motion vector of the current block and the motion vector predictor, can be encoded and signaled to a decoder.

[0158] The decoder can construct a motion vector prediction list, similar to the encoder. Furthermore, it can decode index information from the bitstream and select one of multiple motion vector prediction candidates based on the decoded index information. The selected motion vector prediction candidate can be set as the motion vector prediction value of the current block.

[0159] Additionally, the motion vector differential can be decoded from the bitstream. Afterwards, the motion vector of the current block can be derived by combining the motion vector prediction value and the motion vector differential value.

[0160] When bidirectional prediction is applied to the current block, a motion vector prediction list can be generated for each of the L0 and L1 directions. That is, the motion vector prediction list can be composed of motion vectors in the same direction. Accordingly, the motion vector of the current block and the motion vector prediction candidates included in the motion vector prediction list have the same direction.

[0161] When the motion vector prediction mode is selected, reference picture index and prediction direction information can be explicitly encoded and signaled to the decoder. For example, when there are multiple reference pictures in the reference picture list and motion estimation is performed for each of the multiple reference pictures, a reference picture index for specifying a reference picture from which motion information of the current block is derived among the multiple reference pictures can be explicitly encoded and signaled to the decoder.

[0162] At this time, if the reference picture list contains only one reference picture, encoding / decoding of the reference picture index may be omitted.

[0163] The prediction direction information may be an index pointing to one of L0 unidirectional prediction, L1 unidirectional prediction, or bidirectional prediction. Alternatively, an L0 flag indicating whether prediction is performed in the L0 direction and an L1 flag indicating whether prediction is performed in the L1 direction may be encoded and signaled, respectively.

[0164] Motion Information Merge Mode is a mode in which the motion information of the current block is set to be identical to the motion information of neighboring blocks. In Motion Information Merge Mode, motion information can be encoded / decoded using a motion information merge list.

[0165] Motion information merging candidates can be derived based on motion information from neighboring blocks or neighboring samples adjacent to the current block. For example, after defining reference locations around the current block, it is possible to check whether motion information exists at the defined reference locations. If motion information exists at the defined reference locations, the motion information at those locations can be inserted into the motion information merging list as a motion information merging candidate.

[0166] In the example of Fig. 7, the predefined reference positions may include at least one of A0, A1, B0, B1, B5, and Col. Furthermore, motion information merging candidates may be derived in the order of A1, B1, B0, A0, B5, and Col.

[0167] Among the motion information merge candidates included in the motion information merge list, the motion information of the motion information merge candidate with the optimal cost can be set as the motion information of the current block. Furthermore, index information (e.g., a merge index) indicating the motion information merge candidate selected from among the multiple motion information merge candidates can be encoded and transmitted to the decoder.

[0168] In the decoder, a motion information merge list can be constructed in the same manner as in the encoder. Furthermore, motion information merge candidates can be selected based on the merge index decoded from the bitstream. The motion information of the selected motion information merge candidate can be set as the motion information of the current block.

[0169] Unlike the motion vector prediction list, the motion information merge list is composed of a single list regardless of the prediction direction. That is, the motion information merge candidates included in the motion information merge list may have only L0 motion information or only L1 motion information, or may have bidirectional motion information (i.e., L0 motion information and L1 motion information).

[0170] Motion information about the current block can also be derived using the restoration sample area surrounding the current block. Here, the restoration sample area used to derive motion information about the current block can be referred to as a template.

[0171] Figure 8 is a diagram for explaining a template-based motion estimation method.

[0172] In Fig. 4, it is described that the predicted block of the current block is determined based on the cost between the current block and the reference block within the search range. According to the present embodiment, unlike Fig. 4, motion estimation for the current block can be performed based on the cost between a template neighboring the current block (hereinafter referred to as the "current template") and a reference template having the same size and shape as the current template.

[0173] For example, the cost can be calculated based on the absolute sum of the differences between the restored samples in the current template and the restored samples in the reference block. A smaller absolute sum can reduce the cost.

[0174] Once the reference template with the optimal cost is determined relative to the current template within the search range, the reference block neighboring the reference template can be set as the predicted block of the current block.

[0175] And, motion information of the current block can be set based on the distance between the current block and the reference block, the index of the picture to which the reference block belongs, and whether the reference picture is included in the L0 or L1 reference picture list.

[0176] Since the template defines the restored area around the current block as a template, the decoder itself can perform motion estimation in the same manner as the encoder. Accordingly, when deriving motion information using a template, there is no need to encode and signal the motion information other than information indicating whether a template is used.

[0177] The current template may include at least one of a region adjacent to the top of the current block or a region adjacent to the left of the current block. The region adjacent to the top may include at least one row, and the region adjacent to the left may include at least one column.

[0178] Figure 9 shows examples of template configurations.

[0179] The current template can be configured according to one of the examples illustrated in FIG. 9.

[0180] Alternatively, unlike the example illustrated in FIG. 9, the template may be configured only with the area adjacent to the left of the current block, or only with the area adjacent to the top of the current block.

[0181] The size and / or shape of the current template may be predefined in the encoder and decoder.

[0182] Alternatively, after defining a plurality of template candidates having different sizes and / or shapes, index information specifying one of the plurality of template candidates can be encoded and signaled to a decoder.

[0183] Alternatively, one of multiple template candidates can be adaptively selected based on at least one of the size, shape, or position of the current block. For example, if the current block borders the upper boundary of the CTU, the current template can be constructed using only the region adjacent to the left of the current block.

[0184] Template-based motion estimation can be performed for each of the reference pictures stored in the reference picture list. Alternatively, motion estimation can be performed only for some of the reference pictures. For example, motion estimation can be performed only for reference pictures with a reference picture index of 0, or only for reference pictures with a reference picture index less than a threshold, or only for reference pictures with a POC difference from the current picture less than a threshold.

[0185] Alternatively, a reference picture index can be explicitly encoded and signaled, and then motion estimation can be performed only for the reference picture pointed to by the reference picture index.

[0186] Alternatively, motion estimation can be performed on the reference picture of the neighboring block corresponding to the current template. For example, if the template consists of a left adjacent region and an upper adjacent region, at least one reference picture can be selected using at least one of the reference picture index of the left adjacent block or the reference picture index of the upper adjacent block. Thereafter, motion estimation can be performed on the selected at least one reference picture.

[0187] Information indicating whether template-based motion estimation is applied can be encoded and signaled to a decoder. The information can be a 1-bit flag. For example, a true (1) flag indicates that template-based motion estimation is applied in the L0 direction and L1 direction of the current block. On the other hand, a false (0) flag indicates that template-based motion estimation is not applied. In this case, motion information of the current block can be derived based on a motion information merging mode or a motion vector prediction mode.

[0188] Conversely, template-based motion estimation may be applied only when it is determined that neither the motion information merging mode nor the motion vector prediction mode is applied to the current block. For example, if the first flag indicating whether the motion information merging mode is applied and the second flag indicating whether the motion vector prediction mode is applied are both 0, template-based motion estimation may be performed.

[0189] For each of the L0 and L1 directions, information indicating whether template-based motion estimation is applied can be signaled. That is, whether template-based motion estimation is applied in the L0 direction and whether template-based motion estimation is applied in the L1 direction can be determined independently. Accordingly, template-based motion estimation can be applied to one of the L0 and L1 directions, while another mode (e.g., motion information merging mode or motion vector prediction mode) can be applied to the other.

[0190] When template-based motion estimation is applied to both the L0 direction and the L1 direction, the prediction block of the current block can be generated based on a weighted sum operation of the L0 prediction block and the L1 prediction block. Alternatively, even when template-based motion estimation is applied to one of the L0 direction and the L1 direction, but another mode is applied to the other direction, the prediction block of the current block can be generated based on a weighted sum operation of the L0 prediction block and the L1 prediction block.

[0191] Alternatively, a template-based motion estimation method may be inserted as a motion information merging candidate in a motion information merging mode or a motion vector prediction candidate in a motion vector prediction mode. In this case, whether or not to apply a template-based motion estimation method may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate indicates a template-based motion estimation method.

[0192] Based on the two-way matching method, it is also possible to generate movement information of the current block.

[0193] Figure 10 is a diagram for explaining a motion estimation method based on a bilateral matching method.

[0194] The bilateral matching method can be performed only when the temporal order (i.e., POC) of the current picture exists between the temporal order of the L0 reference picture and the temporal order of the L1 reference picture.

[0195] When a bilateral matching method is applied, a search range can be set for each of the L0 reference picture and the L1 reference picture. At this time, an L0 reference picture index for identifying the L0 reference picture and an L1 reference picture index for identifying the L1 reference picture can be encoded and signaled, respectively.

[0196] As another example, only the L0 reference picture index may be encoded and signaled, and an L1 reference picture may be selected based on the distance between the current picture and the L0 reference picture (hereinafter referred to as the L0 POC difference). For example, among the L1 reference pictures included in the L1 reference picture list, an L1 reference picture having an absolute value of the distance from the current picture (hereinafter referred to as the L1 POC difference) equal to the absolute value of the distance between the current picture and the L0 reference picture may be selected. If there is no L1 reference picture having an L1 POC difference equal to the L0 POC difference, an L1 reference picture having an L1 POC difference most similar to the L0 POC difference may be selected among the L1 reference pictures.

[0197] At this time, among the L1 reference pictures, only L1 reference pictures that are temporally different from the L0 reference picture can be used for bilateral matching. For example, if the POC of the L0 reference picture is smaller than that of the current picture, one of the L1 reference pictures that has a POC larger than that of the current picture can be selected.

[0198] Conversely, one could also encode and signal only the L1 reference picture index, and select the L0 reference picture based on the distance between the current picture and the L1 reference picture.

[0199] Alternatively, a bilateral matching method may be performed using the L0 reference picture having the closest distance to the current picture among the L0 reference pictures and the L1 reference picture having the closest distance to the current picture among the L1 reference pictures.

[0200] Alternatively, a bilateral matching method may be performed using an L0 reference picture (e.g., index 0) assigned with a predefined index in the L0 reference picture list and an L1 reference picture (e.g., index 0) assigned with a predefined index in the L1 reference picture list.

[0201] Alternatively, the LX (X is 0 or 1) reference picture may be selected based on an explicitly signaled reference picture index, and the L|X-1| reference picture may be selected as a reference picture having the closest distance to the current picture among the L|X-1| reference pictures, or as a reference picture having a predefined index in the L|X-1| reference picture list.

[0202] As another example, L0 and / or L1 reference pictures can be selected based on motion information of neighboring blocks of the current block. For example, L0 and / or L1 reference pictures to be used for bilateral matching can be selected using the reference picture index of the neighboring block to the left or above the current block.

[0203] The search range can be set within a predetermined range from a collocated block within a reference picture.

[0204] As another example, the search range can be set based on initial motion information. This initial motion information can be derived from neighboring blocks of the current block. For example, the motion information of the left or top neighboring block of the current block can be set as the initial motion information of the current block.

[0205] When the bilateral matching method is applied, the L0 motion vector and the L1 motion vector are set to have opposite directions. This indicates that the signs of the L0 motion vector and the L1 motion vector have opposite signs. In addition, the size of the LX motion vector can be proportional to the distance between the current picture and the LX reference picture (i.e., the POC difference).

[0206] Thereafter, motion estimation can be performed using the cost between a reference block belonging to the search range of the L0 reference picture (hereinafter referred to as an L0 reference block) and a reference block belonging to the search range of the L1 reference picture (hereinafter referred to as an L1 reference block).

[0207] If an L0 reference block whose vector to the current block is (x, y) is selected, an L1 reference block located at a distance of (-Dx, -Dy) from the current block can be selected. Here, D can be determined by the ratio of the distance between the current picture and the L0 reference picture and the distance between the L1 reference picture and the current picture.

[0208] For example, in the example illustrated in FIG. 10, the absolute value of the distance between the current picture (T) and the L0 reference picture (T-1) and the absolute value of the distance between the current picture (T) and the L1 reference picture (T+1) are equal to each other. Accordingly, in the illustrated example, the L0 motion vector (x0, y0) and the L1 motion vector (x1, y1) have equal magnitudes but opposite distances. If an L1 reference picture with a POC of (T+2) were used, the L1 motion vector (x1, y1) would be set to (-2*x0, -2*y0).

[0209] Once the L0 reference block and L1 reference block with the optimal cost are selected, the L0 reference block and the L1 reference block can be set as the L0 prediction block and the L1 prediction block of the current block, respectively. Thereafter, the final prediction block of the current block can be generated through a weighted sum operation of the L0 reference block and the L1 reference block.

[0210] When the bilateral motion matching method is applied, the decoder can perform motion estimation in the same manner as the encoder. Accordingly, information indicating whether the bilateral motion matching method is applied can be explicitly encoded / decoded, while encoding / decoding of motion information such as motion vectors can be omitted. As previously explained, at least one of the L0 reference picture index or L1 reference picture index may also be explicitly encoded / decoded.

[0211] As another example, information indicating whether a bilateral matching method is applied may be explicitly encoded / decoded, and if a bilateral matching method is applied, the L0 motion vector or the L1 motion vector may be explicitly encoded and signaled. If the L0 motion vector is signaled, the L1 motion vector may be derived based on the POC difference between the current picture and the L0 reference picture and the POC difference between the current picture and the L1 reference picture. If the L1 motion vector is signaled, the L0 motion vector may be derived based on the POC difference between the current picture and the L0 reference picture and the POC difference between the current picture and the L1 reference picture. In this case, the encoder may explicitly encode the smaller one of the L0 motion vector and the L1 motion vector.

[0212] Information indicating whether a bilateral matching method has been applied may be a 1-bit flag. For example, a true flag (e.g., 1) may indicate that a bilateral matching method has been applied to the current block. A false flag (e.g., 0) may indicate that a bilateral matching method has not been applied to the current block. In this case, the current block may be subject to motion information merging mode or motion vector prediction mode.

[0213] Conversely, the bilateral matching method may be applied only when it is determined that neither the motion information merging mode nor the motion vector prediction mode is applied to the current block. For example, the bilateral matching method may be applied when both the first flag indicating whether the motion information merging mode is applied and the second flag indicating whether the motion vector prediction mode is applied are 0.

[0214] Alternatively, the bilateral matching method may be inserted as a motion information merging candidate in the motion information merging mode or as a motion vector prediction candidate in the motion vector prediction mode. In this case, whether the bilateral matching method is applied may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate indicates the bilateral matching method.

[0215] In the bidirectional matching method, it is exemplified that the temporal order of the current picture must exist between the temporal order of the L0 reference picture and the temporal order of the L1 reference picture. A unidirectional matching method, which is not subject to the constraints of the above bidirectional matching method, may also be applied to generate a prediction block of the current block. Specifically, in the unidirectional matching method, two reference pictures having a temporal order (i.e., POC) smaller than that of the current block or two reference pictures having a temporal order larger than that of the current block may be used. In this case, both reference pictures may be derived from the L0 reference picture list or the L1 reference picture list. Alternatively, one of the two reference pictures may be derived from the L0 reference picture list and the other may be derived from the L1 reference picture list.

[0216] Figure 11 is a diagram for explaining a motion estimation method based on a one-way matching method.

[0217] The unidirectional matching method can be performed based on two reference pictures having a POC smaller than that of the current picture (i.e., forward reference pictures) or two reference pictures having a POC larger than that of the current picture (i.e., backward reference pictures). In Fig. 11, motion estimation based on the unidirectional matching method is exemplified as being performed based on a first reference picture (T-1) and a second reference picture (T-2) having a POC smaller than that of the current picture (T).

[0218] At this time, a first reference picture index for identifying a first reference picture and a second reference picture index for identifying a second reference picture may be encoded and signaled, respectively. At this time, among the two reference pictures used in the unidirectional matching method, a reference picture having a smaller POC difference from the current picture may be set as the first reference picture. Accordingly, when the first reference picture is selected, only reference pictures included in the reference picture list having a larger POC difference from the current picture than the first reference picture may be set as the second reference picture. The second reference picture index may be set to point to an index of one of the rearranged reference pictures after rearranging reference pictures having the same temporal direction as the first reference picture and having a larger POC difference from the current picture than the first reference picture.

[0219] Conversely, the reference picture with a larger POC difference from the current picture among the two reference pictures can be set as the first reference picture. In this case, the second reference picture index can be set to point to the index of one of the rearranged reference pictures after rearranging the reference pictures that have the same temporal direction as the first reference picture and have a smaller POC difference from the current picture than the first reference picture.

[0220] Alternatively, a one-way matching method may be performed using a reference picture assigned with a predefined index within a reference picture list and a reference picture having the same temporal direction as the reference picture. For example, a reference picture having an index of 0 within the reference picture list may be set as the first reference picture, and a reference picture having the smallest index among the reference pictures having the same temporal direction as the first reference picture within the reference picture list may be selected as the second reference picture.

[0221] Both the first reference picture and the second reference picture can be selected from the L0 reference picture list or the L1 reference picture list. In Fig. 11, two L0 reference pictures are illustrated as being used in the unidirectional matching method. Alternatively, the first reference picture may be selected from the L0 reference picture list, and the second reference picture may be selected from the L1 reference picture list.

[0222] Information indicating whether the first reference picture and / or the second reference picture belongs to the L0 reference picture list or the L1 reference picture list may be additionally encoded / decoded.

[0223] Alternatively, one-way matching can be performed using one of the L0 reference picture list and the L1 reference picture list, whichever is set as default. Alternatively, two reference pictures can be selected from the L0 reference picture list and the L1 reference picture list, whichever has a larger number of reference pictures.

[0224] Afterwards, a search range within the first reference picture and the second reference picture can be set.

[0225] The search range can be set within a predetermined range from a collocated block within a reference picture.

[0226] As another example, the search range can be set based on initial motion information. This initial motion information can be derived from neighboring blocks of the current block. For example, the motion information of the left or top neighboring block of the current block can be set as the initial motion information of the current block.

[0227] Thereafter, motion estimation can be performed using the cost between the first reference block belonging to the search range of the first reference picture and the second reference block belonging to the search range of the second reference picture.

[0228] At this time, under the unidirectional matching method, the size of the motion vector should be set to increase in proportion to the distance between the current picture and the reference picture. Specifically, if a first reference block whose vector with the current picture is (x, y) is selected, the second reference block should be spaced apart from the current block by (Dx, Dy). Here, D can be determined by the ratio of the distance between the current picture and the first reference picture and the distance between the current picture and the second reference picture.

[0229] For example, in the example of FIG. 11, the distance between the current picture and the first reference picture (i.e., the POC difference) is 1, and the distance between the current picture and the second reference picture (i.e., the POC difference) is 2. Accordingly, when the first motion vector for the first reference block in the first reference picture is (x0, y0), the second motion vector (x1, y1) for the second reference block in the second reference picture can be set to (2x0, 2y0).

[0230] Once the first and second reference blocks with optimal costs are selected, the first and second reference blocks can be set as the first and second prediction blocks of the current block, respectively. Thereafter, a weighted sum operation of the first and second prediction blocks can be performed to generate the final prediction block of the current block.

[0231] When a unidirectional motion matching method is applied, the decoder can perform motion estimation in the same manner as the encoder. Accordingly, information indicating whether a unidirectional motion matching method is applied can be explicitly encoded / decoded, while encoding / decoding of motion information such as motion vectors can be omitted. As previously explained, at least one of the first reference picture index or the second reference picture index may also be explicitly encoded / decoded.

[0232] As another example, information indicating whether a unidirectional matching method is applied may be explicitly encoded / decoded, and if a unidirectional matching method is applied, the first motion vector or the second motion vector may be explicitly encoded and signaled. If the first motion vector is signaled, the second motion vector may be derived based on the POC difference between the current picture and the first reference picture and the POC difference between the current picture and the second reference picture. If the second motion vector is signaled, the first motion vector may be derived based on the POC difference between the current picture and the first reference picture and the POC difference between the current picture and the second reference picture. In this case, the encoder may explicitly encode a smaller one of the first motion vector and the second motion vector.

[0233] Information indicating whether a unidirectional matching method has been applied may be a 1-bit flag. For example, a true flag (e.g., 1) may indicate that a unidirectional matching method has been applied to the current block. A false flag (e.g., 0) may indicate that a unidirectional matching method has not been applied to the current block. In this case, the current block may be subject to motion information merging mode or motion vector prediction mode.

[0234] Conversely, the one-way matching method may be applied only when it is determined that neither the motion information merging mode nor the motion vector prediction mode is applied to the current block. For example, the one-way matching method may be applied when the first flag indicating whether the motion information merging mode is applied and the second flag indicating whether the motion vector prediction mode is applied are both 0.

[0235] Alternatively, the unidirectional matching method may be inserted as a motion information merging candidate in the motion information merging mode or as a motion vector prediction candidate in the motion vector prediction mode. In this case, whether the unidirectional matching method is applied may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate indicates the unidirectional matching method.

[0236] By adjusting the precision of motion vectors, it is also possible to detect the inter-frame movement of an object. Specifically, the position of each pixel within a picture is specified as an integer. However, the inter-frame movement of an object may not be expressed as an integer position.

[0237] Taking this into account, we can search for motion vectors in fractional pixel units by performing interpolation on the reference picture.

[0238] Figures 12 and 13 illustrate examples in which prediction blocks are generated according to the precision of a motion vector.

[0239] Figure 12 shows the location of the current block within the current picture, and Figure 13 shows an example in which a prediction block is obtained according to a motion vector.

[0240] Specifically, (a) of Fig. 13 shows an example in which the motion vector precision is in integer pixel units, and (b) and (c) of Fig. 13 show examples in which the motion vector precision is in 1 / 2 pixel units and 1 / 4 pixel units, respectively.

[0241] The motion vector precision can also be set in smaller units than those shown. For example, the motion vector precision can be set in units of 1 / 8 pixel, 1 / 16 pixel, or 1 / 32 pixel.

[0242] When the motion vector of the current block is expressed in integer units, a reference block composed of integer position samples can be set as a prediction block of the current block, as in the example illustrated in (a) of Fig. 13.

[0243] On the other hand, if the motion vector of the current block is expressed in fractional units, a reference block composed of fractional position samples can be set as the prediction block of the current block, as in the examples shown in (b) and (c) of Fig. 13. In this case, the fractional position samples within the reference block can be generated by interpolating integer position samples. The interpolation filter can have a size of 4 or 8 taps.

[0244] As another example, to reduce complexity, fractional position samples can be generated via linear interpolation using only integer position samples adjacent to the fractional position.

[0245] Information indicating the motion vector precision of the current block may be encoded and signaled. For example, after assigning different indices to each of a plurality of motion vector precision candidates, the index of the motion vector precision candidate corresponding to the motion vector precision of the current block may be encoded and signaled.

[0246] At this time, the number and / or types of available motion vector candidates can be determined based on at least one of the size of the current block, the shape of the current block, the reference picture, or the motion compensation model. Here, the motion compensation model can include at least one of a translation model, a zooming model, or a rotation model. A motion compensation model that combines a translation model with at least one of a zooming model or a rotation model can be referred to as an affine model.

[0247] An index indicating one of the available motion vector candidates for the current block may be encoded. Depending on the number of motion vector candidates available for the current block, the maximum number of bits required to encode the index may be determined.

[0248] By adjusting the precision of the motion vector, the motion vector can be searched more precisely, and thus the prediction accuracy for the current block can be improved.

[0249] Meanwhile, the motion vector expressed in fractional positions can be scaled up to integers and encoded.

[0250] Compensation for the motion of an object may be performed based on at least one of a translational model for compensating for linear motion of the object (e.g., motion in the horizontal and / or vertical direction), a zooming model for compensating for changes in the size of the object, and a rotational model for compensating for rotational motion of the object. Here, zooming may refer to enlargement or reduction in size.

[0251] Figure 14 shows an example in which motion compensation based on a translational model and a zooming model is performed for the current block.

[0252] For convenience of explanation, the current block is assumed to have a size of 4x4, as illustrated in Fig. 12.

[0253] In Figure 14, the variable α represents a size adjustment parameter. The size of the reference block can be derived by multiplying the size of the current block by the variable α.

[0254] A size adjustment parameter α less than 1 indicates that the reference block is smaller than the current block, and a size adjustment parameter α greater than 1 indicates that the reference block is larger than the current block.

[0255] Figures 14 (a) and (b) show examples when the size adjustment parameter α is less than 1, and Figure 14 (c) shows examples when the size adjustment parameter α is greater than 1.

[0256] Based on the motion vector of the current block, the upper left position of the reference block can be determined. Specifically, the upper left position of the reference block can be set as a position spaced apart by the motion vector from the position corresponding to the upper left sample of the current block within the reference picture. Thereafter, based on the resizing parameter, a reference block whose width and height are α times the width and height of the current block, respectively, can be set. Fractional position samples within the reference block can be generated by interpolating integer position samples.

[0257] A reference block derived by a motion vector and a scale parameter can be set as a prediction block of the current block.

[0258] Meanwhile, information about the scaling parameter α may be encoded and signaled. Specifically, a different index may be assigned to each of a plurality of scaling parameter candidates, and an index specifying the scaling parameter candidate applied to the current block may be encoded and signaled.

[0259] Alternatively, the resizing parameters of the current block can be derived based on the resizing parameters of neighboring blocks. For example, the resizing parameters of neighboring blocks at a predefined location can be set as the resizing parameters of the current block.

[0260] Alternatively, when sequentially searching multiple neighboring blocks, the size adjustment parameter of the first searched available neighboring block can be set as the size adjustment parameter of the current block.

[0261] Alternatively, the resizing parameter of a neighboring block may be set as a resizing parameter candidate. In this case, a plurality of neighboring blocks may be sequentially searched to generate a resizing parameter candidate list including a plurality of resizing parameter candidates. One of the plurality of resizing parameter candidates included in the plurality of resizing parameter candidate lists may be set as the resizing parameter of the current block. In this case, an index indicating a candidate among the plurality of resizing parameter candidates that is identical to the resizing parameter of the current block may be encoded and signaled.

[0262] Meanwhile, the neighboring blocks used to derive the size adjustment parameters of the current block may include at least one of an upper neighboring block, a left neighboring block, an upper-left neighboring block, an upper-right neighboring block, or a lower-left neighboring block.

[0263] Figure 15 shows an example in which motion compensation based on a translational model and a rotational model is performed for the current block.

[0264] For convenience of explanation, the current block is assumed to have a size of 4x4, as illustrated in Fig. 12.

[0265] First, as illustrated in (a) of Fig. 15, the location of a temporary block within a reference picture can be determined based on the motion vector of the current block. Specifically, a block location that is spaced apart by the motion vector from the location corresponding to the upper left sample of the current block within the reference picture can be determined as the upper left sample.

[0266] Thereafter, the temporary block can be rotated, as in the example illustrated in (b) of Fig. 15. The block at the rotated position can be set as a reference block, and the reference block can be set as a prediction block of the current block.

[0267] Meanwhile, a rotation matrix can be used to rotate a temporary block specified by a motion vector. That is, the prediction sample for the current block can be set to a sample at a position obtained by applying a rotation matrix to the sample position within the temporary block.

[0268] Mathematical expression 1 represents the rotation matrix.

[0269]

[0270] In the above mathematical expression 1, (pos_x, pos_y) represents the position of a sample within a temporary block. That is, (pos_x, pos_y) can be derived by adding a motion vector to the position of the target sample to be predicted within the current block.

[0271] (pos_x', pos_y') represents the rotated position from the position of the sample within the temporary block, and θ represents the rotation angle.

[0272] The sample value at the position (pos_x', pos_y') within the reference picture can be set as the value of the predicted sample for the position of the target sample. If the position (pos_x', pos_y') is a fractional position, the sample at that position can be generated by interpolating integer position samples.

[0273] Meanwhile, information indicating the rotation angle θ can be encoded and signaled. For example, after assigning a different index to each of a plurality of rotation angle candidates, the index of the rotation angle candidate corresponding to the rotation angle of the current block can be encoded and signaled.

[0274] Alternatively, the rotation angle of the current block can be derived based on the rotation angle of a neighboring block. For example, the rotation angle of a neighboring block at a predefined position can be set to the rotation angle of the current block.

[0275] Alternatively, when sequentially searching multiple neighboring blocks, the rotation angle of the first searched available neighboring block can be set to the rotation angle of the current block.

[0276] Alternatively, the rotation angle of a neighboring block may be set as a rotation angle candidate. In this case, a rotation angle candidate list including a plurality of rotation angle candidates may be generated by sequentially searching a plurality of neighboring blocks. One of the plurality of rotation angle candidates included in the plurality of rotation angle candidate lists may be set as the rotation angle of the current block. In this case, an index indicating a candidate among the plurality of rotation angle candidates having the same rotation angle as the current block may be encoded and signaled.

[0277] Meanwhile, the neighboring block used to derive the rotation angle of the current block may include at least one of an upper neighboring block, a left neighboring block, an upper left neighboring block, an upper right neighboring block, or a lower left neighboring block.

[0278] Although not shown, motion compensation for the current block can also be performed by simultaneously applying the translation model, zoom model, and rotation model.

[0279] Meanwhile, the motion vector precision for the current block or the number and / or types of motion vector precision candidates available for the current block may be determined differently depending on the motion compensation model.

[0280] For example, the number and / or type of motion vector precision candidates available for the current block may differ between cases where only the translational model is applied and cases where at least one of the zooming model or the rotational model is applied.

[0281] For example, if a translational model is applied to the current block, candidates larger than 1 / 4 pixel unit may be available for the current block. Conversely, if at least one zooming model or rotational model is additionally applied to the current block along with the translational model, candidates larger than 1 / 16 pixel unit may be available for the current block.

[0282] Alternatively, if a translational model is applied to the current block, the motion vector precision of the current block may be set to 1 / 4 pixel units. On the other hand, if at least one of a zooming model or a rotational model is additionally applied to the current block along with the translational model, the motion vector precision of the current block may be set to 1 / 16 pixel units.

[0283] Meanwhile, the encoder and decoder may have pre-stored the motion vector precisions or motion vector precision candidates available for each motion compensation model. Alternatively, information indicating the motion vector precisions or motion vector precision candidates available for each motion compensation model may be encoded and signaled through the upper header.

[0284] Motion compensation can be performed on an affine model, in which a zooming model and / or a rotation model are added to a translational model, using the motion vectors of the control points. Here, the control points may correspond to corners of the current block. For example, to perform motion compensation based on the affine model, at least one of the motion vector of the upper left corner, the motion vector of the upper right corner, or the motion vector of the lower left corner may be used.

[0285] Hereinafter, the motion vector of a control point will be referred to as a control point motion vector.

[0286] Figures 16 and 17 illustrate examples of generating a prediction block for a current block using control point motion vectors.

[0287] For convenience of explanation, the current block is assumed to have a size of 4x4, as illustrated in Fig. 12.

[0288] In (a) and (b) of FIG. 16, it is exemplified that a prediction block for the current block is derived by a motion vector of a first control point corresponding to the upper left corner of the current block (first control point motion vector, A) and a motion vector of a second control point corresponding to the upper right corner of the current block (second control point motion vector, B).

[0289] In addition to the example shown, it is also possible to derive a prediction block of the current block by additionally using the motion vector of the lower left corner or by using the motion vector of the lower left corner instead of the upper right corner.

[0290] Figure 18 shows an example of generating a prediction block for the current block using three control point motion vectors.

[0291] In (a) and (b) of FIG. 18, it is exemplified that a prediction block for the current block is derived by a motion vector of a first control point corresponding to the upper left corner of the current block (first control point motion vector, A), a motion vector of a second control point corresponding to the upper right corner of the current block (second control point motion vector, B), and a motion vector of a third control point corresponding to the lower left corner of the current block (third control point motion vector, C).

[0292] As in the examples illustrated in FIGS. 16 to 18, translational, zooming, and rotational motion compensation for the current block can be performed using two control point motion vectors or three control point motion vectors.

[0293] Information indicating the number of control point motion vectors may be encoded and signaled. The information may be signaled on a block-by-block basis. For example, the information may indicate whether two or three control point motion vectors are used in the current block.

[0294] Alternatively, the number of control point motion vectors can be adaptively determined based on at least one of the size or shape of the current block.

[0295] Alternatively, if the control point motion vectors of the current block are derived from neighboring blocks, the number of control point motion vectors for the current block may be set equal to the number of control point motion vectors of the neighboring blocks.

[0296] Using control point motion vectors, a motion vector for each sample within the current block can be derived. Equation 2 represents a formula for deriving a motion vector for each sample using two control point motion vectors.

[0297]

[0298] In the above mathematical expression 2, (mv x , mv y ) represents the motion vector at the (x, y) position within the current block. (mv Ax , mv Ay ) represents the first control point motion vector (A), and (mv Bx , mv By ) represents the second control point motion vector (B). W represents the width of the current block.

[0299] When three control point motion vectors are used, a motion vector per sample can be derived by the following mathematical expression 3.

[0300]

[0301] In the above mathematical expression 3, (mv Cx , mv Cy ) represents the third control point motion vector (C).

[0302] Once a motion vector is derived for each sample, motion compensation can be performed for each sample, as in the example illustrated in Fig. 17. Specifically, a reference sample indicated by the motion vector of the prediction target sample can be set as a prediction sample for the prediction target sample.

[0303] Meanwhile, if the motion vector of the prediction target sample is expressed in fractional units, integer position samples can be interpolated to generate fractional position samples, and the generated fractional position samples can be set as prediction samples for the prediction target sample.

[0304] At this time, the precision of the motion vector for each sample may be different. For example, the motion vector for the first prediction target sample may be derived in units of 1 / 2 pixels, while the motion vector for the second prediction target sample may be derived in units of 1 / 4 pixels.

[0305] In this case, fractional position samples can be generated according to the motion vector precision for each prediction target sample. Alternatively, the motion vector of the prediction target sample can be adjusted according to the reference motion vector precision, and then a prediction sample for the prediction target sample can be derived based on the adjusted motion vector. For example, if the reference motion vector precision is 1 / 2, the motion vector for the second prediction target sample can be adjusted in units of 1 / 4 pixels.

[0306] The reference motion vector precision can be determined on a block-by-block basis. Alternatively, the precision of control point motion vectors can be set as the reference motion vector precision. Alternatively, the reference motion vector precision can be predefined in the encoder and decoder.

[0307] As another example, to reduce complexity, motion vectors can be derived on a sub-block basis.

[0308] Figure 19 shows an example in which a motion vector is derived in sub-block units.

[0309] The size and / or shape of a sub-block may be predefined in the encoder and decoder. For example, a sub-block may be a square block of size 2x2 or 4x4.

[0310] Alternatively, the size and / or shape of the sub-block may be adaptively determined based on the size and / or shape of the current block. For example, if the current block is square, the sub-block may also be square. Conversely, if the current block is non-square, the sub-block may also be non-square.

[0311] Alternatively, information regarding at least one of the division method or division shape of the current block may be explicitly encoded and signaled. For example, information regarding at least one of the size of a sub-block, the shape of a sub-block, the position of a division line dividing the current block, or the number of division lines may be explicitly encoded and signaled. The information may be encoded and signaled on a block-by-block basis, or may be encoded and signaled via an upper header.

[0312] In Fig. 19, it is assumed that the sub-block is a square block of size 2x2.

[0313] A motion vector of a sub-block can be derived using coordinates of a predefined position within a sub-block. Here, the predefined position can be one of the positions of the upper left sample, the upper right sample, the lower left sample, the lower right sample, or the center position within the sub-block.

[0314] By substituting the coordinates of a predefined position within a sub-block into (x, y) in Equation 2, the motion vector of the sub-block can be derived.

[0315] As in the example described above, motion vectors can be derived in sub-block units based on the affine motion model.

[0316] Meanwhile, motion vectors can also be derived for each sub-block using collocated pictures. As described above, deriving motion vectors for each sub-block using collocated pictures can be called SbTMVP (Sub-block Temporal Motion Vector Prediction).

[0317] A collocated picture may be one of the reference pictures included in a reference picture list. For example, a picture with an index of 0 in the reference picture list may be selected as a collocated picture.

[0318] Alternatively, information indicating the index of a reference picture to be set as a collocated picture within the reference picture list may be explicitly encoded and signaled.

[0319] Figures 20 and 21 illustrate examples in which motion vectors are derived in units of sub-blocks within the current block when SbTMVP is applied.

[0320] The size and / or shape of the sub-block may be predefined in the encoder and decoder.

[0321] Alternatively, the size and / or shape of the sub-block may be adaptively determined based on the size and / or shape of the current block. For example, if at least one of the width or height of the current block is greater than a threshold, the size of the sub-block may be set to 8x8. Otherwise, the size of the sub-block may be set to 4x4.

[0322] Alternatively, information indicating the size and / or shape of the sub-block may be explicitly encoded and signaled.

[0323] In the example shown in Fig. 20, it is assumed that the size of the current block is 16x16 and the size of the sub-block is 4x4.

[0324] When SbTMVP is applied, an initial motion vector of the current block can be derived. The initial motion vector can be derived based on at least one of a motion vector prediction list or a motion information merge list. For example, an index indicating one of the motion vector prediction candidates included in the motion vector prediction list can be encoded and signaled. The initial motion vector can be derived by adding a motion vector differential value to the motion vector prediction candidate indicated by the index. Meanwhile, the motion vector differential value can also be explicitly encoded and signaled.

[0325] Alternatively, encoding of the index may be omitted, and a motion vector prediction candidate having a predefined index in the motion vector prediction list may be set as a prediction value for the initial motion vector. Here, the motion vector prediction candidate having a predefined index may be a motion vector prediction candidate having an index of 0 or a motion vector prediction candidate having the largest index.

[0326] Alternatively, an index indicating one of the motion information merging candidates included in the motion information merging list may be encoded and signaled. The initial motion vector may be set to be identical to the motion vector of the motion information merging candidate indicated by the index.

[0327] Alternatively, encoding of the index can be omitted, and the initial motion vector can be derived based on a motion information merge candidate having a predefined index in the motion information merge list. Here, the motion information merge candidate having a predefined index can be a motion information merge candidate having an index of 0 or a motion information merge candidate having the largest index.

[0328] Alternatively, the initial motion vector can be derived using the motion vector of a neighboring block at a predefined position. Here, the neighboring block at the predefined position can be a left neighboring block or an upper neighboring block.

[0329] The motion vector of a neighboring block at a predefined position can be set as a predicted value of the initial motion vector, and the difference value can be added to the predicted value to derive the initial motion vector.

[0330] Alternatively, the motion vector of a neighboring block at a predefined position can be set as the initial motion vector.

[0331] Alternatively, the initial motion vector can be derived using a template-based motion estimation method (i.e., template matching method) or bilateral matching.

[0332] The precision of the initial motion vector may be predefined in the encoder and decoder. For example, the precision of the initial motion vector may be fixed in integer pixel units.

[0333] Alternatively, information indicating the precision of the initial motion vector may be explicitly encoded and signaled. The information may be an index indicating one of a plurality of motion vector precision candidates.

[0334] When deriving an initial motion vector using a motion vector prediction candidate, the motion vector prediction candidates can be derived based on the motion vector precision of the initial motion vector. That is, the motion vector prediction candidate can be adjusted according to the motion vector precision of the initial motion vector, and then the adjusted initial motion vector prediction candidate can be inserted into the motion vector prediction list.

[0335] When deriving an initial motion vector using a motion information merge candidate, the motion information merge candidates can be derived based on the motion vector precision of the initial motion vector. That is, the motion information merge candidate can be adjusted according to the motion vector precision of the initial motion vector, and then the adjusted initial motion information merge candidate can be inserted into the motion information merge list.

[0336] Meanwhile, among the motion information merge candidates included in the motion information merge list, only those candidates whose reference picture is identical to the collocated picture of the current block can be used to derive the initial motion vector. That is, if the reference picture of a motion information merge candidate is different from the collocated picture of the current block, the initial motion vector may not be derived from the motion information merge candidate.

[0337] If there are multiple candidates among motion information merging candidates whose reference pictures are identical to the collocated pictures of the current block, an index indicating one of the multiple candidates may be encoded and signaled. Alternatively, if there are multiple candidates among motion information merging candidates whose reference pictures are identical to the collocated pictures of the current block, the initial motion vector may be derived from the candidate with the smallest index or the candidate with the largest index among the multiple candidates.

[0338] If a motion information merging candidate has both motion information in the L0 direction and motion information in the L1 direction, one of the motion information in the L0 direction and the motion information in the L1 direction can be selected according to a preset priority, and an initial motion vector can be derived from the selected motion information.

[0339] The priority may be determined based on at least one of the magnitude of the motion vector of the motion merging candidate, the index of the reference picture of the motion merging candidate, or whether the reference picture of the motion merging candidate is the same as the collocated picture.

[0340] Alternatively, it can be set to always derive the initial motion vector based on the motion information in the L0 direction.

[0341] When the initial motion vector is derived based on a template matching method, motion estimation can be performed according to the precision of the initial motion vector. For example, if the precision of the initial motion vector is integer pixel units, motion estimation based on template matching can also be performed only at integer positions.

[0342] Similarly, when the initial motion vector is derived based on bilateral matching, motion estimation can be performed according to the precision of the initial motion vector.

[0343] Meanwhile, as a result of the bilateral matching, a motion vector in the L0 direction (L0 motion vector) and a motion vector in the L1 direction (L1 motion vector) are derived. In this case, one of the L0 motion vector and the L1 motion vector can be set as the initial motion vector according to the preset priority.

[0344] Alternatively, it can be set to always derive the initial motion vector based on the motion information in the L0 direction.

[0345] Alternatively, information indicating which of the L0 motion vector and the L1 motion vector is set as the initial motion vector may be encoded and signaled.

[0346] Once the initial motion vector is derived, the position of the collocated block within the collocated block can be determined using the initial motion vector. For example, a block located at a position spaced apart by the initial motion vector from a position corresponding to the current block within the reference picture can be set as the collocated block. At this time, the position of the collocated block can be determined based on a predefined position within the current block. Here, the predefined position can be an upper left position, an upper right position, a lower left position, a lower right position, or a center position.

[0347] Depending on the division method of the current block, the collocated block can be divided into multiple collocated sub-blocks. In addition, the motion vector of each collocated sub-block within the collocated block can be set to the motion vector of each sub-block within the current block.

[0348] As another example, the initial motion vector can be used to determine the positions of the collocated sub-blocks corresponding to each sub-block within the current block within the collocated picture. In this case, the positions of the collocated sub-blocks can be derived based on predefined positions within the sub-blocks. Here, the predefined positions can be the upper left position, the upper right position, the lower left position, the lower right position, or the center position.

[0349] Thereafter, the motion vector of the collocated sub-block corresponding to the sub-block can be set as the motion vector of the sub-block. Specifically, the motion vector stored at a position corresponding to a predefined position within the collocated sub-block can be set as the motion vector of the sub-block.

[0350] Meanwhile, if the motion information of the collocated sub-block is unavailable, a predefined motion vector can be set as the motion vector of the sub-block. Here, the predefined motion vector can be a zero vector (i.e., (0, 0)) or an initial motion vector.

[0351] Alternatively, if the motion information of the collocated sub-block corresponding to the sub-block is not available, the motion vector of the sub-block may be derived from another location within the collocated sub-block.

[0352] Specifically, if a position corresponding to a predefined position within a collocated sub-block is encoded using intra prediction, then there is no motion vector at that position. For example, assuming that the predefined position is a central position (e.g., c10 in FIG. 21), if no motion vector is stored at the central position, the motion vector of the sub-block cannot be derived.

[0353] In this case, the motion vector of the sub-block can be derived based on the motion vector stored at a location different from the central location. Specifically, the motion vector of the sub-block can be derived based on the motion vector stored at a location adjacent to the central location (e.g., the upper adjacent location c6, the left adjacent location c9, or the upper left adjacent location c5).

[0354] Alternatively, if the center position is not available, the samples within the collocated sub-blocks may be searched according to the scan order, and the first available motion vector found may be set as the motion vector of the sub-block. Here, the scan order may be a horizontal scan, a vertical scan, a diagonal scan, or a raster scan.

[0355] Alternatively, if motion information of a collocated sub-block is unavailable, the motion vector of the sub-block may be set as the motion vector of the collocated block. For example, a motion vector stored at a position corresponding to a predefined position within the current block within the collocated block may be set as the motion vector of the sub-block.

[0356] As in the example described above, motion vectors can be derived for each sub-block using the affine motion model or SbTMVP. When motion vectors are derived for each sub-block, motion compensation can be performed for each sub-block based on the motion vector of each sub-block.

[0357] By performing motion compensation for each sub-block, a prediction block for the current block can be obtained. That is, the prediction block may be composed of prediction samples of each of the sub-blocks.

[0358] When detecting motion between frames, the precision of the motion vector can be adjusted. Specifically, the location of each sample within a picture is defined as an integer position. However, the position reflecting the motion may be a decimal position, not an integer position.

[0359] Taking this into account, we can search for motion vectors more precisely through reference picture interpolation.

[0360] Figures 22 and 23 are diagrams showing examples in which prediction blocks are derived according to motion vector precision.

[0361] Figure 22 shows the position of the current block within the current picture, and Figure 23 shows the position of the reference block according to the motion vector precision.

[0362] As in the examples illustrated in FIGS. 22 and 23, the motion vector of the current block can be defined as the distance from the sample corresponding to the upper left position of the current block in the reference picture to the sample corresponding to the upper left position of the reference block in the reference picture.

[0363] Figure 23 (a) illustrates a case where the motion vector precision of the current block is an integer pel, Figure 23 (b) illustrates a case where the motion vector precision of the current block is 1 / 2 pel, and Figure 23 (c) illustrates a case where the motion vector precision of the current block is 1 / 4 pel.

[0364] In Figure 23, motion vectors are expressed up to 1 / 4 vector precision, but motion vectors can also be expressed more precisely, such as 1 / 8, 1 / 16, or 1 / 32.

[0365] Meanwhile, information indicating the motion vector precision of the current block may be encoded and signaled. For example, the information may be an index identifying one of the motion vector precision candidates. Specifically, each of the motion vector precision candidates may be assigned a different index, and the information may indicate the index of the motion vector precision candidate applied to the current block.

[0366] By adjusting the precision of the motion vector used for inter-screen prediction, more precise motion vector search can be achieved. If the reference block indicated by the motion vector exists at a real number location, the samples at the real number location can be generated using samples at integer locations and an interpolation filter. Furthermore, motion vectors expressed as real numbers can be scaled up to integers and then encoded / decoded.

[0367] In this way, motion vectors (MV), motion vector predictors (MVPs), and motion vector differences (MVDs) can be encoded / decoded as integer values ​​through integerization. Specifically, motion vectors, motion vector predictors, and / or motion vector differences can be integerized based on motion vector precision.

[0368] For example, if the motion vector precision is 1 / N, integerization can be performed by multiplying the motion vector difference MVD by N. For example, if the motion vector difference MVD is (4 / 16, 8 / 16), the motion vector difference MVD can be integerized by multiplying by 16. That is, the integerized motion vector difference MVD can be expressed as (4, 8).

[0369] Based on the motion vector precision, the actual MVD can be derived from the integerized MVD. For example, if the motion vector precision is 1 / N, the actual MVD can be derived by dividing the integerized MVD by N. For example, if the integerized MVD is (4, 8) and the motion vector precision is 1 / 8, the actual MVD can be (4 / 8, 8 / 8). Alternatively, if the integerized MVD is (4, 8) and the motion vector precision is 1 / 4, the actual MVD can be (4 / 4, 8 / 4).

[0370] Depending on the motion vector precision, the range of expression of the integerized MVD may vary. For example, assume that the motion vector difference MVD is (4 / 16, 8 / 16) (i.e., (1 / 4, 2 / 4)). If the motion vector precision is 1 / 16, the integerized MVD is derived as (4, 8). On the other hand, if the motion vector precision is 1 / 4, the integerized MVD is derived as (1, 2).

[0371] Comparing the two cases above, if the motion vector precision is adjusted from 1 / 16 to 1 / 4, the integerized MVD value can be reduced from (4, 8) to (1, 2).

[0372] As a result, depending on the motion vector precision, the number of bits required to encode / decode the integerized motion vector difference MVD may vary. Accordingly, when encode / decode the motion vector difference MVD, a motion vector precision that can minimize the number of bins can be selected. Then, based on the selected motion vector precision, the motion vector difference MVD can be integerized, and the integerized motion vector difference MVD can be encode / decoded. In addition, information regarding the motion vector precision can be additionally encode / decode.

[0373] In the decoder, the actual MVD can be reconstructed from the decoded MVD based on the motion vector precision. Then, the motion vector MV can be derived by combining the reconstructed MVD and the motion vector prediction value MVP.

[0374] As above, the method of adjusting the value of the motion vector difference MVD to be encoded / decoded based on the motion vector precision is called the AMVR (Adaptive Motion Vector Resolution) method.

[0375] Figures 24 and 25 are diagrams for explaining the process of encoding and decoding a motion vector difference value, respectively, when the AMVR method is applied.

[0376] For convenience of explanation, it is assumed that the motion vector and motion vector difference are expressed in units of 1 / 16 before integerization is performed, and 1 / 16 is expressed as the original motion vector precision.

[0377] The motion vector difference value MVD can be derived by differentiating the motion vector prediction value MVP from the motion vector MV (S2410).

[0378] The motion vector difference MVD may be composed of a horizontal direction component (i.e., x-axis component) and a vertical direction component (i.e., y-axis component).

[0379] If the motion vector difference value is 0, that is, if both the horizontal direction component and the vertical direction component are 0, the value of the motion vector difference value MVD to be encoded becomes 0 regardless of the motion vector precision. Therefore, if the motion vector difference value MVD is 0, encoding of AMVR-related information can be omitted (S2420).

[0380] On the other hand, if the motion vector difference is not 0, i.e., if at least one of the horizontal direction component and the vertical direction component is not 0, the motion vector precision can be determined (S2430). Meanwhile, the motion vector precision can be encoded as AMVR-related information.

[0381] Information related to AMVR may include at least one of a flag (e.g., amvr_flag) indicating whether the AMVR method is applied to the current block, and an index (e.g., amvr_prec_idx) indicating one of a plurality of motion precision candidates when the AMVR method is applied.

[0382] If the AMVR method is not applied to the current block, the motion vector precision can be set to a default value. In this case, amvr_flag can be encoded as a value of 0. Meanwhile, the default value can be 1, 1 / 2, 1 / 4, 1 / 8, or 1 / 16.

[0383] When the AMVR method is applied to the current block, an index indicating one of multiple motion vector precision candidates, i.e., amvr_prec_idx, may be additionally decoded. In this case, amvr_flag may be encoded with a value of 1, and amvr_prec_idx may be encoded with a value from 0 to (n-1), where n represents the number of motion vector precision candidates. For example, the multiple motion vector precision candidates may include at least one of 4, 2, 1, 1 / 2, 1 / 4, 1 / 8, or 1 / 16. Meanwhile, the default value may not be set to the multiple motion vector precision candidates indicated by the index. That is, when the motion vector precision of the current block is the default value, it is encoded and signaled with the value of amvr_flag, 0, and encoding of amvr_prec_idx may be omitted.

[0384] In the encoder, the optimal motion vector precision can be determined by performing Rate Distortion Optimization (RDO) for each combination of amvr_flag and amvr_prec_idx. That is, the combination with the optimal cost can be selected by performing RDO for the following cases.

[0385] 1) If amvr_flag is 0

[0386] 2) If amvr_flag is 1 and amvr_prec_idx is 0

[0387] 3) If amvr_flag is 1 and amvr_prec_idx is 1

[0388] 4) If amvr_flag is 1 and amvr_prec_idx is 2

[0389] Depending on the motion vector precision of the current block, a variable for scaling the motion vector difference, i.e., a scaling parameter, can be set. As an example, Table 1 illustrates the values ​​of the variable amvrshift according to the motion vector precision.

[0390] amvr_flagamvr_prec_idxamvrshift0 (1 / 4)-210 (1 / 2)311 (1-pel)412 (4-pel)6

[0391] If the finest motion vector precision that can be applied to the current block is 1 / 16, the motion vector precision can be expressed as in the following mathematical expression 4.

[0392]

[0393] As shown in Table 1, when the value of amvr_flag is 0, the variable amvrshift is set to 2. This indicates that the motion vector precision is 1 / 4 according to Equation 4.

[0394] When the value of amvr_flag is 1, the variable amvrshft can be determined according to the value of amvr_prec_idx. For example, when amvr_prec_idx is 1, the variable amvrshift is set to 4. This indicates that the motion vector precision is 1 according to Equation 4.

[0395] In the encoder, the motion vector difference value MVD can be scaled down and encoded using the variable amvrshift according to the motion vector precision. As an example, mathematical expression 5 shows an example of performing a scale down operation on the motion vector difference value MVD.

[0396]

[0397] In the above mathematical expression 5, MVD_x represents the horizontal component of the motion vector difference, and MVD_y represents the vertical component of the motion vector difference. MVD'_x and MVD'_y represent the results of the scale down operation.

[0398] The encoder can encode motion vector difference values ​​and AMVR information with changed precision (S2440).

[0399] In the decoder, the motion vector difference value MVD can be decoded (S2510).

[0400] If the motion vector difference value is 0, decoding of AMVR-related information is omitted, and the motion vector MV of the current block can be set to be the same as the motion vector prediction value (S2520).

[0401] On the other hand, if the motion vector difference is not 0, i.e., if at least one of the horizontal direction component and the vertical direction component is not 0, information related to AMVR can be additionally decoded (S2530).

[0402] Based on the AMVR information, a variable amvrshift can be derived for scaling the motion vector difference. For example, as shown in Table 1, a variable amvrshfit can be derived based on amvr_flag and / or amvr_prec_idx.

[0403] Thereafter, by using the variable amvrshift, the decoded MVD can be scaled up to obtain a motion vector difference MVD restored to its original precision (S2540). Mathematical expression 6 shows an example in which a scale-up operation is applied to the decoded MVD.

[0404]

[0405] In mathematical expression 6, MVD' represents the decoded motion vector difference. MVD represents the motion vector difference restored to its original precision, i.e., 1 / 16, through a scale-up operation.

[0406] Afterwards, the motion vector MV can be obtained by combining the motion vector difference MVD restored to the original precision and the motion vector prediction value MVP.

[0407] As in the example described above, when the motion vector prediction mode is applied, the decoder can derive the motion vector MV by combining the motion vector prediction value MVP and the motion vector difference value MVD.

[0408] Prediction using the aforementioned motion information is performed using a reconstructed picture decoded prior to the current picture, i.e., a reference picture. Alternatively, the current block can also be predicted by referencing a previously decoded region within the current picture. In other words, the current picture can be used as a reference picture, and a prediction block for the current block can be derived from a previously decoded region within the current picture. This type of prediction method can be referred to as the Intra Block Copy mode.

[0409] When the intra block copy mode is applied, the location of the reference block within the current picture may be indicated by a block vector. The size of the block vector may be limited to point to a previously decoded area prior to the current block within the current picture.

[0410] For example, if the block vector of the current block is (x0, y0), a block located at a position (x0, y0) away from the position (x, y) of the current block (i.e., (x-x0, y-y0)) can be set as the reference block of the current block.

[0411] Meanwhile, a block vector region (BV region) can be set, and a reference block of the current block can be determined from the BV region. The BV region contains previously decoded samples within the current picture and can be updated whenever a block is decoded.

[0412] Meanwhile, a picture may be divided into a plurality of reference blocks. Here, the reference block may represent a Coding Tree Unit (CTU) or a Coding Tree Block (CTB). Information indicating the size of the reference block may be encoded and signaled through an upper header.

[0413] The size of the BV area can be determined based on the size of the reference blocks.

[0414] Meanwhile, each of the reference blocks can be split based on a tree structure such as quad tree splitting, binary tree splitting, or triple tree splitting.

[0415] Additionally, encoding / decoding may be performed between blocks according to the raster scan order. However, the encoding / decoding order between blocks may also be determined according to a scanning method other than raster scanning, such as diagonal scanning, horizontal scanning, or vertical scanning.

[0416] The BV region may include at least one of the reconstructed samples of a decoded reference block prior to the current reference block and the reconstructed samples of a decoded block prior to the current block within the current reference block. Here, the current reference block may indicate a reference block including the current block.

[0417] Figure 26 is an example diagram to explain the configuration of the BV area.

[0418] For convenience of explanation, in the example shown in Fig. 26, it is assumed that the block at position d is the current block.

[0419] Assuming that the order of encoding / decoding between blocks is determined according to the raster scan order, the order of encoding / decoding of blocks (a to h) within the current reference block follows the order of a to h.

[0420] Accordingly, when starting to encrypt / decrypt the current block d, blocks a to c may have already been encrypt / decrypted.

[0421] The BV region may include at least one reference block that was sub-encoded / decoded before the current reference block and at least one block that was sub-encoded / decoded before the current block within the current reference block. That is, the BV region may include reconstruction samples of the reference block that was sub-encoded / decoded before the current reference block and reconstruction samples of blocks that were sub-encoded / decoded before the current reference block (i.e., blocks a to c).

[0422] Alternatively, for efficient buffer management, the BV area may be configured in units of reference blocks. That is, even if a block is encoded / decoded before the current block, if it is a block included in the same reference block as the current block (i.e., a to c in FIG. 26), it may not be included in the BV area.

[0423] Using the block vector candidate list, the block vector of the current block can be encoded / decoded.

[0424] For example, when the merge mode is applied, the block vector of the current block may be set to be identical to one of the block vector candidates selected from the block vector candidate list. In this case, information indicating one of the block vector candidates may be encoded / decoded.

[0425] On the other hand, when the block vector prediction mode is applied, one of the block vector candidates included in the block vector candidate list can be selected and set as the block vector prediction value of the current block. The block vector of the current block can be derived by adding a block vector differential value to the block vector prediction value. In this case, information indicating one of the block vector candidates and a block vector differential value indicating the difference between the block vector of the current block and the block vector prediction value can be explicitly encoded and signaled.

[0426]

[0427] When constructing a candidate list for inter prediction, a candidate generated by summing multiple motion vectors can also be inserted into the candidate list. Below, an example of adding a candidate generated by summing multiple motion vectors to the candidate list is described in detail. For convenience of explanation, the candidate list is assumed to be a motion information merge list.

[0428] FIG. 27 is a flowchart of a method for configuring a motion information merge list according to one embodiment of the present disclosure.

[0429] The example illustrated in Fig. 27 can be applied not only to encoding an image, but also to decoding an image.

[0430] If the number of motion information merge candidates included in the motion information merge list of the current block is less than a threshold (S2710), a chained merge candidate can be derived and the chained merge candidate can be inserted into the motion information merge list (S2720, S2730).

[0431] Here, the threshold value may be a value that is the maximum number of motion information merging candidates that the motion information merging list can include, or a value that is a difference from the maximum number.

[0432] A chain merge candidate can be added to a motion information merge list after inserting at least one of a motion information merge candidate derived from a spatial neighboring block of the current block, a motion information merge candidate derived from a temporal neighboring block, a motion information merge candidate derived from a history-based motion information buffer, or a pair-wise merge candidate into the motion information merge list.

[0433] For example, if the number of motion information merge candidates included in the motion information merge list is less than a threshold even though motion information merge candidates derived from spatial neighboring blocks and temporal neighboring blocks are added to the motion information merge list, the derived chain merge candidate can be added to the motion information merge list.

[0434] Figure 28 shows an example of deriving a chain merge candidate.

[0435] For convenience of explanation, it is assumed that there are three reference pictures (a first reference picture, a second reference picture, and a third reference picture) before the current picture, and that encoding / decoding is performed in the order of the third reference picture, the second reference picture, and the first reference picture.

[0436] To derive chain merge candidates, initial motion information is set. For example, the motion information of the motion information merge candidate with the smallest index among the motion information merge candidates included in the motion information merge list of the current block can be selected as the initial motion information.

[0437] Meanwhile, if a chain merge candidate cannot be derived from the selected motion information merge candidate, or if a chain merge candidate derived based on a motion information merge candidate with the smallest index is added to the motion information merge list, but the number of motion information merge candidates is less than a threshold, a new chain merge candidate can be derived by selecting the next motion information merge candidate.

[0438] That is, the motion information merge candidates can be derived in ascending order of the index to derive the chain merge candidates.

[0439] Alternatively, one of the merge candidates with unidirectional motion information in the motion information merge list may be selected, or one of the merge candidates with bidirectional motion information may be selected. Here, a merge candidate with unidirectional information refers to a merge candidate with motion information only in the L0 direction or the L1 direction, and a merge candidate with bidirectional information refers to a motion information merge candidate with motion information in the L0 direction and motion information in the L1 direction.

[0440] Alternatively, an index indicating one of the motion information merging candidates within the motion information merging list can be explicitly encoded / decoded. The initial motion information can be derived from the motion information merging candidate indicated by the index.

[0441] Alternatively, the motion information of a spatial neighboring block or a temporal neighboring block at a predefined location may be set as the initial motion information. Here, the spatial neighboring block at the predefined location may include at least one of a left neighboring block or an upper neighboring block.

[0442] Alternatively, the motion information of a candidate stored in a history-based motion information buffer can be set as the initial motion information. For example, the motion information of the candidate with the smallest index or the candidate with the largest index among the candidates stored in the history-based motion information buffer can be set as the initial motion information.

[0443] In the example illustrated in Fig. 28, the reference picture indicated by the initial motion information is the second reference picture, and the motion vector of the initial motion information is MV_1.

[0444] Based on the initial motion information of the current block, a first reference block is selected from the second reference picture (step 1). For example, if the reference picture index of the initial motion information points to the second reference picture and the motion vector of the initial motion information (i.e., the initial motion vector) is MV_1, a block located at a position that is spaced apart by MV_1 from a block at the same position as the current block within the second reference picture can be selected as the first reference block.

[0445] If there is motion information available for the first reference block, a second reference block is selected based on the first motion information of the first reference block (step 2). For example, if the first reference picture index of the first motion information points to the first reference picture and the motion vector of the first motion information (i.e., the first motion vector) is MV_2, a block located at a position MV_2 away from the position of the first reference block within the first reference picture can be selected as the second reference block.

[0446] If there is motion information available for the second reference block, a third reference block is selected based on the second motion information of the second reference block (step 3). For example, if the reference picture index of the second motion information points to the third reference picture and the motion vector of the second motion information (i.e., the second motion vector) is MV_3, a block located at a position MV_3 away from the position of the second reference block within the third reference picture can be selected as the third reference block.

[0447] If there is no motion information available for the third reference block, the reference picture index of the second motion information that was ultimately used to derive the third reference block (i.e., the index assigned to the third reference picture) can be set as the reference picture index of the concatenation candidate. In addition, the sum (i.e., MV_1 + MV_2 + MV_3) of the motion vectors that were concatenated to specify the third reference block (i.e., the initial motion vector, the first motion vector, and the second motion vector) can be set as the motion vector of the concatenation candidate.

[0448] That is, based on the initial motion information, a chain merge candidate can be derived by sequentially searching for motion information until no more motion information is available. Specifically, the reference picture index of the last searched motion information among the searched motion information can be set as the reference picture index of the chain merge candidate, and the motion vectors of the initial motion information and the searched motion information can be added together to be set as the motion vector of the chain merge candidate.

[0449] Alternatively, the sum of the initial motion vector and at least one motion vector searched so far may be compared with a threshold to determine whether to perform an additional search. For example, if the sum of the motion vectors is greater than the threshold, the additional search may not be performed. Meanwhile, the sum of the motion vectors being greater than the threshold may mean that either the x-component or the y-component of the sum is greater than the threshold, or that both the x-component and the y-component of the sum are greater than the threshold. Meanwhile, the thresholds for the x-component and the y-component may be the same, or the thresholds for the x-component and the y-component may be set individually.

[0450] Meanwhile, the size of the searched reference blocks can be set to be the same as the current block.

[0451] Whether or not motion information is available can be determined based on whether the position indicated by the motion information of the reference block is outside the boundary of a picture, subpicture, tile, or slice. For example, if the position indicated by the motion information of the reference block is outside the boundary of a picture, subpicture, tile, or slice, the motion information of the reference block can be determined to be unavailable.

[0452] Alternatively, it is also possible to determine whether available motion information exists based on whether motion information in the same direction as the initial motion information exists in the reference block.

[0453] For example, if the initial motion information has motion information in the L0 direction, the motion information of the reference block can be determined to be available only if the reference block has motion information in the L0 direction. For example, in the example of FIG. 28, it is assumed that the motion vector MV_1 of the initial motion information is in the L0 direction, and the motion vector MV_2 existing in the first reference block is in the L1 direction with respect to the reference block. If the search for deriving the concatenation candidate is set to be performed only in the same direction, the motion vector of the first reference block may not be available as a concatenation candidate.

[0454] Accordingly, the initial motion vector MV1 is selected as the final motion vector of the concatenation candidate, and the index pointing to the second reference picture can be set as the reference picture index of the concatenation candidate. In this example, if the reference block has bidirectional motion information, only the motion information in the L0 direction among the bidirectional motion information can be used to derive the concatenation candidate.

[0455] Meanwhile, if the motion information of the reference block searched in the first step is not available, the motion information of the concatenation candidate is set to be the same as the initial motion information. If the initial motion information is derived from a merge candidate in the concatenation motion information merge list, the motion information of the concatenation candidate is set to be the same as the merge candidate used to derive the initial motion information. To prevent identical merge candidates from being included in the motion information merge list, if the available motion information is not searched in the first step, it may be determined that the initial motion information cannot be used to derive an available concatenation candidate. In this case, the motion information of the next-ranked merge candidate may be set as the initial motion information, and the concatenation candidate derivation process may be restarted. Alternatively, for simplicity, if the initial motion information and the concatenation candidate have the same value, the method of deriving the concatenation candidate may not be used.

[0456] When the initial motion information is bidirectional, a chain search can be applied to each of the L0 and L1 directions. At this time, based on the L0 motion information of the initial motion information, motion information in the L0 direction can be searched, and based on the L1 motion information of the initial motion information, motion information in the L1 direction can be searched. At this time, as described above, among the motion information existing in the reference block, only motion information having the same directionality as that of the initial motion information can be used for chain merging.

[0457] If it is determined that the initial motion information is bidirectional and the direction of the motion information within the reference block must be the same as the direction of the initial motion information, the L0 reference picture index of the concatenation candidate is set to the reference picture index of the last searched L0 motion information, and the L0 motion vector of the concatenation candidate can be derived by adding the initial L0 motion vector and the searched L0 motion vectors. In addition, the L1 reference picture index of the concatenation candidate is set to the reference picture index of the last searched L1 motion information, and the L1 motion vector of the concatenation candidate can be derived by adding the initial L1 motion vector and the searched L1 motion vectors. Alternatively, whether or not available motion information exists can be determined based on whether or not the reference block indicated by the motion information has been encoded / decoded with inter prediction. That is, whether or not available motion information exists can be determined based on whether or not motion information is stored in the reference block indicated by the motion information.

[0458] For example, if a reference block specified by motion information is encoded / decoded with inter prediction, it may be determined that there is motion information available for the reference block, and if the reference block is not encoded / decoded with inter prediction, it may be determined that there is no motion information available for the reference block.

[0459] Meanwhile, if the reference block indicated by the motion information is encoded / decoded in intra block copy mode, the motion vector of the concatenated merge candidate can also be derived using the block vector of the reference block.

[0460] Figure 29 shows an example of deriving a chain merging candidate using a block vector.

[0461] For example, as in the example illustrated in FIG. 29, when the first reference block indicated by the initial motion information (i.e., MV_1, the second reference picture) is encoded / decoded in intra block copy mode, the second reference block within the second reference picture 1 can be specified based on the block vector (BV_1) of the reference block 1.

[0462] Thereafter, depending on whether motion information or block vectors are available in the second reference block, it can be determined whether to additionally utilize the motion information or block vectors of the second reference block when deriving a concatenated merge candidate. In the example illustrated in Fig. 29, it is exemplified that motion information (MV2, the third reference picture) within the second reference block is additionally searched.

[0463] If a reference block encoded / decoded in intra block copy mode is also set to be available for deriving a concatenation candidate, the motion vector of the concatenation candidate can be derived as the sum of the initial motion vector, the searched motion vector, and the searched block vector. For example, in the example illustrated in FIG. 29, the motion vector of the concatenation candidate can be derived by adding the motion vector MV_1 of the initial motion information, the block vector BV_1 of the first reference block, and the motion vector MV_2 of the second reference block.

[0464] Additionally, the reference picture index of the concatenated merge candidate can be derived from the motion information of the most recently searched reference block. For example, in the example illustrated in FIG. 29, the reference picture index of the concatenated merge candidate can be set to the index of the third reference picture based on the motion information of the second reference block.

[0465] Meanwhile, if the most recently searched reference block is encoded / decoded in intra block copy mode, the reference picture index of the concatenated merge candidate may be set to point to the reference picture to which the most recently searched reference block belongs.

[0466] The movement information of the reference block may be stored at a predefined location within the reference block.

[0467] Figure 30 shows the location from which the motion information of the reference block is derived.

[0468] The motion information of the reference block can be derived from a predefined position within the reference block. Here, the predefined position can be at least one of the upper left position (LT), the upper right position (RT), the lower left position (LB), the lower right position (RB), or the center position (C0), as in the example illustrated in FIG. 30.

[0469] Alternatively, the motion information of the reference block can be derived from the collocated blocks of the reference block.

[0470] Alternatively, the predefined positions within the reference block can be sequentially scanned, and the first available motion information found can be set as the motion information of the reference block.

[0471] Meanwhile, after correcting the initial motion vector, chain merge candidates can be derived based on the corrected initial motion vector. For example, a search area can be set centered around the location indicated by the initial motion information, and then the location within the search area with the lowest cost relative to the templates surrounding the current block can be determined. The initial motion vector of the current block can then be updated with the corrected motion vector to point to the location with the lowest cost.

[0472] Whether to correct the initial motion vector can be determined based on whether motion information is available at the location indicated by the initial motion vector. For example, if no motion information is available at the first location indicated by the initial motion information, the initial motion vector can be corrected, and a chain merge candidate can be derived based on the corrected initial motion vector.

[0473] On the other hand, if there is motion information available at the first position indicated by the initial motion information, no correction may be performed on the initial motion vector.

[0474] Additionally, the motion vector of the reference block may be corrected and used to derive a concatenation candidate. For example, if the first motion information of the first reference block points to a second reference location within a third reference picture, a search area may be set centered around the second reference location. Thereafter, the first motion vector may be corrected to the location with the lowest cost relative to the first reference block within the search area, and the corrected first motion vector may be used to derive a concatenation candidate.

[0475] Meanwhile, it is determined whether motion information is available for the second reference block pointed to by the corrected first motion vector, and if motion information is available for the second reference block, the second motion information of the second reference block may also be used to derive a concatenated merge candidate. In addition, a correction process may also be performed for the second motion vector of the second reference block.

[0476] Whether to correct the Nth motion information of the Nth reference block can also be determined based on whether there is motion information available for the (N+1)th reference block indicated by the Nth motion information. For example, if there is no motion information available for the second reference block indicated by the first motion information, the first motion information (specifically, the first motion vector) can be corrected, and a concatenation candidate can be derived based on the corrected first motion information. On the other hand, if there is motion information available for the second reference block indicated by the first motion information, the concatenation candidate can be derived without correcting the first motion information.

[0477] Alternatively, for simplicity, the correction of motion vectors may be performed only on at least one of the initial motion vectors or the most recently searched motion vectors.

[0478] Alternatively, the correction of the motion vector may be performed only for at least one of the initial motion vector or the first motion vector of the first reference block searched by the initial motion information.

[0479] Alternatively, the correction of the motion vector can be performed only for the initial motion vector and the N motion vectors searched thereafter. Here, N is a value predefined in the encoder and decoder, and can have a value such as 0, 1, 2, or 3.

[0480] In the above example, at least one motion information is searched for the purpose of searching for the aggregate motion information until the available motion information is exhausted. For example, in the example illustrated in FIG. 28, the first motion information is derived through a first search based on the initial motion information, and the second motion information is derived through a second search based on the first motion information.

[0481] Meanwhile, for simplicity, the search may be stopped when the number of searches reaches a threshold. For example, when the threshold is 1, a motion vector sum candidate may be derived based on the initial motion information and the first motion information derived through the first search based on the initial motion information. That is, the reference picture index of the motion vector sum candidate is set to be the same as the reference picture index of the first motion information, and the motion vector of the motion vector sum candidate may be derived by summing the initial motion vector and the first motion vector of the first motion information.

[0482] Alternatively, when the threshold is 2, a motion vector summation candidate can be derived based on initial motion information, first motion information derived through a first search based on the initial motion information, and second motion information derived through a second search based on the first motion information. That is, the reference picture index of the motion vector summation candidate is set to be the same as the reference picture index of the second motion information, and the motion vector of the motion vector summation candidate can be derived by summing the initial motion vector, the first motion vector of the first motion information, and the second motion vector of the second motion information.

[0483] The threshold may be predefined in the encoder and decoder.

[0484] Alternatively, information indicating the threshold value may be encoded and signaled via the bitstream. For example, information indicating the threshold value may be encoded / decoded via at least one of a picture header, a slice header, or a sequence parameter set.

[0485] Alternatively, the threshold may be adaptively determined based on at least one of the size / shape of the current block or the number of reference pictures in the reference picture list. For example, if the initial motion information has an L0 direction, the threshold may be determined based on the number of reference pictures that are in an active state in the L0 reference picture list. Alternatively, if the initial motion information has an L1 direction, the threshold may be determined based on the number of reference pictures that are in an active state in the L1 reference picture list.

[0486] Meanwhile, when multiple motion information pieces are searched, multiple chain merge candidates can be derived based on the multiple motion information pieces. For example, when N motion information pieces are searched based on the initial motion information, up to N chain merge candidates can be derived based on the N motion information pieces.

[0487] For example, a first chain merge candidate can be derived based on initial motion information and N pieces of motion information, and a second chain merge candidate can be derived based on (N-1) pieces of motion information excluding the initial motion information and the last searched motion information. The Nth chain merge candidate can be derived based on the initial motion information and the first searched motion information.

[0488] In this way, we can iteratively derive chain merge candidates by excluding the last searched motion information.

[0489] For example, in the example illustrated in FIG. 28, the first concatenated merge candidate can be derived based on initial motion information, motion information of the first reference block, and motion information of the second reference block. That is, the reference picture index of the first concatenated merge candidate is set to be the same as the reference picture index of the second reference block, and the motion vector of the first concatenated merge candidate can be derived as the sum of the motion vector of the initial motion information, the motion vector of the first reference block, and the motion vector of the second reference block.

[0490] In addition, the second concatenated merge candidate can be derived based on the initial motion information and the motion information of the first reference block, excluding the motion information of the second reference block that was last searched. That is, the reference picture index of the second concatenated merge candidate is set to be the same as the reference picture index of the first reference block, and the motion vector of the second concatenated merge candidate can be derived by adding the motion vector of the initial motion information and the motion vector of the first reference block.

[0491] Alternatively, after deriving multiple chain merge candidates, the cost of each chain merge candidate (specifically, the template matching cost) may be derived. Thereafter, M chain merge candidates with the smallest cost may be inserted into the motion information merge candidates. Here, M may be a natural number such as 1, 2, or 3. M may be defined in the encoder and decoder. Alternatively, M may be derived by the difference between the maximum number of motion information merge candidates that the motion information merge list can include and the number of motion information merge candidates included in the motion information merge list.

[0492] The cost of a chain merge candidate may be the Sum of Absolute Difference (SAD) between a template adjacent to the current block and a template adjacent to the reference block pointed to by the chain merge candidate (i.e., the most recently explored reference block).

[0493] Alternatively, the insertion order of the chain merge candidates can be determined by considering the cost of each chain merge candidate. That is, even if the chain merge candidate with the lowest cost is added to the motion information merge list, if the number of motion information merge candidates included in the motion information merge list does not reach a threshold, the next highest-priority chain merge candidate can be added to the motion information merge list.

[0494] Accordingly, a chain merge candidate with a small cost can be added to the motion information merge list first, and a chain merge candidate with a large cost can be added to the motion information merge list later.

[0495] As another example, if at least one piece of motion information is searched based on initial motion information, a concatenation candidate may be derived by excluding the initial motion information. For example, in the example illustrated in FIG. 28, the reference picture index of the concatenation candidate is set to be the same as the reference picture index of the second reference block, and the motion vector of the concatenation candidate may be derived by combining the motion vector of the first reference block and the motion vector of the second reference block.

[0496] Alternatively, instead of deriving a chain merge candidate, the motion information of the reference block indicated by the initial motion information may be used to derive a new merge candidate.

[0497] For example, in the example illustrated in FIG. 28, if reference block 1 is specified by initial motion information, the motion information of reference block 1 can be added to the motion information merge list as a new merge candidate.

[0498] Meanwhile, the candidate list construction method of FIG. 27 can also be applied to deriving motion vector prediction candidates within a motion vector prediction list. Specifically, when the number of motion vector prediction candidates included in the motion vector prediction list is less than a threshold, a chain motion vector prediction candidate can be derived, and the derived chain motion vector prediction candidate can be added to the motion vector prediction list.

[0499] The candidate for chain motion vector prediction may be identical to the one used to derive the motion vector of the chain merge candidate. If at least one motion information is retrieved based on the initial motion information, the motion vector of the initial motion information and the motion vector of the retrieved motion information may be combined to derive the candidate for chain motion vector prediction.

[0500] Meanwhile, when the motion vector prediction mode is applied, information indicating the reference picture index of the current block can be explicitly encoded / decoded. Accordingly, instead of setting the sum of motion vectors as the concatenated motion vector prediction candidate, the concatenated motion vector prediction candidate can be derived by scaling the summed motion vector.

[0501] Mathematical expression 7 shows an example of scaling the summed motion vector.

[0502]

[0503] In the above mathematical expression 7, a represents the distance between the current picture and the reference picture of the current block. Additionally, b represents the distance between the current picture and the reference picture indicated by the most recently searched motion information.

[0504] Alternatively, scaling can be performed based on a reference picture specified by the most recently searched motion information or a reference picture specified by the initial motion information.

[0505] The initial motion information can be derived from motion vector prediction candidates within the motion vector prediction list. That is, the motion vector prediction candidate with the smallest index within the motion vector prediction list is set as the motion vector of the initial motion information, and the reference picture index of the initial motion information can be set to be the same as the reference picture index of the current block.

[0506] Alternatively, the motion information of a spatial neighboring block or the motion information of a temporal neighboring block at a predefined location can be set as the initial motion information.

[0507] Meanwhile, the method for deriving chain merge candidates of Fig. 27 can also be applied when constructing a block vector candidate list under merge mode or prediction mode.

[0508] Specifically, when the number of block vector candidates included in the block vector candidate list is less than a threshold, a concatenated block vector candidate derived by combining the initial block vector and at least one searched block vector can be inserted into the block vector candidate list.

[0509] A concatenated block vector candidate can be added to the block vector candidate list after inserting at least one of a block vector candidate derived from a spatial neighboring block of the current block, a block vector candidate derived from a temporal neighboring block, a block vector candidate derived from a history-based block vector buffer, or a pair-wise block vector candidate into the block vector candidate list.

[0510] The initial block vector may be one of the block vector candidates included in the block vector candidate list. For example, the block vector candidate with the smallest index among the block vector candidates may be set as the initial block vector.

[0511] Alternatively, the initial block vector can be derived from the spatial or temporal neighboring blocks of the current block. For example, if a block vector exists in the left neighboring block or the upper neighboring block of the current block, the block vector of that block can be set as the initial block vector.

[0512] Figure 31 shows an example of deriving a chain block vector candidate.

[0513] When deriving a chain merge candidate, motion information of a reference block belonging to a reference picture different from the current picture is used.

[0514] In contrast, when deriving a chain block vector candidate, motion information of a reference block existing in the current picture is used.

[0515] Specifically, once the initial block vector (i.e., BV_1) is set, the first reference block in the current picture can be specified based on the initial block vector.

[0516] If there is a block vector available for the first reference block, a second reference block within the current picture can be determined based on the first block vector (i.e., BV_2) of the first reference block. Specifically, a block located at a position spaced apart from the first reference block by the first block vector (i.e., BV_2) can be set as the second reference block.

[0517] The above chain search process can be repeated until no block vectors are available for the searched reference block. For example, in the example illustrated in FIG. 31, if no block vectors are available for the third reference block indicated by the second block vector of the second reference block (i.e., BV_3), the search can be terminated and a chain block vector candidate can be derived.

[0518] Specifically, the concatenated block vector candidate can be derived by combining the initial block vector (BV_1), the first block vector (i.e., BV_2), and the second block vector (i.e., BV_3).

[0519] Meanwhile, the block vector of the reference block can be derived from a predefined location within the reference block. For example, a block vector candidate of the reference block can be derived from among the locations illustrated in FIG. 30.

[0520] Alternatively, as in the embodiment involving chain merge candidates, the search may be stopped when the number of searches reaches a threshold.

[0521] As in the embodiment involving the chain merge candidate, correction may also be performed on at least one of the initial motion vector or the motion vector of the reference block.

[0522] As in the embodiment related to the chain merge candidate, when multiple block vectors are searched, multiple chain block vector candidates can be derived based on the multiple block vectors. For example, when N block vectors are searched based on the initial block vector, up to N chain block vector candidates can be derived based on the N block vectors.

[0523] For example, the first block vector candidate can be set as the sum of the initial block vector, the first block vector, and the second block vector.

[0524] The second block vector candidate can be set as the sum of the initial block vector and the first block vector, excluding the second block vector that was last searched.

[0525] As in the embodiment related to the chain merge candidate, after deriving multiple chain block vector candidates, the cost (specifically, template matching cost) of each chain block vector candidate can be derived. Thereafter, the M chain block vector candidates with the highest cost can be inserted into the block vector candidate list.

[0526] Alternatively, the insertion order of the concatenated block vector candidates can be determined by considering the cost of each of the concatenated block vector candidates.

[0527] Meanwhile, the block vector of the current block can be restricted to point within the BV region. Accordingly, if the sum of the initial block vector and the searched block vectors points to a location outside the BV region, a chained block vector candidate can be derived, excluding the last searched block vector among the searched block vectors.

[0528] Meanwhile, even if a specific description is missing, the embodiments described focusing on the construction of a motion information merge list can also be used to construct a motion vector prediction list or a block vector candidate list.

[0529] In addition, it is also possible to use the embodiments described with a focus on configuring a motion vector prediction list to configure a motion information merge list or a block vector candidate list, or to use the embodiments described with a focus on configuring a block vector candidate list to configure a motion information merge list or a motion vector prediction list.

[0530] Additionally, a list for deriving motion information of the current block, such as a motion information merge list or a motion vector prediction list, may be called an inter prediction list, and a candidate (e.g., a motion information merge candidate or a motion vector prediction candidate) included in the inter prediction list may be called an inter prediction candidate.

[0531] Furthermore, a list for deriving a vector (i.e., a motion vector or a block vector) for prediction of the current block, such as an inter prediction list for deriving motion information of the current block or a block vector list for deriving a block vector of the current block, may be called a candidate list, and a candidate included in the candidate list may be called a prediction candidate.

[0532] Meanwhile, the above-described embodiments relate to a method for generating a new candidate for insertion into a candidate list based on motion information or block vectors of reference blocks.

[0533] However, instead of generating a new candidate based on the motion information or block vector of the reference blocks, it is also possible to obtain a prediction block of the current block by weighting the concatenated reference blocks.

[0534] For example, as in the example illustrated in FIG. 28, if a first reference block, a second reference block, and a third reference block are specified based on the initial motion information of the current block, the final predicted block of the current block can also be obtained based on a weighted sum operation of the serially searched first reference block, second reference block, and third reference block.

[0535] At this time, the weight assigned to each reference block can be determined according to the ratio of the distance between the current picture and the reference picture to which each reference block belongs (i.e., POC difference).

[0536] Alternatively, a reference block explored earlier may be assigned a higher weight than a reference block explored later.

[0537] Alternatively, it is also possible to perform a weighted sum by applying the same weights.

[0538]

[0539] Applying the embodiments described above, focusing on the decoding or encoding process, to the encoding or decoding process is within the scope of the present disclosure. Changing the embodiments described above, in a given order, to a different order is also within the scope of the present disclosure.

[0540] Although the above-described disclosure is described based on a series of steps or a flowchart, this does not limit the chronological order of the invention, and may be performed simultaneously or in a different order as needed. In addition, each component (e.g., unit, module, etc.) constituting the block diagram in the above-described disclosure may be implemented as a hardware device or software, or multiple components may be combined to be implemented as a single hardware device or software. For example, the hardware device may include at least one of a processor for performing calculations, a memory for storing data, a transmitter for transmitting data, and a receiver for receiving data.

[0541] The above-described disclosure may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination.

[0542] In addition, according to the present disclosure, a computer-readable recording medium can be provided that stores a bitstream generated by the above-described encoding method. The bitstream can be transmitted by an encoding device, and a decoding device can receive the bitstream and decode an image.

[0543] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memories. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present disclosure, and vice versa.

[0544] The present disclosure may be applied to a computing or electronic device capable of encoding / decoding a video signal.

Claims

1. Step of constructing an inter prediction list of the current block; A step of deriving a motion vector of the current block from the above inter prediction list; and A step of obtaining a prediction block for the current block based on the above motion vector, A video decoding method, characterized in that the inter prediction list includes a chain inter prediction candidate derived based on the initial motion information of the current block and at least one motion information searched in a chain based on the initial motion information.

2. In paragraph 1, A method for decoding an image, characterized in that the initial motion information is derived from an inter prediction candidate having the smallest index in the inter prediction list.

3. In paragraph 1, An image decoding method, characterized in that the initial motion information is derived from a spatial neighboring block or a temporal neighboring block at a predefined position of the current block.

4. In paragraph 1, If at least one piece of motion information is searched based on the above initial motion information, The reference picture index of the above chain inter prediction candidate is set to the reference picture index of the most recently searched motion information, An image decoding method, characterized in that the motion vector of the above-mentioned chain inter prediction candidate is derived by combining the motion vectors of each of the initial motion information and the searched motion information.

5. In paragraph 4, A method for decoding an image, characterized in that the search for motion information to derive the above-mentioned chain inter prediction candidate is repeatedly performed until no motion information is available for a reference block specified by the last searched motion information.

6. In paragraph 5, An image decoding method, characterized in that the motion information of the reference block is obtained from a predefined position within the reference block.

7. In paragraph 4, An image decoding method, characterized in that the search for motion information to derive the above-mentioned chain inter prediction candidate is performed repeatedly until the number of search times reaches a threshold.

8. In paragraph 4, If there is an available block vector for the reference block indicated by the above initial motion information or the searched motion information, An image decoding method, characterized in that the block vector is additionally used when deriving the above chain inter prediction candidate.

9. In paragraph 4, Based on the above initial movement information, if multiple movement information is searched, An image decoding method, characterized in that a plurality of chain inter prediction candidates are derived based on the plurality of pieces of motion information.

10. In paragraph 9, Among the above multiple chain inter prediction candidates, the first chain inter prediction candidate is derived using the above multiple motion information, An image decoding method, characterized in that a second inter prediction candidate among the plurality of chain inter prediction candidates is derived by excluding the last motion information among the plurality of motion information.

11. In paragraph 1, Based on the above initial movement information, if multiple movement information is searched, An image decoding method, characterized in that the motion vector of the above-mentioned chain inter prediction candidate is derived by scaling the sum of the motion vectors of the initial motion information and each of the plurality of motion information.

12. In paragraph 11, A method for decoding an image, characterized in that the scaling is performed based on the distance between the current picture and the reference picture of the current block and the distance between the current picture and the reference picture indicated by the last searched motion information.

13. In paragraph 1, A video decoding method, characterized in that the above inter prediction list is a motion information merge list or a motion vector prediction list.

14. Step of constructing an inter prediction list of the current block; A step of deriving a motion vector of the current block from the above inter prediction list; and A step of obtaining a prediction block for the current block based on the above motion vector, A video encoding method, characterized in that the inter prediction list includes a chain inter prediction candidate derived based on the initial motion information of the current block and at least one motion information searched in a chain based on the initial motion information.

15. Step of constructing an inter prediction list of the current block; A step of deriving a motion vector of the current block from the above inter prediction list; and A step of obtaining a prediction block for the current block based on the above motion vector, A computer-readable recording medium storing a bitstream encoded by a video encoding method, characterized in that the inter prediction list includes a chain inter prediction candidate derived based on initial motion information of the current block and at least one motion information searched in a chain based on the initial motion information.

Citation Information

Patent Citations

  • Derivation of motion vectors in video coding

    KR102520296B1

  • Graphite sheet and manufacturing method thereof

    KR102575637B1

  • Encoding and decoding method and device, encoding unit and decoding unit

    KR102619925B1

  • Ceramic susceptor and manufacturing method thereof

    KR102681418B1

  • Multi-frame motion compensation synthesis for video coding

    WO2023172243A1

Cited By

  • Motion vector predictor by using motion vector predictor lookahead / lookbehind

    US20250240449A1

  • Scaling and reordering of chained motion vector prediction for video coding

    US20250317570A1