Video encoding / decoding method and recording medium storing bitstream

WO2024205224A3PCT designated stage expired Publication Date: 2025-06-19KT CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/003843
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2024-03-27
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video data leads to increased data volume, resulting in higher transmission and storage costs, necessitating efficient video compression technologies, particularly for high-resolution and ultra-high-resolution stereoscopic video content.

Method used

A method and apparatus for deriving a prediction block by modifying a reference block through inter or intra-screen prediction, using weight parameters based on information from previously decoded regions, and applying filters to improve prediction accuracy without increasing signaling overhead.

Benefits of technology

This approach enhances prediction accuracy for video encoding and decoding, effectively managing high-resolution video data without increasing transmission or storage costs, by using weight parameters to modify reference blocks and generate prediction blocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024003843_19062025_PF_FP_ABST
    Figure KR2024003843_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A video decoding method according to the present disclosure comprises the steps of: deriving a block vector or motion information for a current block; deriving a reference block for the current block on the basis of the motion information or the block vector; deriving a weight parameter on the basis of a current template around the current block and a reference template around the reference block; and generating a prediction block for the current block on the basis of the weight parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and recording medium for storing bitstream

[0001] The present disclosure relates to a video signal processing method and device.

[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, is increasing across various application fields. As image data becomes higher in resolution and quality, the relative amount of data increases compared to conventional image data. Therefore, transmitting image data using existing media such as wired and wireless broadband lines or storing it using existing storage media leads to increased transmission and storage costs. To address these issues arising from the increasing resolution and quality of image data, high-efficiency image compression technologies can be utilized.

[0003] There are various technologies for image compression, such as inter-picture prediction technology that predicts pixel values ​​included in the current picture from pictures before or after the current picture, intra-picture prediction technology that predicts pixel values ​​included in the current picture using pixel information in the current picture, and entropy encoding technology that assigns short codes to values ​​with high frequency of appearance and long codes to values ​​with low frequency of appearance. Using these image compression technologies, image data can be effectively compressed and transmitted or stored.

[0004] Meanwhile, as demand for high-resolution video grows, so does the demand for stereoscopic video content as a new video service. Discussions are underway on video compression technologies to effectively deliver high-resolution and ultra-high-resolution stereoscopic video content.

[0005] The present disclosure aims to provide a method for deriving a prediction block by modifying a reference block derived through inter-prediction or intra-screen block copying, and a device therefor.

[0006] The present disclosure aims to provide a method and a device for deriving weight parameters for modifying a reference block based on information of a previously decoded area.

[0007] The present disclosure aims to provide a method for deriving a prediction block using a plurality of reference blocks and a device therefor.

[0008] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0009] A video decoding method according to the present disclosure may include: a step of deriving motion information or a block vector for a current block; a step of deriving a reference block for the current block based on the motion information or the block vector; a step of deriving a weight parameter based on a current template around the current block and a reference template around the reference block; and a step of generating a prediction block for the current block based on the weight parameter.

[0010] A video encoding method according to the present disclosure may include: a step of deriving motion information or a block vector for a current block; a step of deriving a reference block for the current block based on the motion information or the block vector; a step of deriving a weight parameter based on a current template around the current block and a reference template around the reference block; and a step of generating a prediction block for the current block based on the weight parameter.

[0011] In a video encoding / decoding method according to one embodiment of the present disclosure, the weight parameter may include filter coefficients, and the filter coefficients may be such that, when a filter is applied to a reference sample in the reference template, a difference from the reference sample in the current template is minimized.

[0012] In a video encoding / decoding method according to one embodiment of the present disclosure, the prediction sample of the current block can be derived by modifying a sample in the reference block based on the weight parameter.

[0013] In a video encoding / decoding method according to one embodiment of the present disclosure, the modified sample in the reference block can be obtained by filtering the sample using adjacent samples adjacent to the sample.

[0014] In a video encoding / decoding method according to one embodiment of the present disclosure, when an adjacent sample adjacent to the sample is unavailable, the value of the adjacent sample may be set to be equal to the value of the sample.

[0015] In a video encoding / decoding method according to one embodiment of the present disclosure, if an adjacent sample adjacent to the sample is located outside the reference block, the value may be set to be the same as the value of the restored sample at the corresponding location.

[0016] In a video encoding / decoding method according to one embodiment of the present disclosure, the filter applied to the sample may be a cross-shaped filter having the same width and height.

[0017] In a video encoding / decoding method according to one embodiment of the present disclosure, a filter type applied to the sample can be adaptively determined depending on the location of the sample.

[0018] According to the present disclosure, a computer-readable recording medium for storing a bitstream generated by the image encoding method can be provided.

[0019] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0020] According to the present disclosure, there is an effect of improving prediction accuracy by deriving a prediction block by modifying a reference block derived through inter-prediction or intra-screen block copying.

[0021] According to the present disclosure, by deriving weight parameters for modifying a reference block based on information of a previously decoded area, there is an effect of improving prediction accuracy without increasing signaling overhead.

[0022] According to the present disclosure, there is an effect of improving prediction accuracy by deriving a prediction block using a plurality of reference blocks.

[0023] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0024] FIG. 1 is a block diagram illustrating an image encoding device according to an embodiment of the present disclosure.

[0025] FIG. 2 is a block diagram showing an image decoding device according to an embodiment of the present disclosure.

[0026] FIG. 3 illustrates an image encoding / decoding method performed by an image encoding / decoding device according to the present disclosure.

[0027] FIG. 4 and FIG. 5 illustrate examples of multiple intra prediction modes according to the present disclosure.

[0028] FIG. 6 illustrates an intra prediction method based on a planar mode according to the present disclosure.

[0029] FIG. 7 illustrates an intra prediction method based on DC mode according to the present disclosure.

[0030] FIG. 8 illustrates an intra prediction method based on a directional mode according to the present disclosure.

[0031] Figure 9 illustrates a method for deriving samples of fractional positions.

[0032] Figures 10 and 11 illustrate that the tangent value for the angle is scaled by a factor of 32 for each intra prediction mode.

[0033] Figure 12 is a diagram illustrating an intra prediction aspect when the directional mode is one of modes 34 to 49.

[0034] Figure 13 is a drawing for explaining an example of generating an upper reference sample by interpolating left reference samples.

[0035] Figure 14 shows an example in which intra prediction is performed using reference samples arranged in a 1D array.

[0036] Figure 15 is a diagram schematically illustrating the process of performing inter prediction in an encoder and decoder.

[0037] Figure 16 shows an example in which motion estimation is performed.

[0038] Figures 17 and 18 illustrate examples in which a prediction block of a current block is generated based on motion information generated through motion estimation.

[0039] Figure 19 shows the locations referenced to derive motion vector prediction values.

[0040] Figure 20 is a diagram for explaining a template-based motion estimation method.

[0041] Figure 21 shows examples of template configurations.

[0042] Figure 22 is a diagram for explaining a motion estimation method based on a bilateral matching method.

[0043] Figure 23 is a diagram for explaining a motion estimation method based on a one-way matching method.

[0044] Figures 24 and 25 are diagrams showing examples in which prediction blocks are derived according to motion vector precision.

[0045] Figure 26 is a flowchart of a method for copying blocks within a screen performed by a video encoding / decoding device.

[0046] Figure 27 is for explaining an example of deriving a block vector using template matching.

[0047] Figure 28 shows an example of deriving a prediction block of the current block using a block vector.

[0048] Figure 29 illustrates a process for deriving a modified reference block.

[0049] Figure 30 illustrates the configuration of the current template and reference template.

[0050] Figure 31 illustrates a convolution filter applied to a reference sample.

[0051] Figure 32 shows an example in which padding is performed for unavailable samples.

[0052] Figure 33 shows an example in which values ​​are generated for samples adjacent to a reference block through padding.

[0053] Figure 34 illustrates an example in which padding is performed only for sample locations that do not belong to the reference template among samples located outside the reference block.

[0054] FIG. 35 is a diagram for explaining an example of obtaining a prediction block of a current block using multiple reference blocks.

[0055] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.

[0056] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.

[0057] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0058] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0059] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereinafter, identical components in the drawings will be designated by the same reference numerals, and redundant descriptions of identical components will be omitted.

[0060] FIG. 1 is a block diagram illustrating an image encoding device according to an embodiment of the present disclosure.

[0061] Referring to FIG. 1, a video encoding device (100) may include a picture segmentation unit (110), a prediction unit (120, 125), a transformation unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse transformation unit (145), a filter unit (150), and a memory (155).

[0062] Each component shown in Fig. 1 is independently depicted to represent different characteristic functions in the video encoding device, and does not mean that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or one component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present disclosure as long as they do not deviate from the essence of the present disclosure.

[0063] Additionally, some components may not be essential components that perform the essential functions of the present disclosure, but may be optional components merely used to enhance performance. The present disclosure may be implemented by including only components essential to implementing the essence of the present disclosure, excluding components used solely for performance enhancement. A structure that includes only essential components, excluding optional components used solely for performance enhancement, is also within the scope of the present disclosure.

[0064] The picture splitting unit (110) can split the input picture into at least one processing unit. At this time, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The picture splitting unit (110) can split one picture into a combination of multiple coding units, prediction units, and transform units, and select one combination of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function) to encode the picture.

[0065] For example, a picture can be split into multiple coding units. A recursive tree structure such as a quad tree, a ternary tree, or a binary tree can be used to split a coding unit in a picture. A coding unit that is split into other coding units starting from an image or the largest coding unit as the root can be split into as many child nodes as the number of split coding units. A coding unit that cannot be split any further according to a certain restriction becomes a leaf node. For example, assuming that a quad tree split is applied to a coding unit, a coding unit can be split into at most four different coding units.

[0066] Hereinafter, in the embodiments of the present disclosure, the encoding unit may be used to mean a unit that performs encoding or may be used to mean a unit that performs decoding.

[0067] A prediction unit may be divided into at least one square or rectangular shape of the same size within a single coding unit, or may be divided such that one prediction unit among the divided prediction units within a single coding unit has a different shape and / or size from another prediction unit.

[0068] When predicting within a screen, the transformation unit and the prediction unit can be set to be the same. In this case, the encoding unit can be divided into multiple transformation units, and then intra-screen prediction can be performed for each transformation unit. The encoding unit can be divided in the horizontal or vertical direction. The number of transformation units generated by dividing the encoding unit can be 2 or 4, depending on the size of the encoding unit.

[0069] The prediction unit (120, 125) may include an inter-prediction unit (120) that performs inter-prediction and an intra-prediction unit (125) that performs intra-prediction. It may be determined whether to use inter-prediction or intra-prediction for an encoding unit, and specific information (e.g., intra-prediction mode, motion vector, reference picture, etc.) according to each prediction method may be determined. At this time, the processing unit where prediction is performed and the processing unit where the prediction method and specific contents are determined may be different. For example, the prediction method and prediction mode, etc. are determined in the encoding unit, and the prediction may be performed in the prediction unit or the transformation unit. The residual value (residual block) between the generated prediction block and the original block may be input to the transformation unit (130). In addition, the prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value in the entropy encoding unit (165) and transmitted to the decoding device. When using a specific encoding mode, it is also possible to encode the original block as is and transmit it to the decoding unit without generating a prediction block through the prediction unit (120, 125).

[0070] The inter-screen prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous or subsequent pictures of the current picture, and in some cases, may predict a prediction unit based on information of a portion of an encoded region within the current picture. The inter-screen prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.

[0071] The reference picture interpolation unit can receive reference picture information from the memory (155) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.

[0072] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods can be used to derive a motion vector, such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm). The motion vector can have a motion vector value of 1 / 2 or 1 / 4 pixel unit based on the interpolated pixel. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various methods can be used as motion prediction methods, such as the Skip method, the Merge method, the AMVP (Advanced Motion Vector Prediction) method, and the Intra Block Copy method.

[0073] The on-screen prediction unit (125) can generate a prediction block based on reference pixel information, which is pixel information within the current picture. The reference pixel information can be derived from one selected from among a plurality of reference pixel lines. The Nth reference pixel line among the plurality of reference pixel lines can include left pixels having an x-axis difference of N from the upper left pixel within the current block and upper pixels having a y-axis difference of N from the upper left pixel. The number of reference pixel lines that the current block can select can be 1, 2, 3, or 4.

[0074] If the neighboring blocks of the current prediction unit are blocks that have performed inter-screen prediction and the reference pixel is a pixel that has performed inter-screen prediction, the reference pixel included in the block that has performed inter-screen prediction can be replaced with the reference pixel information of the neighboring block that has performed intra-screen prediction. That is, if the reference pixel is unavailable, the unavailable reference pixel information can be replaced with information from at least one of the available reference pixels.

[0075] In intra-screen prediction, the prediction mode can have a directional prediction mode that uses reference pixel information according to the prediction direction, and a non-directional mode that does not use directional information when performing prediction. The mode for predicting luminance information and the mode for predicting chrominance information can be different, and the intra-screen prediction mode information used to predict luminance information or the predicted luminance signal information can be utilized to predict chrominance information.

[0076] When performing intra-screen prediction, if the size of the prediction unit and the size of the transformation unit are the same, intra-screen prediction for the prediction unit can be performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side.

[0077] The on-screen prediction method can generate prediction blocks by applying a smoothing filter to reference pixels according to the prediction mode. Depending on the selected reference pixel line, whether or not the smoothing filter is applied can be determined.

[0078] In order to perform an intra-screen prediction method, the intra-screen prediction mode of the current prediction unit can be predicted from the intra-screen prediction modes of prediction units existing around the current prediction unit. When the prediction mode of the current prediction unit is predicted using mode information predicted from the surrounding prediction units, if the intra-screen prediction modes of the current prediction unit and the surrounding prediction units are the same, information indicating that the prediction modes of the current prediction unit and the surrounding prediction units are the same can be transmitted using predetermined flag information, and if the prediction modes of the current prediction unit and the surrounding prediction units are different, entropy encoding can be performed to encode the prediction mode information of the current block.

[0079] Additionally, a residual block containing residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130).

[0080] In the transformation unit (130), the residual block including the residual value information of the prediction unit generated through the original block and the prediction unit (120, 125) can be transformed using a transformation method such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on at least one of the size of the transformation unit, the shape of the transformation unit, the prediction mode of the prediction unit, or the prediction mode information within the screen of the prediction unit.

[0081] The quantization unit (135) can quantize values ​​converted to the frequency domain by the transformation unit (130). The quantization coefficients can vary depending on the block or the importance of the image. The values ​​produced by the quantization unit (135) can be provided to the dequantization unit (140) and the reordering unit (160).

[0082] The rearrangement unit (160) can perform rearrangement of coefficient values ​​for quantized residual values.

[0083] The rearrangement unit (160) can change a two-dimensional block-shaped coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the rearrangement unit (160) can change the two-dimensional block-shaped coefficient into a one-dimensional vector form by scanning from the DC coefficient to the coefficient of the high-frequency region using a zig-zag scan method. Depending on the size of the transformation unit and the intra-screen prediction mode, a vertical scan that scans the two-dimensional block-shaped coefficient in the column direction, a horizontal scan that scans the two-dimensional block-shaped coefficient in the row direction, or a diagonal scan that scans the two-dimensional block-shaped coefficient in the diagonal direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra-screen prediction mode, it is possible to determine which scan method among the zig-zag scan, the vertical scan, the horizontal scan, or the diagonal scan is to be used.

[0084] The entropy encoding unit (165) can perform entropy encoding based on the values ​​produced by the rearrangement unit (160). Entropy encoding can use various encoding methods such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).

[0085] The entropy encoding unit (165) can encode various information such as residual value coefficient information of the encoding unit, block type information, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information from the rearrangement unit (160) and the prediction unit (120, 125).

[0086] The entropy encoding unit (165) can entropy encode the coefficient values ​​of the encoding unit input from the rearrangement unit (160).

[0087] The inverse quantization unit (140) and the inverse transformation unit (145) inversely quantize the values ​​quantized in the quantization unit (135) and inversely transform the values ​​transformed in the transformation unit (130). The residual values ​​generated in the inverse quantization unit (140) and the inverse transformation unit (145) can be combined with the predicted prediction units predicted through the motion estimation unit, motion compensation unit, and intra-screen prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.

[0088] The filter unit (150) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter).

[0089] A deblocking filter can remove block distortion caused by boundaries between blocks in a reconstructed picture. To determine whether to perform deblocking, a deblocking filter can be applied to the current block based on the pixels contained in several columns or rows within the block. When applying a deblocking filter to a block, a strong filter or a weak filter can be applied depending on the required deblocking filtering strength. Furthermore, when applying a deblocking filter, horizontal and vertical filtering can be processed in parallel when performing vertical and horizontal filtering.

[0090] The offset correction unit can correct the offset from the original image on a pixel-by-pixel basis for an image that has undergone deblocking. To perform offset correction for a specific picture, the pixels contained in the image can be divided into a certain number of regions, the regions to be offset can be determined, and the offset can be applied to those regions. Alternatively, the offset can be applied by considering the edge information of each pixel.

[0091] Adaptive Loop Filtering (ALF) can be performed based on the comparison of the filtered restored image with the original image. After dividing the pixels included in the image into predetermined groups, a filter to be applied to each group can be determined, and filtering can be performed differentially for each group. Information regarding whether to apply ALF can be transmitted by luminance signal for each coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied can vary depending on each block. Furthermore, an ALF filter of the same form (fixed form) can be applied regardless of the characteristics of the target block.

[0092] The memory (155) can store a restoration block or picture produced through the filter unit (150), and the stored restoration block or picture can be provided to the prediction unit (120, 125) when performing inter-screen prediction.

[0093] FIG. 2 is a block diagram showing an image decoding device according to an embodiment of the present disclosure.

[0094] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (210), a rearrangement unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).

[0095] When a video bitstream is input to a video encoding device, the input bitstream can be decoded in the opposite procedure to that of the video encoding device.

[0096] The entropy decoding unit (210) can perform entropy decoding in a procedure opposite to that of the entropy encoding unit of the video encoding device. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied in response to the method performed in the video encoding device.

[0097] The entropy decoding unit (210) can decode information related to intra-screen prediction and inter-screen prediction performed in the encoding device.

[0098] The reordering unit (215) can perform reordering based on the method in which the bitstream entropy-decoded by the entropy decoding unit (210) is reordered by the encoding unit. The coefficients expressed in the form of a one-dimensional vector can be reordered by restoring them back to coefficients in the form of a two-dimensional block. The reordering unit (215) can perform reordering by receiving information related to the coefficient scanning performed by the encoding unit and performing reverse scanning based on the scanning order performed by the corresponding encoding unit.

[0099] The dequantization unit (220) can perform dequantization based on the quantization parameters provided from the encoding device and the coefficient values ​​of the rearranged block.

[0100] The inverse transform unit (225) can perform inverse transform, i.e., inverse DCT, inverse DST, and inverse KLT, on the transforms, i.e., DCT, DST, and KLT, performed by the transform unit on the quantization result performed by the image encoding device. The inverse transform can be performed based on the transmission unit determined by the image encoding device. In the inverse transform unit (225) of the image decoding device, a transform technique (e.g., DCT, DST, KLT) can be selectively performed according to a plurality of pieces of information, such as a prediction method, the size and shape of the current block, the prediction mode, and the prediction direction within the screen.

[0101] The prediction unit (230, 235) can generate a prediction block based on prediction block generation related information provided from the entropy decoding unit (210) and previously decoded block or picture information provided from the memory (245).

[0102] As described above, when performing intra-screen prediction in the same manner as the operation in the video encoding device, if the size of the prediction unit and the size of the transformation unit are the same, intra-screen prediction for the prediction unit is performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side. However, when performing intra-screen prediction, if the size of the prediction unit and the size of the transformation unit are different, intra-screen prediction can be performed using reference pixels based on the transformation unit. In addition, intra-screen prediction using NxN division only for the minimum coding unit can be used.

[0103] The prediction unit (230, 235) may include a prediction unit determination unit, an inter-screen prediction unit, and an intra-screen prediction unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (210), prediction mode information of an intra-screen prediction method, and motion prediction-related information of an inter-screen prediction method, and may distinguish a prediction unit from a current encoding unit and determine whether the prediction unit performs inter-screen prediction or intra-screen prediction. The inter-screen prediction unit (230) may perform inter-screen prediction on the current prediction unit based on information included in at least one of a previous picture or a subsequent picture of the current picture including the current prediction unit, using information necessary for inter-screen prediction of the current prediction unit provided from the video encoding device. Alternatively, inter-screen prediction may be performed based on information on a pre-restored portion of the current picture including the current prediction unit.

[0104] In order to perform inter-screen prediction, it is possible to determine whether the motion prediction method of the prediction unit included in the encoding unit is Skip Mode, Merge Mode, AMVP Mode, or Intra-screen Block Copy Mode based on the encoding unit.

[0105] The intra-screen prediction unit (235) can generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit that has performed intra-screen prediction, intra-screen prediction can be performed based on intra-screen prediction mode information of the prediction unit provided by the video encoding device. The intra-screen prediction unit (235) can include an AIS (Adaptive Intra Smoothing) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a part that performs filtering on the reference pixels of the current block, and can determine and apply whether to apply the filter according to the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the video encoding device. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.

[0106] The reference pixel interpolation unit can generate a reference pixel of a pixel unit less than an integer value by interpolating the reference pixel when the prediction mode of the prediction unit is a prediction unit that performs intra-screen prediction based on the pixel value interpolated from the reference pixel. If the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.

[0107] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.

[0108] Information regarding whether a deblocking filter has been applied to a corresponding block or picture may be received from a video encoding device, and if a deblocking filter has been applied, information regarding whether a strong or weak filter has been applied. The deblocking filter of the video decoding device may receive information related to the deblocking filter provided by the video encoding device, and the video decoding device may perform deblocking filtering on the corresponding block.

[0109] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and offset value information.

[0110] ALF can be applied to an encoding unit based on information such as whether ALF is applied and ALF coefficient information provided from an encoding device. This ALF information can be provided by being included in a specific parameter set.

[0111] The memory (245) can store a restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to an output unit.

[0112] As described above, in the following embodiments of the present disclosure, for convenience of explanation, the term coding unit is used as an encoding unit, but it may also be a unit that performs not only encoding but also decoding.

[0113] In addition, the current block represents a block to be encoded / decoded, and may represent a coding tree block (or coding tree unit), an encoding block (or encoding unit), a transform block (or transform unit), a prediction block (or prediction unit), or a block to which an in-loop filter is applied, depending on the encoding / decoding step. In this specification, a 'unit' represents a basic unit for performing a specific encoding / decoding process, and a 'block' may represent a pixel array of a predetermined size. Unless otherwise distinguished, 'block' and 'unit' may be used with the same meaning. For example, in the embodiment described below, an encoding block (coding block) and an encoding unit (coding unit) may be understood to have the same meaning.

[0114] Furthermore, we will refer to the picture that contains the current block as the current picture.

[0115] Prediction for the current block can be performed based on intra prediction or inter prediction.

[0116] Intra prediction is to remove redundant data for the current block based on the similarity between the current block and reference samples.

[0117] FIG. 3 illustrates an intra prediction method performed by an image encoding / decoding device according to the present disclosure.

[0118] Referring to FIG. 3, a reference line for intra prediction of the current block can be determined (S300).

[0119] The current block can use one or more of a plurality of pre-defined reference line candidates in the video encoding / decoding device as reference lines for intra prediction. Here, the plurality of pre-defined reference line candidates can include neighboring reference lines adjacent to the current block to be decoded and N non-neighboring reference lines that are 1 to N samples away from the boundary of the current block. N can be 1, 2, 3, or an integer greater than or equal to 1. For convenience of explanation, it is assumed hereafter that the plurality of reference line candidates available to the current block are composed of neighboring reference line candidates and three non-neighboring reference line candidates, but the present invention is not limited thereto. That is, it goes without saying that the plurality of reference line candidates available to the current block can include four or more non-neighboring reference line candidates.

[0120] An image encoding device can determine an optimal reference line candidate from among a plurality of reference line candidates and encode an index for specifying the optimal reference line candidate. An image decoding device can determine a reference line of a current block based on an index signaled through a bitstream. The index can specify any one of the plurality of reference line candidates. The reference line candidate specified by the index can be used as a reference line of the current block.

[0121] The number of indexes signaled to determine the reference line of the current block may be 1, 2, or more. For example, when the number of indexes signaled is 1, the current block can perform intra prediction using only a single reference line candidate specified by the signaled index among a plurality of reference line candidates. Alternatively, when the number of indexes signaled is 2 or more, the current block can perform intra prediction using a plurality of reference line candidates specified by a plurality of indexes among a plurality of reference line candidates.

[0122] Referring to FIG. 3, the intra prediction mode of the current block can be determined (S310).

[0123] The intra prediction mode of the current block can be determined from among multiple intra prediction modes pre-defined in the video encoding / decoding device. The multiple pre-defined intra prediction modes will be described with reference to FIGS. 4 and 5.

[0124] FIG. 4 illustrates an example of multiple intra prediction modes according to the present disclosure.

[0125] Referring to FIG. 4, the multiple intra prediction modes pre-defined in the video encoding / decoding device may be configured as a non-directional mode and a directional mode. The non-directional mode may include at least one of a planar mode or a DC mode. The directional mode may include directional modes 2 to 66.

[0126] The directional mode can be further expanded than that shown in Fig. 4. Fig. 5 shows an example in which the directional mode is expanded.

[0127] In Fig. 5, modes -1 to -14 and modes 67 to 80 are exemplified as being added. These directional modes may be referred to as wide-angle intra prediction modes. Whether to use the wide-angle intra prediction mode may be determined depending on the shape of the current block. For example, if the current block is a non-square block whose width is greater than its height, some directional modes (e.g., 2 to 15) may be converted to wide-angle intra prediction modes between 67 and 80. On the other hand, if the current block is a non-square block whose height is greater than its width, some directional modes (e.g., 53 to 66) may be converted to wide-angle intra prediction modes between -1 and -14.

[0128] The range of available wide-angle intra prediction modes can be adaptively determined based on the width-to-height ratio of the current block. Table 1 shows the range of available wide-angle intra prediction modes based on the width-to-height ratio of the current block.

[0129] Available Wide Angle Intra Prediction Mode Ranges W / H = 1667~80 W / H = 867~78 W / H = 467~76 W / H = 267~74 W / H = 1 None W / H = 1 / 2-1~-8 W / H = 1 / 4-1~-10 W / H = 1 / 8-1~-12 W / H = 1 / 16-1~-14

[0130] Among the above multiple intra prediction modes, K candidate modes (most probable modes, MPMs) can be selected. A candidate list including the selected candidate modes can be generated. An index indicating one of the candidate modes in the candidate list can be signaled. The intra prediction mode of the current block can be determined based on the candidate mode indicated by the index. For example, the candidate mode indicated by the index can be set as the intra prediction mode of the current block. Alternatively, the intra prediction mode of the current block can be determined based on a value of the candidate mode indicated by the index and a predetermined difference value. The difference value can be defined as a difference between a value of the intra prediction mode of the current block and a value of the candidate mode indicated by the index. The difference value can be signaled through a bitstream. Alternatively, the difference value may be a pre-defined value in the video encoding / decoding device. Alternatively, the intra prediction mode of the current block may be determined based on a flag indicating whether a mode identical to the intra prediction mode of the current block exists in the candidate list. For example, when the flag has a first value, the intra prediction mode of the current block may be determined from the candidate list. In this case, an index indicating any one of a plurality of candidate modes belonging to the candidate list may be signaled. The candidate mode indicated by the index may be set as the intra prediction mode of the current block. On the other hand, when the flag has a second value, any one of the remaining intra prediction modes may be set as the intra prediction mode of the current block. The remaining intra prediction mode may mean a mode excluding a candidate mode belonging to the candidate list among the plurality of pre-defined intra prediction modes. When the flag has a second value, an index indicating any one of the remaining intra prediction modes may be signaled.An intra prediction mode indicated by a signaled index may be set as the intra prediction mode of the current block. The intra prediction mode of a chroma block may be selected from intra prediction mode candidates of multiple chroma blocks. To this end, index information indicating one of the intra prediction mode candidates of the chroma block may be explicitly encoded and signaled through the bitstream. Table 2 illustrates examples of intra prediction mode candidates of chroma blocks.

[0131] Intra prediction mode candidates for indexed chroma blocks Luma mode: 0 Luma mode: 50 Luma mode: 18 Luma mode: 1 Other 0 6 6 0 0 0 1 5 0 6 6 5 0 5 0 5 0 2 1 8 1 8 6 6 1 8 1 8 3 1 1 6 6 1 4 DM

[0132] In the example of Table 2, DM (Direct Mode) means setting the intra prediction mode of the luma block existing at the same position as the chroma block to the intra prediction mode of the chroma block. Meanwhile, the luma block existing at the same position as the chroma block can be determined based on the position of the upper left sample or the position of the center sample of the chroma block. For example, if the intra prediction mode (luma mode) of the luma block is 0 (planar mode) and the index points to 2, the intra prediction mode of the chroma block can be determined as the horizontal mode (18). For example, if the intra prediction mode (luma mode) of the luma block is 1 (DC mode) and the index points to 0, the intra prediction mode of the chroma block can be determined as the planar mode (0).

[0133] Consequently, the intra prediction mode of the chroma block may also be set to one of the intra prediction modes illustrated in FIG. 4 or FIG. 5. The intra prediction mode of the current block may also be used to determine a reference line of the current block, in which case step S310 may be performed before step S300.

[0134] Referring to FIG. 3, intra prediction can be performed for the current block based on the reference line and intra prediction mode of the current block (S320).

[0135] Hereinafter, with reference to FIGS. 6 through 8, we will examine in detail the intra prediction method for each intra prediction mode. However, for convenience of explanation, it is assumed that a single reference line is used for intra prediction of the current block. However, even when multiple reference lines are used, the intra prediction method described below can be applied in the same / similar manner.

[0136] FIG. 6 illustrates an intra prediction method based on a planar mode according to the present disclosure.

[0137] Referring to Fig. 6, T represents a reference sample located at the upper right corner of the current block, and L represents a reference sample located at the lower left corner of the current block. P1 can be generated through horizontal interpolation. For example, P1 can be generated by interpolating T with a reference sample located on the same horizontal line as P1. P2 can be generated through vertical interpolation. For example, P2 can be generated by interpolating L with a reference sample located on the same vertical line as P2. The current sample within the current block can be predicted through a weighted sum of P1 and P2, as in the following mathematical expression 1.

[0138]

[0139] In Equation 1, weights α and β can be determined by considering the width and height of the current block. Depending on the width and height of the current block, weights α and β may have the same value or different values. If the width and height of the current block are the same, weights α and β can be set to the same value, and the prediction sample of the current sample can be set to the average value of P1 and P2. If the width and height of the current block are not the same, weights α and β can have different values. For example, if the width is greater than the height, a smaller value can be set for the weight corresponding to the width of the current block, and a larger value can be set for the weight corresponding to the height of the current block. Conversely, if the width is greater than the height, a larger value can be set for the weight corresponding to the width of the current block, and a smaller value can be set for the weight corresponding to the height of the current block. Here, the weight corresponding to the width of the current block can mean β, and the weight corresponding to the height of the current block can mean α.

[0140] FIG. 7 illustrates an intra prediction method based on DC mode according to the present disclosure.

[0141] Referring to FIG. 7, the average value of neighboring samples adjacent to the current block can be calculated, and the calculated average value can be set as the predicted value of all samples in the current block. Here, the neighboring samples can include the upper reference sample and the left reference sample of the current block. However, depending on the shape of the current block, the average value can be calculated using only the upper reference sample or the left reference sample. For example, if the width of the current block is greater than the height, the average value can be calculated using only the upper reference sample of the current block. Alternatively, if the ratio of the width to the height of the current block is greater than or equal to a predetermined threshold, the average value can be calculated using only the upper reference sample of the current block. Alternatively, if the ratio of the width to the height of the current block is less than or equal to a predetermined threshold, the average value can be calculated using only the upper reference sample of the current block. On the other hand, if the width of the current block is less than the height, the average value can be calculated using only the left reference sample of the current block. Alternatively, if the ratio of the width to the height of the current block is less than or equal to a predetermined threshold, the average value can be calculated using only the left reference sample of the current block. Alternatively, if the ratio of the width and height of the current block is greater than or equal to a predetermined threshold, the average value can be calculated using only the left reference sample of the current block.

[0142] FIG. 8 illustrates an intra prediction method based on a directional mode according to the present disclosure.

[0143] If the intra prediction mode of the current block is a directional mode, projection can be performed to a reference line according to the angle of the directional mode. If a reference sample exists at the projected position, the reference sample can be set as a prediction sample of the current sample. If a reference sample does not exist at the projected position, a sample corresponding to the projected position can be generated using one or more neighboring samples neighboring the projected position. For example, a sample corresponding to the projected position can be generated by performing interpolation based on two or more neighboring samples neighboring in both directions with respect to the projected position. Alternatively, one neighboring sample neighboring the projected position can be set as the sample corresponding to the projected position. In this case, among the plurality of neighboring samples neighboring the projected position, the neighboring sample closest to the projected position can be used. The sample corresponding to the projected position can be set as a prediction sample of the current sample.

[0144] Referring to FIG. 8, for the current sample B, when projection is performed with a reference line according to the angle of the intra prediction mode at the corresponding position, a reference sample exists at the projected position (i.e., a reference sample at an integer position, R3). In this case, the reference sample at the projected position can be set as a prediction sample of the current sample B. For the current sample A, when projection is performed with a reference line according to the angle of the intra prediction mode at the corresponding position, a reference sample (i.e., a reference sample at an integer position) does not exist at the projected position. In this case, a sample (r) at a fractional position can be generated by performing interpolation based on neighboring samples (e.g., R2 and R3) adjacent to the projected position. The generated sample (r) at the fractional position can be set as a prediction sample of the current sample A.

[0145] Figure 9 illustrates a method for deriving samples of fractional positions.

[0146] In the example of Fig. 9, the variable h represents the vertical distance (i.e., vertical distance) between the position of the predicted sample A and the reference sample line, and the variable w represents the horizontal distance (i.e., horizontal distance) between the position of the predicted sample A and the fractional position sample. In addition, the variable θ represents a predefined angle according to the directionality of the intra prediction mode, and the variable x represents the fractional position.

[0147] The variable w can be derived as shown in the following mathematical expression 2.

[0148]

[0149] Afterwards, by removing the integer position from the variable w, the fractional position can be finally derived.

[0150] Fractional position samples can be generated by interpolating adjacent integer position reference samples. For example, integer position reference sample R2 and integer position reference sample R3 can be interpolated to generate fractional position reference samples at the x position.

[0151] To avoid floating-point operations when deriving fractional position samples, a scaling factor can be used. For example, if the scaling factor f is set to 32, the distance between neighboring integer reference samples can be set to 32 instead of 1, as in the example illustrated in (b) of Fig. 8.

[0152] Additionally, the tangent value for the angle θ determined by the directionality of the intra prediction mode can also be scaled up using the same scaling factor (e.g., 32).

[0153] Figures 10 and 11 illustrate that the tangent value for the angle is scaled by a factor of 32 for each intra prediction mode.

[0154] Figure 10 shows the scaled results of tangent values ​​for the non-wide angle intra prediction mode, and Figure 11 shows the scaled results of tangent values ​​for the wide angle intra prediction mode.

[0155] If the tangent value (tanθ) for the angle value of the intra prediction mode is positive, intra prediction can be performed using only one of the reference samples belonging to the upper line of the current block (i.e., upper reference samples) or the reference samples belonging to the left line of the current block (i.e., left reference samples). On the other hand, if the tangent value for the angle value of the intra prediction mode is negative, both the reference samples located at the upper side and the reference samples located at the left side are used.

[0156] At this time, to simplify the implementation, the reference samples may be arranged in a 1D array form by projecting the left reference samples upward or the top reference samples to the left, and intra prediction may be performed using the reference samples in the 1D array form.

[0157] Figure 12 is a diagram illustrating an intra prediction aspect when the directional mode is one of modes 34 to 49.

[0158] If the intra prediction mode of the current block is one of modes 34 to 49, intra prediction is performed using not only the upper reference samples of the current block but also the left reference samples. At this time, as in the example illustrated in Fig. 12, the reference sample located on the left side of the current block can be copied to the position of the upper line, or the reference samples located on the left can be interpolated to generate the reference sample of the upper line.

[0159] For example, in case of obtaining a reference sample for position A at the top of the current block, projection can be performed from position A of the top line to the left line of the current block, considering the directionality of the intra prediction mode of the current block. If the projected position is a, the value corresponding to position a can be copied, or a fractional position value corresponding to a can be generated and set as the value of position A. For example, if position a is an integer position, the value of position A can be generated by copying the integer position reference sample. On the other hand, if position a is a fractional position, the reference sample located above position a and the reference sample located below position a can be interpolated, and the interpolated value can be set as the value of position A. Meanwhile, the direction of projection from position A at the top of the current block to the left line of the current block can be parallel to and opposite to the direction of the intra prediction mode of the current block.

[0160] Figure 13 is a drawing for explaining an example of generating an upper reference sample by interpolating left reference samples.

[0161] In Fig. 13, the variable h represents the horizontal distance between position A of the upper line and position a of the left line. The variable w represents the vertical distance between position A of the upper line and position a of the left line. In addition, the variable θ represents a predefined angle according to the directionality of the intra prediction mode, and the variable x represents a fractional position.

[0162] The variable h can be derived as shown in the following mathematical expression 3.

[0163]

[0164] Afterwards, by removing the integer position from the variable h, the fractional position can be finally derived.

[0165] To avoid real-valued operations when deriving fractional position samples, a scaling factor can be used. For example, the tangent value for variable θ can be scaled using the scaling factor f1. Here, since the direction projected to the left line is parallel and opposite to the directional prediction model, the scaled tangent value illustrated in FIGS. 10 and 11 can also be used.

[0166] When the scaling factor f1 is applied, Equation 3 can be transformed into Equation 4 below.

[0167]

[0168] In this manner, a 1D reference sample array can be constructed using only the reference samples belonging to the upper line. As a result, intra prediction for the current block can be performed using only the upper reference samples formed in the 1D array.

[0169] Figure 14 shows an example in which intra prediction is performed using reference samples arranged in a 1D array.

[0170] As in the example illustrated in Fig. 14, by projecting the left reference samples to generate the upper reference samples, prediction samples of the current block can be obtained using only the reference samples belonging to the upper line.

[0171] Contrary to what is shown in FIGS. 12 and 14, a 1D reference sample array can also be constructed using only the reference samples belonging to the left line by projecting the upper reference sample onto the left line. Specifically, for modes 19 to 33 among the directional modes in which the tangent value (tanθ) for the angle of the directional mode is negative, the reference samples belonging to the upper line can be projected onto the left line to generate the left reference sample.

[0172]

[0173] Inter prediction is to generate a prediction block for the current block based on the similarity between the current block and a reference block within a reference picture. That is, when encoding the current picture, redundant data between pictures can be removed through inter prediction. Inter prediction can be performed on a block-by-block basis. Specifically, a prediction block of the current block can be generated from a reference picture using motion information of the current block. Here, the motion information can include at least one of a motion vector, a reference picture index, and a prediction direction.

[0174] Figure 15 is a diagram schematically illustrating the process of performing inter prediction in an encoder and decoder.

[0175] As in the example illustrated in FIG. 15, in order to perform inter prediction, motion information for the current block can be acquired (S1510). Here, the motion information can include at least one of a motion vector, a reference picture index, or a weight applied to the prediction block. For the current block, motion information for at least one of the L0 direction or the L1 direction can be acquired.

[0176] In the encoder, motion information of the current block can be derived through motion estimation, and the derived motion information can be encoded and signaled to the decoder. Meanwhile, the encoding / decoding of motion information can be based on a motion information merging mode, a motion vector prediction mode, a template-based motion estimation method, or a bilateral matching method, which will be described later.

[0177] In the decoder, motion information of the current block can be derived based on the information transmitted from the encoder.

[0178] Alternatively, the motion information of the current block can be derived from the decoder in the same manner as in the encoder. This method can be referred to as decoder-side motion estimation.

[0179] Once motion information for the current block is derived, a prediction block for the current block can be obtained based on the derived motion information (S1520). For example, a reference block spaced apart by a motion vector from the current block's position within the reference picture can be set as the prediction block for the current block.

[0180] Below, we will explain in more detail the process of calculating inter predictions.

[0181] The motion information of the current block can be generated through motion estimation.

[0182] Figure 16 shows an example in which motion estimation is performed.

[0183] In Fig. 16, it is assumed that the POC (Picture Order Count) of the current picture is T, and the POC of the reference picture is (T-1).

[0184] A search range for motion estimation can be set from the same location as the reference point of the current block within the reference picture. Here, the reference point may be the location of the upper left sample of the current block.

[0185] For example, in Fig. 16, it is illustrated that a rectangle of size (w0+w01) and (h0+h1) is set as a search range centered on a reference point. In the above example, w0, w1, h0, and h1 may have the same value. Alternatively, at least one of w0, w1, h0, and h1 may be set to have a different value from the other. Alternatively, the sizes of w0, w1, h0, and h1 may be determined so as not to exceed a Coding Tree Unit (CTU) boundary, a slice boundary, a tile boundary, or a picture boundary.

[0186] Within the search range, reference blocks of the same size as the current block can be set, and the cost of each reference block relative to the current block can be measured. The cost can be calculated using the similarity between the two blocks.

[0187] For example, the cost can be calculated based on the absolute sum of the differences between the original samples in the current block and the original samples (or reconstructed samples) in the reference block. A smaller absolute sum can reduce the cost.

[0188] Afterwards, the cost of each reference block is compared, and the reference block with the optimal cost can be set as the prediction block of the current block.

[0189] Additionally, the distance between the current block and the reference block can be set as a motion vector. Specifically, the x-coordinate difference and the y-coordinate difference between the current block and the reference block can be set as the motion vector.

[0190] Furthermore, the index of the picture containing the reference block identified through motion estimation is set as the reference picture index.

[0191] Additionally, the prediction direction can be set based on whether the reference picture belongs to the L0 reference picture list or the L1 reference picture list.

[0192] Additionally, motion estimation can be performed for each of the L0 direction and the L1 direction. If prediction is performed for both the L0 direction and the L1 direction, motion information in the L0 direction and motion information in the L1 direction can be generated, respectively.

[0193] Figures 17 and 18 illustrate examples in which a prediction block of a current block is generated based on motion information generated through motion estimation.

[0194] Figure 17 shows an example of generating a prediction block with unidirectional (i.e., L0 direction) prediction, and Figure 18 shows an example of generating a prediction block with bidirectional (i.e., L0 and L1 direction) prediction.

[0195] In the case of unidirectional prediction, a prediction block of the current block is generated using a single motion information. For example, the motion information may include an L0 motion vector, an L0 reference picture index, and prediction direction information indicating the L0 direction.

[0196] In the case of bidirectional prediction, a prediction block is generated using two pieces of motion information. For example, a reference block in the L0 direction, determined based on motion information for the L0 direction (L0 motion information), can be set as an L0 prediction block, and an L1 prediction block can be generated based on a reference block in the L1 direction, determined based on motion information for the L1 direction (L1 motion information). Thereafter, the L0 prediction block and the L1 prediction block can be weighted and combined to generate a prediction block of the current block.

[0197] In the examples illustrated in FIGS. 16 to 18, the L0 reference picture is illustrated as existing in the previous direction of the current picture (i.e., having a POC value smaller than that of the current picture), and the L1 reference picture is illustrated as existing in the subsequent direction of the current picture (i.e., having a POC value larger than that of the current picture).

[0198] However, unlike the illustrated example, the L0 reference picture may exist in the subsequent direction of the current picture, or the L1 reference picture may exist in the previous direction of the current picture. For example, both the L0 reference picture and the L1 reference picture may exist in the previous direction of the current picture, or both may exist in the subsequent direction of the current picture. Alternatively, bidirectional prediction may be performed using the L0 reference picture existing in the subsequent direction of the current picture and the L1 reference picture existing in the previous direction of the current picture.

[0199] Motion information for blocks for which inter prediction has been performed can be stored in memory. At this time, the motion information can be stored on a sample-by-sample basis. Specifically, the motion information for a block to which a specific sample belongs can be stored as motion information for that specific sample. The stored motion information can be used to derive motion information for neighboring blocks to be encoded / decoded in the future.

[0200] In the encoder, information encoding residual samples corresponding to the difference between the sample of the current block (i.e., the original sample) and the predicted sample, and motion information required to generate a predicted block can be signaled to the decoder. The decoder can decode information about the signaled difference value to derive a difference sample, and add a prediction sample within the predicted block generated using the motion information to the difference sample to generate a restored sample.

[0201] At this time, in order to effectively compress the motion information signaled to the decoder, one of a plurality of inter prediction modes may be selected. Here, the plurality of inter prediction modes may include a motion information merging mode and a motion vector prediction mode.

[0202] The motion vector prediction mode is a mode that signals by encoding the difference between a motion vector and a motion vector prediction value. Here, the motion vector prediction value can be derived based on motion information of neighboring blocks or neighboring samples adjacent to the current block.

[0203] Figure 19 shows the locations referenced to derive motion vector prediction values.

[0204] For convenience of explanation, the current block is assumed to have a size of 4x4.

[0205] In the illustrated example, 'LB' represents a sample contained in the leftmost column and bottommost row within the current block. 'RT' represents a sample contained in the rightmost column and topmost row within the current block. A0 to A4 represent samples neighboring to the left of the current block, and B0 to B5 represent samples neighboring to the top of the current block. For example, A1 represents a sample neighboring to the left of LB, and B1 represents a sample neighboring to the top of RT.

[0206] Col indicates the location of a sample neighboring the lower right of the current block within a co-located picture. A co-located picture is a picture different from the current picture, and information for specifying the co-located picture (e.g., a co-located picture index) can be explicitly encoded and signaled in the bitstream. Alternatively, a reference picture having a predefined reference picture index can be set as the co-located picture.

[0207] A block covering a Col position within a collocated picture may be referred to as a collocated block. If the position of the lower-right neighboring sample of the current block within the reference picture is unavailable, the Col position may be set to the position of the center sample or the upper-left sample of the current block.

[0208] The motion vector prediction value of the current block can be derived from at least one motion vector prediction candidate included in a motion vector prediction list.

[0209] The number of motion vector prediction candidates that can be inserted into the motion vector prediction list (i.e., the size of the list) may be predefined in the encoder and decoder. For example, the maximum number of motion vector prediction candidates may be 2.

[0210] A motion vector stored at the location of a neighboring sample adjacent to the current block or a scaled motion vector derived by scaling the motion vector can be inserted into the motion vector prediction list as a motion vector prediction candidate. At this time, the motion vector prediction candidates can be derived by scanning the neighboring samples adjacent to the current block in a predefined order.

[0211] For example, it is possible to check whether a motion vector is stored at each location in the order of A0 to A4. Then, according to the above scanning order, the first available motion vector found can be inserted into the motion vector prediction list as a motion vector prediction candidate.

[0212] As another example, in the order of A0 to A4, it is checked whether a motion vector is stored at each position, and the motion vector of the position that is found first and has the same reference picture as the current block can be inserted into the motion vector prediction list as a motion vector prediction candidate. If there is no neighboring sample that has the same reference picture as the current block, a motion vector prediction candidate can be derived based on the first found available vector. Specifically, the first found available motion vector can be scaled, and then the scaled motion vector can be inserted into the motion vector prediction list as a motion vector prediction candidate. At this time, the scaling can be performed based on the output order difference between the current picture and the reference picture (i.e., the POC difference) and the output order difference between the current picture and the reference picture of the neighboring sample (i.e., the POC difference).

[0213] Furthermore, it is possible to check whether a motion vector is stored at each location in the order of B0 to B5. Then, according to the above scanning order, the first available motion vector found can be inserted into the motion vector prediction list as a motion vector prediction candidate.

[0214] As another example, in the order of B0 to B5, it is checked whether a motion vector is stored at each position, and the motion vector of the position that has the same reference picture as the current block that is found first can be inserted into the motion vector prediction list as a motion vector prediction candidate. If there is no neighboring sample that has the same reference picture as the current block, a motion vector prediction candidate can be derived based on the first found available vector. Specifically, the first found available motion vector can be scaled, and then the scaled motion vector can be inserted into the motion vector prediction list as a motion vector prediction candidate. At this time, the scaling can be performed based on the output order difference between the current picture and the reference picture (i.e., the POC difference) and the output order difference between the current picture and the reference picture of the neighboring sample (i.e., the POC difference).

[0215] As in the example described above, a motion vector prediction candidate can be derived from a sample adjacent to the left of the current block, and a motion vector prediction candidate can be derived from a sample adjacent to the top of the current block.

[0216] At this time, the motion vector prediction candidate derived from the left sample may be inserted into the motion vector prediction list before the motion vector prediction candidate derived from the upper sample. In this case, the index assigned to the motion vector prediction candidate derived from the left sample may have a smaller value than the motion vector prediction candidate derived from the upper sample.

[0217] Conversely, the motion vector prediction candidate derived from the top sample may be inserted into the motion vector prediction list before the motion vector prediction candidate derived from the left sample.

[0218] Among the motion vector prediction candidates included in the above motion vector prediction list, the motion vector prediction candidate with the highest encoding efficiency can be set as the motion vector predictor (MVP) of the current block. In addition, index information indicating the motion vector prediction candidate set as the motion vector predictor of the current block among the plurality of motion vector prediction candidates can be encoded and signaled to a decoder. When the number of motion vector prediction candidates is two, the index information can be a 1-bit flag (e.g., an MVP flag). In addition, a motion vector difference (MVD), which is the difference between the motion vector of the current block and the motion vector predictor, can be encoded and signaled to a decoder.

[0219] The decoder can construct a motion vector prediction list, similar to the encoder. Furthermore, it can decode index information from the bitstream and select one of multiple motion vector prediction candidates based on the decoded index information. The selected motion vector prediction candidate can be set as the motion vector prediction value of the current block.

[0220] Additionally, the motion vector differential can be decoded from the bitstream. Then, the motion vector of the current block can be derived by combining the motion vector prediction value and the motion vector differential value.

[0221] When bidirectional prediction is applied to the current block, a motion vector prediction list can be generated for each of the L0 and L1 directions. That is, the motion vector prediction list can be composed of motion vectors in the same direction. Accordingly, the motion vector of the current block and the motion vector prediction candidates included in the motion vector prediction list have the same direction.

[0222] When the motion vector prediction mode is selected, reference picture index and prediction direction information can be explicitly encoded and signaled to the decoder. For example, when there are multiple reference pictures in the reference picture list and motion estimation is performed for each of the multiple reference pictures, a reference picture index for specifying a reference picture from which motion information of the current block is derived among the multiple reference pictures can be explicitly encoded and signaled to the decoder.

[0223] At this time, if the reference picture list contains only one reference picture, encoding / decoding of the reference picture index may be omitted.

[0224] The prediction direction information may be an index pointing to one of L0 unidirectional prediction, L1 unidirectional prediction, or bidirectional prediction. Alternatively, an L0 flag indicating whether prediction is performed in the L0 direction and an L1 flag indicating whether prediction is performed in the L1 direction may be encoded and signaled, respectively.

[0225] Motion Information Merge Mode is a mode in which the motion information of the current block is set to be identical to the motion information of neighboring blocks. In Motion Information Merge Mode, motion information can be encoded / decoded using a motion information merge list.

[0226] Motion information merging candidates can be derived based on motion information from neighboring blocks or neighboring samples adjacent to the current block. For example, after defining reference locations around the current block, it is possible to check whether motion information exists at the defined reference locations. If motion information exists at the defined reference locations, the motion information at those locations can be inserted into the motion information merging list as a motion information merging candidate.

[0227] In the example of Fig. 19, the predefined reference positions may include at least one of A0, A1, B0, B1, B5, and Col. Furthermore, motion information merging candidates may be derived in the order of A1, B1, B0, A0, B5, and Col.

[0228] Among the motion information merge candidates included in the motion information merge list, the motion information of the motion information merge candidate with the optimal cost can be set as the motion information of the current block. Furthermore, index information (e.g., a merge index) indicating the motion information merge candidate selected from among the multiple motion information merge candidates can be encoded and transmitted to the decoder.

[0229] In the decoder, a motion information merge list can be constructed in the same manner as in the encoder. Furthermore, motion information merge candidates can be selected based on the merge index decoded from the bitstream. The motion information of the selected motion information merge candidate can be set as the motion information of the current block.

[0230] Unlike the motion vector prediction list, the motion information merge list is composed of a single list regardless of the prediction direction. That is, the motion information merge candidates included in the motion information merge list may have only L0 motion information or only L1 motion information, or may have bidirectional motion information (i.e., L0 motion information and L1 motion information).

[0231] Motion information about the current block can also be derived using the restoration sample area surrounding the current block. Here, the restoration sample area used to derive motion information about the current block can be referred to as a template.

[0232] Figure 20 is a diagram for explaining a template-based motion estimation method.

[0233] In Fig. 16, it is described that the predicted block of the current block is determined based on the cost between the current block and the reference block within the search range. According to the present embodiment, unlike Fig. 16, motion estimation for the current block can be performed based on the cost between a template neighboring the current block (hereinafter referred to as the "current template") and a reference template having the same size and shape as the current template.

[0234] For example, the cost can be calculated based on the absolute sum of the differences between the restored samples in the current template and the restored samples in the reference block. A smaller absolute sum can reduce the cost.

[0235] Once the reference template with the optimal cost is determined within the search range, the reference block neighboring the reference template can be set as the predicted block of the current block.

[0236] And, motion information of the current block can be set based on the distance between the current block and the reference block, the index of the picture to which the reference block belongs, and whether the reference picture is included in the L0 or L1 reference picture list.

[0237] Since the template defines the restored area around the current block as a template, the decoder itself can perform motion estimation in the same manner as the encoder. Accordingly, when deriving motion information using a template, there is no need to encode and signal the motion information other than information indicating whether a template is used.

[0238] The current template may include at least one of a region adjacent to the top of the current block or a region adjacent to the left of the current block. The region adjacent to the top may include at least one row, and the region adjacent to the left may include at least one column.

[0239] Figure 21 shows examples of template configurations.

[0240] The current template can be configured according to one of the examples illustrated in FIG. 21.

[0241] Alternatively, unlike the example illustrated in FIG. 21, the template may be configured only with the area adjacent to the left of the current block, or only with the area adjacent to the top of the current block.

[0242] The size and / or shape of the current template may be predefined in the encoder and decoder.

[0243] Alternatively, after defining a plurality of template candidates having different sizes and / or shapes, index information specifying one of the plurality of template candidates can be encoded and signaled to a decoder.

[0244] Alternatively, one of multiple template candidates can be adaptively selected based on at least one of the size, shape, or position of the current block. For example, if the current block borders the upper boundary of the CTU, the current template can be constructed using only the region adjacent to the left of the current block.

[0245] Template-based motion estimation can be performed for each of the reference pictures stored in the reference picture list. Alternatively, motion estimation can be performed only for some of the reference pictures. For example, motion estimation can be performed only for reference pictures with a reference picture index of 0, or only for reference pictures with a reference picture index less than a threshold, or only for reference pictures with a POC difference from the current picture less than a threshold.

[0246] Alternatively, a reference picture index can be explicitly encoded and signaled, and then motion estimation can be performed only for the reference picture pointed to by the reference picture index.

[0247] Alternatively, motion estimation can be performed on the reference picture of the neighboring block corresponding to the current template. For example, if the template consists of a left adjacent region and an upper adjacent region, at least one reference picture can be selected using at least one of the reference picture index of the left adjacent block or the reference picture index of the upper adjacent block. Thereafter, motion estimation can be performed on the selected at least one reference picture.

[0248] Information indicating whether template-based motion estimation is applied can be encoded and signaled to a decoder. The information can be a 1-bit flag. For example, a true (1) flag indicates that template-based motion estimation is applied in the L0 direction and L1 direction of the current block. On the other hand, a false (0) flag indicates that template-based motion estimation is not applied. In this case, motion information of the current block can be derived based on a motion information merging mode or a motion vector prediction mode.

[0249] Conversely, template-based motion estimation may be applied only when it is determined that neither the motion information merging mode nor the motion vector prediction mode is applied to the current block. For example, if the first flag indicating whether the motion information merging mode is applied and the second flag indicating whether the motion vector prediction mode is applied are both 0, template-based motion estimation may be performed.

[0250] For each of the L0 and L1 directions, information indicating whether template-based motion estimation is applied can be signaled. That is, whether template-based motion estimation is applied in the L0 direction and whether template-based motion estimation is applied in the L1 direction can be determined independently. Accordingly, template-based motion estimation can be applied to one of the L0 and L1 directions, while another mode (e.g., motion information merging mode or motion vector prediction mode) can be applied to the other.

[0251] When template-based motion estimation is applied to both the L0 direction and the L1 direction, the prediction block of the current block can be generated based on a weighted sum operation of the L0 prediction block and the L1 prediction block. Alternatively, even when template-based motion estimation is applied to one of the L0 direction and the L1 direction, but another mode is applied to the other direction, the prediction block of the current block can be generated based on a weighted sum operation of the L0 prediction block and the L1 prediction block.

[0252] Alternatively, a template-based motion estimation method may be inserted as a motion information merging candidate in a motion information merging mode or a motion vector prediction candidate in a motion vector prediction mode. In this case, whether or not to apply a template-based motion estimation method may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate indicates a template-based motion estimation method.

[0253] Based on the two-way matching method, it is also possible to generate movement information of the current block.

[0254] Figure 22 is a diagram for explaining a motion estimation method based on a bilateral matching method.

[0255] The bilateral matching method can be performed only when the temporal order (i.e., POC) of the current picture exists between the temporal order of the L0 reference picture and the temporal order of the L1 reference picture.

[0256] When a bilateral matching method is applied, a search range can be set for each of the L0 reference picture and the L1 reference picture. At this time, an L0 reference picture index for identifying the L0 reference picture and an L1 reference picture index for identifying the L1 reference picture can be encoded and signaled, respectively.

[0257] As another example, only the L0 reference picture index may be encoded and signaled, and an L1 reference picture may be selected based on the distance between the current picture and the L0 reference picture (hereinafter referred to as the L0 POC difference). For example, among the L1 reference pictures included in the L1 reference picture list, an L1 reference picture having an absolute value of the distance from the current picture (hereinafter referred to as the L1 POC difference) equal to the absolute value of the distance between the current picture and the L0 reference picture may be selected. If there is no L1 reference picture having an L1 POC difference equal to the L0 POC difference, an L1 reference picture having an L1 POC difference most similar to the L0 POC difference may be selected among the L1 reference pictures.

[0258] At this time, among the L1 reference pictures, only L1 reference pictures that are temporally different from the L0 reference picture can be used for bilateral matching. For example, if the POC of the L0 reference picture is smaller than that of the current picture, one of the L1 reference pictures that has a POC larger than that of the current picture can be selected.

[0259] Conversely, one could also encode and signal only the L1 reference picture index, and select the L0 reference picture based on the distance between the current picture and the L1 reference picture.

[0260] Alternatively, a bilateral matching method may be performed using the L0 reference picture having the closest distance to the current picture among the L0 reference pictures and the L1 reference picture having the closest distance to the current picture among the L1 reference pictures.

[0261] Alternatively, a bilateral matching method may be performed using an L0 reference picture (e.g., index 0) assigned with a predefined index in the L0 reference picture list and an L1 reference picture (e.g., index 0) assigned with a predefined index in the L1 reference picture list.

[0262] Alternatively, the LX (X is 0 or 1) reference picture may be selected based on an explicitly signaled reference picture index, and the L|X-1| reference picture may be selected as a reference picture having the closest distance to the current picture among the L|X-1| reference pictures, or as a reference picture having a predefined index in the L|X-1| reference picture list.

[0263] As another example, L0 and / or L1 reference pictures can be selected based on motion information of neighboring blocks of the current block. For example, L0 and / or L1 reference pictures to be used for bilateral matching can be selected using the reference picture index of the neighboring block to the left or above the current block.

[0264] The search range can be set within a predetermined range from a collocated block within a reference picture.

[0265] As another example, the search range can be set based on initial motion information. This initial motion information can be derived from neighboring blocks of the current block. For example, the motion information of the left or top neighboring block of the current block can be set as the initial motion information of the current block.

[0266] When the bilateral matching method is applied, the L0 motion vector and the L1 motion vector are set to have opposite directions. This indicates that the signs of the L0 motion vector and the L1 motion vector have opposite signs. In addition, the size of the LX motion vector can be proportional to the distance between the current picture and the LX reference picture (i.e., the POC difference).

[0267] Thereafter, motion estimation can be performed using the cost between a reference block belonging to the search range of the L0 reference picture (hereinafter referred to as an L0 reference block) and a reference block belonging to the search range of the L1 reference picture (hereinafter referred to as an L1 reference block).

[0268] If an L0 reference block whose vector to the current block is (x, y) is selected, an L1 reference block located at a distance of (-Dx, -Dy) from the current block can be selected. Here, D can be determined by the ratio of the distance between the current picture and the L0 reference picture and the distance between the L1 reference picture and the current picture.

[0269] For example, in the example illustrated in FIG. 22, the absolute value of the distance between the current picture (T) and the L0 reference picture (T-1) and the absolute value of the distance between the current picture (T) and the L1 reference picture (T+1) are equal to each other. Accordingly, in the illustrated example, the L0 motion vector (x0, y0) and the L1 motion vector (x1, y1) have equal magnitudes but opposite distances. If an L1 reference picture with a POC of (T+2) were used, the L1 motion vector (x1, y1) would be set to (-2*x0, -2*y0).

[0270] Once the L0 reference block and L1 reference block with the optimal cost are selected, the L0 reference block and the L1 reference block can be set as the L0 prediction block and the L1 prediction block of the current block, respectively. Thereafter, the final prediction block of the current block can be generated through a weighted sum operation of the L0 reference block and the L1 reference block.

[0271] When the bilateral motion matching method is applied, the decoder can perform motion estimation in the same manner as the encoder. Accordingly, information indicating whether the bilateral motion matching method is applied can be explicitly encoded / decoded, while encoding / decoding of motion information such as motion vectors can be omitted. As previously explained, at least one of the L0 reference picture index or L1 reference picture index may also be explicitly encoded / decoded.

[0272] As another example, information indicating whether a bilateral matching method is applied may be explicitly encoded / decoded, and if a bilateral matching method is applied, the L0 motion vector or the L1 motion vector may be explicitly encoded and signaled. If the L0 motion vector is signaled, the L1 motion vector may be derived based on the POC difference between the current picture and the L0 reference picture and the POC difference between the current picture and the L1 reference picture. If the L1 motion vector is signaled, the L0 motion vector may be derived based on the POC difference between the current picture and the L0 reference picture and the POC difference between the current picture and the L1 reference picture. In this case, the encoder may explicitly encode the smaller one of the L0 motion vector and the L1 motion vector.

[0273] Information indicating whether a bilateral matching method has been applied may be a 1-bit flag. For example, a true flag (e.g., 1) may indicate that a bilateral matching method has been applied to the current block. A false flag (e.g., 0) may indicate that a bilateral matching method has not been applied to the current block. In this case, the current block may be subject to motion information merging mode or motion vector prediction mode.

[0274] Conversely, the bilateral matching method may be applied only when it is determined that neither the motion information merging mode nor the motion vector prediction mode is applied to the current block. For example, the bilateral matching method may be applied when both the first flag indicating whether the motion information merging mode is applied and the second flag indicating whether the motion vector prediction mode is applied are 0.

[0275] Alternatively, the bilateral matching method may be inserted as a motion information merging candidate in the motion information merging mode or as a motion vector prediction candidate in the motion vector prediction mode. In this case, whether the bilateral matching method is applied may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate indicates the bilateral matching method.

[0276] In the bidirectional matching method, it is exemplified that the temporal order of the current picture must exist between the temporal order of the L0 reference picture and the temporal order of the L1 reference picture. A unidirectional matching method, which is not subject to the constraints of the above bidirectional matching method, may also be applied to generate a prediction block of the current block. Specifically, in the unidirectional matching method, two reference pictures having a temporal order (i.e., POC) smaller than that of the current block or two reference pictures having a temporal order larger than that of the current block may be used. In this case, both reference pictures may be derived from the L0 reference picture list or the L1 reference picture list. Alternatively, one of the two reference pictures may be derived from the L0 reference picture list and the other may be derived from the L1 reference picture list.

[0277] Figure 23 is a diagram for explaining a motion estimation method based on a one-way matching method.

[0278] The unidirectional matching method can be performed based on two reference pictures having a POC smaller than that of the current picture (i.e., forward reference pictures) or two reference pictures having a POC larger than that of the current picture (i.e., backward reference pictures). In Fig. 23, motion estimation based on the unidirectional matching method is exemplified as being performed based on a first reference picture (T-1) and a second reference picture (T-2) having a POC smaller than that of the current picture (T).

[0279] At this time, a first reference picture index for identifying a first reference picture and a second reference picture index for identifying a second reference picture may be encoded and signaled, respectively. At this time, among the two reference pictures used in the unidirectional matching method, a reference picture having a smaller POC difference from the current picture may be set as the first reference picture. Accordingly, when the first reference picture is selected, only reference pictures included in the reference picture list having a larger POC difference from the current picture than the first reference picture may be set as the second reference picture. The second reference picture index may be set to point to an index of one of the rearranged reference pictures after rearranging reference pictures having the same temporal direction as the first reference picture and having a larger POC difference from the current picture than the first reference picture.

[0280] Conversely, the reference picture with a larger POC difference from the current picture among the two reference pictures can be set as the first reference picture. In this case, the second reference picture index can be set to point to the index of one of the rearranged reference pictures after rearranging the reference pictures that have the same temporal direction as the first reference picture and have a smaller POC difference from the current picture than the first reference picture.

[0281] Alternatively, a one-way matching method may be performed using a reference picture assigned with a predefined index within a reference picture list and a reference picture having the same temporal direction as the reference picture. For example, a reference picture having an index of 0 within the reference picture list may be set as the first reference picture, and a reference picture having the smallest index among the reference pictures having the same temporal direction as the first reference picture within the reference picture list may be selected as the second reference picture.

[0282] Both the first reference picture and the second reference picture can be selected from the L0 reference picture list or the L1 reference picture list. In Fig. 23, two L0 reference pictures are illustrated as being used in the unidirectional matching method. Alternatively, the first reference picture may be selected from the L0 reference picture list, and the second reference picture may be selected from the L1 reference picture list.

[0283] Information indicating whether the first reference picture and / or the second reference picture belongs to the L0 reference picture list or the L1 reference picture list may be additionally encoded / decoded.

[0284] Alternatively, one-way matching can be performed using one of the L0 reference picture list and the L1 reference picture list, whichever is set as default. Alternatively, two reference pictures can be selected from the L0 reference picture list and the L1 reference picture list, whichever has a larger number of reference pictures.

[0285] Afterwards, a search range within the first reference picture and the second reference picture can be set.

[0286] The search range can be set within a predetermined range from a collocated block within a reference picture.

[0287] As another example, the search range can be set based on initial motion information. This initial motion information can be derived from neighboring blocks of the current block. For example, the motion information of the left or top neighboring block of the current block can be set as the initial motion information of the current block.

[0288] Thereafter, motion estimation can be performed using the cost between the first reference block belonging to the search range of the first reference picture and the second reference block belonging to the search range of the second reference picture.

[0289] At this time, under the unidirectional matching method, the size of the motion vector should be set to increase in proportion to the distance between the current picture and the reference picture. Specifically, if a first reference block whose vector with the current picture is (x, y) is selected, the second reference block should be spaced apart from the current block by (Dx, Dy). Here, D can be determined by the ratio of the distance between the current picture and the first reference picture and the distance between the current picture and the second reference picture.

[0290] For example, in the example of FIG. 23, the distance between the current picture and the first reference picture (i.e., the POC difference) is 1, and the distance between the current picture and the second reference picture (i.e., the POC difference) is 2. Accordingly, when the first motion vector for the first reference block in the first reference picture is (x0, y0), the second motion vector (x1, y1) for the second reference block in the second reference picture can be set to (2x0, 2y0).

[0291] Once the first and second reference blocks with optimal costs are selected, the first and second reference blocks can be set as the first and second prediction blocks of the current block, respectively. Thereafter, a weighted sum operation of the first and second prediction blocks can be performed to generate the final prediction block of the current block.

[0292] When a unidirectional motion matching method is applied, the decoder can perform motion estimation in the same manner as the encoder. Accordingly, information indicating whether a unidirectional motion matching method is applied can be explicitly encoded / decoded, while encoding / decoding of motion information such as motion vectors can be omitted. As previously explained, at least one of the first reference picture index or the second reference picture index may also be explicitly encoded / decoded.

[0293] As another example, information indicating whether a unidirectional matching method is applied may be explicitly encoded / decoded, and if a unidirectional matching method is applied, the first motion vector or the second motion vector may be explicitly encoded and signaled. If the first motion vector is signaled, the second motion vector may be derived based on the POC difference between the current picture and the first reference picture and the POC difference between the current picture and the second reference picture. If the second motion vector is signaled, the first motion vector may be derived based on the POC difference between the current picture and the first reference picture and the POC difference between the current picture and the second reference picture. In this case, the encoder may explicitly encode a smaller one of the first motion vector and the second motion vector.

[0294] Information indicating whether a unidirectional matching method has been applied may be a 1-bit flag. For example, a true flag (e.g., 1) may indicate that a unidirectional matching method has been applied to the current block. A false flag (e.g., 0) may indicate that a unidirectional matching method has not been applied to the current block. In this case, the current block may be subject to motion information merging mode or motion vector prediction mode.

[0295] Conversely, the one-way matching method may be applied only when it is determined that neither the motion information merging mode nor the motion vector prediction mode is applied to the current block. For example, the one-way matching method may be applied when the first flag indicating whether the motion information merging mode is applied and the second flag indicating whether the motion vector prediction mode is applied are both 0.

[0296] Alternatively, the unidirectional matching method may be inserted as a motion information merging candidate in the motion information merging mode or as a motion vector prediction candidate in the motion vector prediction mode. In this case, whether the unidirectional matching method is applied may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate indicates the unidirectional matching method.

[0297] When detecting motion between frames, the precision of the motion vector can be adjusted. Specifically, the location of each sample within a picture is defined as an integer position. However, the position reflecting the motion may be a decimal position, not an integer position.

[0298] Taking this into account, we can search for motion vectors more precisely through reference picture interpolation.

[0299] Figures 24 and 25 are diagrams showing examples in which prediction blocks are derived according to motion vector precision.

[0300] Fig. 24 shows the position of the current block within the current picture, and Fig. 13 shows the position of the reference block according to the motion vector precision.

[0301] As in the examples illustrated in FIGS. 24 and 25, the motion vector of the current block can be defined as the distance from the sample corresponding to the upper left position of the current block in the reference picture to the sample corresponding to the upper left position of the reference block in the reference picture.

[0302] Figure 25 (a) illustrates a case where the motion vector precision of the current block is an integer pel, Figure 25 (b) illustrates a case where the motion vector precision of the current block is 1 / 2 pel, and Figure 25 (c) illustrates a case where the motion vector precision of the current block is 1 / 4 pel.

[0303] In Figure 25, motion vectors are expressed up to 1 / 4 vector precision, but motion vectors can also be expressed more precisely, such as 1 / 8, 1 / 16, or 1 / 32.

[0304] Meanwhile, information indicating the motion vector precision of the current block may be encoded and signaled. For example, the information may be an index identifying one of the motion vector precision candidates. Specifically, each of the motion vector precision candidates may be assigned a different index, and the information may indicate the index of the motion vector precision candidate applied to the current block.

[0305] By adjusting the precision of the motion vector used for inter-screen prediction, more precise motion vector search can be achieved. If the reference block indicated by the motion vector exists at a real number location, the samples at the real number location can be generated using samples at integer locations and an interpolation filter. Furthermore, motion vectors expressed as real numbers can be scaled up to integers and then encoded / decoded.

[0306] In a similar manner to inter-screen prediction, a block vector for the current block can be derived, and a predicted block for the current block can be obtained using the block vector. Specifically, after deriving a block vector for the current block, the block vector can be used to specify a reference block within the current picture. The reference block specified by the block vector can be set as the predicted block for the current block. In this way, the method of predicting the current block using a block vector can be referred to as an intra-screen block copy mode.

[0307] Figure 26 is a flowchart of a method for copying blocks within a screen performed by a video encoding / decoding device.

[0308] Referring to Fig. 26, first, a block vector for the current block can be derived (S2610).

[0309] In the encoder, the reference block with the highest similarity to the current block within the current picture can be searched for. The distance between the searched reference block and the current block can be defined as a block vector.

[0310] In deriving a block vector, all restoration areas within the current picture can be set as search areas.

[0311] Alternatively, for simplicity, the size of the search area can be limited to a predefined size / shape. Specifically, a search area of ​​a predefined size / shape can be set within the current picture, and reference blocks can be searched only within the set search area.

[0312] As another example, the size of the search area can be set based on the size of the coding tree unit. For example, a reference block can be searched only within a search area that is N times the size of the coding tree unit.

[0313] Meanwhile, at least one of the motion information merging mode, motion vector prediction mode, or template matching mode used for inter prediction can be applied to encoding / decoding a block vector.

[0314] For example, when the block vector merging mode is applied to the current block, a merging candidate can be derived from a spatial neighboring block adjacent to the current block. Here, the spatial neighboring block can refer to a neighboring block that covers a sample at a location adjacent to the current block (e.g., one of A0 to A4 or B0 to B5 in FIG. 14). For example, the spatial neighboring block can include at least one of an upper neighboring block, a left neighboring block, an upper-right neighboring block, an upper-left neighboring block, or a lower-left neighboring block.

[0315] If a block vector is stored at the location of a spatial neighboring block of the current block (i.e., if the neighboring block is encoded / decoded in the block copy mode within the screen), the block vector stored at that location can be set as the block vector of the merge candidate.

[0316] On the other hand, if a block vector is not stored at the location of a spatial neighboring block, the spatial neighboring block may not be available as a merge candidate.

[0317] A block vector merge list can be generated by inserting at least one merge candidate derived from at least one spatial neighboring block. If the number of merge candidates included in the block vector merge list is less than a threshold, at least one of the zero block vector, the pair-wise block vector candidate, or the candidate included in the IBC HMVP list can be inserted into the block vector merge list as a merge candidate.

[0318] When the block vector merge mode is applied to the current block, the block vector of the current block may be identical to the block vector of the merge candidate. Accordingly, an index indicating one of the merge candidates included in the block vector merge list may be encoded and signaled.

[0319] As another example, when a block prediction mode is applied to the current block, block vector prediction candidates can be derived from spatial neighboring blocks adjacent to the current block. The spatial neighboring blocks adjacent to the current block may include at least one of an upper neighboring block, a left neighboring block, an upper-right neighboring block, an upper-left neighboring block, or a lower-left neighboring block.

[0320] If a block vector is stored at the location of a spatial neighboring block of the current block (i.e., if the neighboring block is encoded / decoded in the block copy mode within the screen), the block vector stored at that location can be set as a block vector prediction candidate.

[0321] On the other hand, if a block vector is not stored at the location of a spatial neighboring block, the spatial neighboring block may not be available as a block vector prediction candidate.

[0322] A block vector prediction list can be generated by inserting at least one block vector prediction candidate derived from at least one spatial neighboring block. If the number of block vector prediction candidates included in the block vector prediction list is less than a threshold, at least one of the zero block vector, the pair-wise block vector candidate, or the candidate included in the IBC HMVP list can be inserted into the block vector merge list as a block vector prediction candidate.

[0323] When the block vector prediction mode is applied to the current block, a block vector prediction candidate may be set as the block vector prediction value of the current block. When there are multiple block vector prediction candidates included in the block vector merge list, an index indicating one of the block vector prediction candidates may be encoded and signaled.

[0324] The block vector of the current block can be derived by adding the block vector prediction value and the block vector differential value. In the encoder, the block vector differential value can be encoded and signaled.

[0325] Meanwhile, based on at least one of the size / shape or color components of the current block, it can be determined whether to generate a merge candidate or a block vector prediction candidate using a spatial neighboring block.

[0326] For example, if the size of the current block is 4x4, the process of deriving a merge candidate or a block vector prediction candidate using a spatial neighboring block may be omitted. On the other hand, if the size of the current block is larger than 4x4, the process of generating a merge candidate or a block vector prediction candidate using a spatial neighboring block may be performed. That is, only when the size of the current block is larger than 4x4, a merge candidate or a block vector prediction candidate derived from a spatial neighboring block may be inserted into the list.

[0327] Temporal neighboring blocks (i.e., collocated blocks within a collocated picture) may not be used to derive merge candidates or block vector prediction candidates. That is, block vectors may be derived based on data (i.e., spatial data) contained in the same picture as the current block.

[0328] As another example, the availability of temporal neighboring blocks may be determined based on the type of the slice or picture containing the current block. For example, if the slice or picture containing the current block is of type P or B, the temporal neighboring blocks may be available for deriving merge candidates or block vector prediction candidates. On the other hand, if the slice or picture containing the current block is of type I, the temporal neighboring blocks may not be available for deriving merge candidates or block vector prediction candidates.

[0329] As another example, template matching can be used to derive the block vector of the current block.

[0330] Figure 27 is for explaining an example of deriving a block vector using template matching.

[0331] As in the example illustrated in Fig. 27, in the previously restored area within the current picture, a reference template similar to the current template that includes the previously restored area around the current block is searched.

[0332] Once the reference template with the lowest template matching cost is determined, the distance between the current template and the reference template can be set as a block vector.

[0333] Meanwhile, the search for the reference template can be performed targeting the entire restored region within the current picture.

[0334] Alternatively, reference templates can be searched only within the search area. As in the example described above, the search area can be set to have a predefined size / shape, or can be set to N times the size of the coding tree unit.

[0335] Once the block vector of the current block is determined, a predicted block of the current block can be derived based on the block vector (S2620).

[0336] Figure 28 shows an example of deriving a prediction block of the current block using a block vector.

[0337] As in the example illustrated in Fig. 28, a reference block within the current block can be specified using a block vector. The reference block indicated by the block vector can be set as a predicted block of the current block.

[0338]

[0339] When inter prediction is applied, a reference block within a reference picture can be specified based on the motion vector of the current block, and the reference block can be set as the prediction block of the current block. When intra block copy is applied, a reference block within the current picture can be specified based on the block vector of the current block, and the reference block can be set as the prediction block of the current block.

[0340] Both inter prediction and intra block copying have in common that they set a reference block, which is specified based on the vector components of the current block, as the prediction block of the current block.

[0341] At this time, the reference block can be modified based on the weight parameter, and the modified reference block can be set as the prediction block of the current block.

[0342] The weight parameters can be derived based on the previously reconstructed region around the current block and the previously reconstructed region around the reference block. The weight parameters can include filter coefficients for filtering samples within the reference block. For convenience of explanation, in the embodiments and drawings described below, it is assumed that the reference block is included in the current picture (i.e., the intra block copy mode is applied). However, even when the reference block is included in the reference picture (i.e., the inter prediction is applied), the prediction block acquisition method according to the embodiments described below can be applied.

[0343] Figure 29 illustrates a process for deriving a modified reference block.

[0344] Referring to Fig. 29, first, weight parameters can be derived through a filter coefficient learning step (S2910).

[0345] The weight parameters can be derived using the current template around the current block and the reference template around the reference block. For example, the current template may include a pre-restored region located above the current block and a pre-restored region located to the left of the current block, and the reference template may include a pre-restored region located above the reference block and a pre-restored region located to the left of the reference block.

[0346] Figure 30 illustrates the configuration of the current template and reference template.

[0347] As in the example illustrated in FIG. 30, the shape and size of the current template and the reference template may be the same. For convenience of explanation, it is assumed that the coordinates of the restoration samples included in the current template are determined by assuming that the position of the upper left sample of the current block is (0, 0). In addition, it is assumed that the coordinates of the restoration samples included in the reference template are determined by assuming that the position of the upper left sample of the reference block is (0, 0). In the illustrated example, it is exemplified that the current template and the reference template each include left restoration samples with x-axis coordinates of -1 to -4 and top restoration samples with y-axis coordinates of -1 to -4.

[0348] The restoration samples included in the current template and reference template are referred to as "reference samples." Unless otherwise specified, in the embodiments described below, the term "template" may be understood to refer to at least one of the current template and the reference template. Furthermore, the term "template" will also be used when describing embodiments that can be commonly applied to the current template and reference template.

[0349] Filter coefficients can be derived that minimize the differences between reference samples included in a reference template and reference samples within the current template. Specifically, when a filter is applied to a reference sample within a reference template, filter coefficients can be derived that minimize the differences with reference samples at the same location within the current template. In this case, the filter applied to the reference sample may be a convolution filter.

[0350] Figure 31 illustrates a convolution filter applied to a reference sample.

[0351] As in the example illustrated in Fig. 31, the output value for the convolution filter can be determined using a reference sample (C) and four adjacent samples (N, W, S, E).

[0352] Mathematical expression 5 shows an example for deriving filter coefficients.

[0353]

[0354] In Equation 5, [i, j] represents the coordinates of the reference sample. T represents the template region. That is, [i, j]∈T represents the coordinates of the reference sample within the current template or the reference template.

[0355] C represents a reference sample to which a filter is applied within the reference template. That is, C can represent a reference sample at position (i, j) within the reference template. Y represents a reference sample at the same position as C within the current template, that is, a reference sample at position (i, j) within the current template.

[0356] N, S, E, and W represent samples adjacent to the target reference sample C within the reference template. For example, N may represent a sample adjacent to the top of the target reference sample C, i.e., a sample at position [i, j-1]. S may represent a sample adjacent to the bottom of the target reference sample C, i.e., a sample at position [i, j+1]. W may represent a sample adjacent to the left of the target reference sample C, i.e., a sample at position [i-1, j]. E may represent a sample adjacent to the right of the target reference sample C, i.e., a sample at position [i+1, j].

[0357] B may be a value derived based on the bit depth of the picture. For example, mathematical expression 6 shows an example of deriving the variable B.

[0358]

[0359] In mathematical expression 6, D represents the bit depth. For example, if the bit depth is 10 bits, B can be set to 512, which is the midpoint of the range that can be expressed with 10 bits. Alternatively, if the bit depth is 8 bits, B can be set to 128, which is the midpoint of the range that can be expressed with 8 bits.

[0360] As another example, when applying a filter to a reference sample within the current template, a filter coefficient that minimizes the difference with respect to the reference sample at the same location within the reference template can be derived. In this case, in Equation 5, Y may represent a reference sample within the reference template, and C, N, S, W, and E may represent reference samples within the current template.

[0361] In mathematical expression 5, filter coefficients w0 to w5 that minimize the value of E can be derived. Regression analysis can be used to derive the filter coefficients.

[0362] At this time, after applying a smoothing filter to pixels existing in the reference template and the current template, filter coefficients may be derived. By applying the derived filter coefficients to the reference block, the reference block can be modified (S2920). Specifically, by performing a convolution using the filter coefficients on samples within the reference block, samples within the reference block can be modified. The modified samples can be set as predicted samples of the current block. Meanwhile, the filter applied to samples within the reference block may be the same as that illustrated in FIG. 31.

[0363] Mathematical expression 7 shows an example of modifying a restoration sample within a reference block.

[0364]

[0365] In the above mathematical expression 7, M represents a reference block. That is, [i,j]∈M represents the coordinates of a sample within the reference block.

[0366] Y' represents an improved (modified) sample within the reference block. That is, Y' can represent an improved sample at position (i, j) within the reference block, and the improved sample can be set as a predicted sample at position (i, j) within the current block.

[0367] C' represents a sample within the reference block. That is, C' can represent a sample at position (i, j) within the reference block.

[0368] N', S', E', and W' represent samples adjacent to sample C' within the reference block. For example, N' may represent a sample adjacent to the top of sample C, i.e., a sample at position [i, j-1]. S' may represent a sample adjacent to the bottom of sample C', i.e., a sample at position [i, j+1]. W' may represent a sample adjacent to the left of sample C', i.e., a sample at position [i-1, j]. E' may represent a sample adjacent to the right of target reference sample C', i.e., a sample at position [i+1, j].

[0369] Meanwhile, the filter applied to the reference template or reference block is not limited to the example illustrated in FIG. 31. For example, a 1D filter (1xN or Nx1 filter) or a square-shaped filter (e.g., NxN) may be used to derive filter coefficients or utilize improved samples. Here, N is a natural number and may be 1, 2, 3, 4, or 5.

[0370] An index indicating one of multiple filter candidates may be encoded and signaled. Alternatively, the filter type may be adaptively determined based on at least one of the size / shape of the current block or the encoding mode of the reference block.

[0371] Depending on the filter type, the number of filter coefficients included in the weight parameter may vary.

[0372] For example, the number of filter coefficients may be set differently depending on the size of the current block. For example, if the size of the current block is larger than a threshold, a filter type including a larger number of filter coefficients may be used compared to a case where the size of the current block is equal to or smaller than the threshold.

[0373] Alternatively, multiple thresholds may be defined to determine the filter type for the current block. Each section defined by the multiple thresholds may be mapped to a different filter type.

[0374] By comparing the current block size with multiple thresholds, we determine which interval the current block size falls within. After determining the interval to which the current block size belongs, we can derive a prediction sample for the current block using the filter type mapped to that interval.

[0375] As another example, the number of filter coefficients can be adaptively determined depending on the size of the current picture containing the current block.

[0376] Alternatively, instead of using a multi-tap filter as illustrated in FIG. 31, a single weight can be derived that minimizes the difference between the sum of the reference samples included in the reference template and the sum of the reference samples included in the current template. In this case, a predicted block of the current block can be obtained by multiplying the reference block by the weight.

[0377] In the above example, it is assumed that the filter applied to the reference template and the filter applied to the reference block are of the same type, but the types of the filter applied to the reference template and the filter applied to the reference block may be different.

[0378] As described above, the filter coefficient learning step is performed using the previously reconstructed region within the current picture. Accordingly, filter coefficients can be derived in the same manner in the encoder and decoder.

[0379] As another example, information about filter coefficient sets may be encoded and signaled to the decoder. The filter coefficient sets may include w0 to w5 as described above. As an example, an index indicating one of the multiple filter coefficient set candidates may be encoded and signaled.

[0380] Convolution is used during filter learning and filter application. Depending on the filter type, samples outside the template may need to be used.

[0381] For example, when applying the cross-shaped filter shown in Fig. 31, filtering should be performed on the reference sample at the (-4, -4) position located at the upper left of the template, using the sample at the (-5, -4) position and the sample at the (-4. -5) position outside the reference template.

[0382] For buffer optimization, samples outside the template can be treated as unavailable, or filtering can be performed using only reference samples within the template.

[0383] For example, a filter may not be applied to reference samples belonging to an edge within a reference template. That is, reference samples belonging to an edge within a reference template are not utilized as samples at position C in Equation 5, and only reference samples at other positions may be utilized as samples at position C in Equation 5.

[0384] As another example, the type of filter applied to reference samples located at the edge of the template may be different from the type of filter applied to reference samples located elsewhere.

[0385] For example, a horizontal 1D filter, for example, an Nx1 filter, may be applied to a reference sample that touches the upper or lower boundary of the template. A vertical 1D filter, for example, a 1xN filter, may be applied to a reference sample that touches the left or right boundary of the template.

[0386] As another example, filtering can be performed by replacing values ​​in samples that fall outside the reference template with values ​​from available reference samples.

[0387] Additionally, at least part of the reference template may lie outside a slice, tile, or picture boundary, resulting in unavailable samples within the reference template.

[0388] As above, if there are unavailable samples, padding can be performed to replace the values ​​of the unavailable samples with the values ​​of the available samples.

[0389] Figure 32 shows an example in which padding is performed for unavailable samples.

[0390] For convenience of explanation, it is assumed that samples outside the template are set as unavailable samples.

[0391] Padding can be performed by copying the sample closest to the unavailable sample. For example, the values ​​of pixels located to the left, right, top, or bottom of the unavailable sample can be copied to the unavailable sample location.

[0392] Meanwhile, when the number of adjacent available samples is multiple, padding for unavailable samples can be performed by averaging, weighting, or interpolating the adjacent available samples.

[0393] For example, for a location A outside the reference template, the left adjacent sample and the top adjacent sample are available. In this case, the average of the left adjacent sample and the top adjacent sample can be set as the sample value for location A.

[0394] Similarly, for position A' outside the current template, left-adjacent samples and top-adjacent samples are available. In this case, the average of the left-adjacent samples and top-adjacent samples can be set as the sample value at position A'.

[0395] As another example, the values ​​of unavailable samples can be derived by copying the values ​​of samples located in a predefined direction. Here, the predefined direction can be left, right, top, or bottom.

[0396] Meanwhile, the predefined direction may vary depending on the relative position between the unavailable sample and the template. For example, if the unavailable sample is located at the top of the template, the value of the bottom adjacent sample may be set to the value of the unavailable sample. Alternatively, if the unavailable sample is located on the left side of the template, the value of the right adjacent sample may be set to the value of the unavailable sample.

[0397] As another example, padding may be performed considering the encoding mode of the reference block. For example, if the reference block is encoded in a directional intra prediction mode, the values ​​of samples located in the forward or backward direction according to the directional intra prediction mode may be set to the values ​​of unavailable samples.

[0398] For example, when the upper right diagonal intra prediction mode is applied to the reference block, the value of a sample located in the upper right diagonal direction or the upper left diagonal direction at an unavailable sample location can be set to the value of the unavailable sample.

[0399] The same problem can arise when applying a filter to a reference block. For example, when filtering samples located at the edge of the reference block (i.e., samples located at the boundary of the reference block), samples located outside the reference block must be used. For example, applying the cross-shaped filter illustrated in Figure 31 to a 4x4 reference block requires samples within a 6x6 area.

[0400] If samples located outside the reference block are available, samples located at that location can be used to filter samples located at the edge within the reference block.

[0401] As another example, padding can be used to generate values ​​for samples located outside the reference block. That is, samples located outside the reference block (e.g., samples included in the reference template) can be set as unavailable, and values ​​for these unavailable samples can be generated through padding.

[0402] Figure 33 shows an example in which values ​​are generated for samples adjacent to a reference block through padding.

[0403] Padding can be performed by copying samples at adjacent locations within the reference block. For example, the value of a sample adjacent to the top of the reference block can be generated by copying the value of a sample bordering the top boundary within the reference block, and the value of a sample adjacent to the bottom of the reference block can be generated by copying the value of a sample bordering the bottom boundary within the reference block. Similarly, the value of a sample adjacent to the left of the reference block can be generated by copying the value of a sample bordering the left boundary within the reference block, and the value of a sample adjacent to the right of the reference block can be generated by copying the value of a sample bordering the right boundary within the reference block.

[0404] As another example, padding can be performed only on samples that are located outside the reference block but do not belong to the reference template. That is, samples that are located outside the reference block but belong to the reference template can be set to be available for filtering of the reference sample.

[0405] Figure 34 illustrates an example in which padding is performed only for sample locations that do not belong to the reference template among samples located outside the reference block.

[0406] When a reference template includes an upper base restoration area and a left base restoration area of ​​a reference block, samples bordering the upper boundary or the left boundary within the reference block can be filtered using reference samples belonging to the reference template. On the other hand, in order to filter samples bordering the right boundary or the bottom boundary within the reference block, samples bordering the right or bottom of the reference block can be generated through padding.

[0407] In another example, the type of filter to be applied to a sample within a reference block can be adaptively determined based on the availability of neighboring samples outside the reference block. For example, if the left or right neighboring samples of a target sample to which a filter is to be applied within the reference block are not available, a filter in the form of a vertical 1D filter (e.g., 1xN) can be applied to filter the target sample.

[0408] On the other hand, if the upper or lower neighboring samples of the target sample to which the filter is to be applied within the reference block are not available, a filter in the form of a horizontal 1D filter (e.g., 1xN) can be applied to filter the target sample. Here, N can be a natural number such as 1, 3, or 5.

[0409] In the above example, the modified reference block based on the weight parameter was set as the prediction block of the current block. Unlike the example described above, the reference block can be set as the prediction block of the current block, and then the prediction block can be modified using the weight parameter. The modified prediction block can be set as the final prediction block of the current block. In other words, the current block can be reconstructed by adding the residual block to the modified prediction block.

[0410] Meanwhile, a filter may be applied to the prediction sample of the current block to obtain a modified prediction sample. At this time, as described with reference to FIG. 33, the values ​​of the samples outside the current block may be derived by copying the values ​​of the outermost prediction samples within the current block. For example, the values ​​of the samples adjacent to the top of the current block may be generated by copying the values ​​of the prediction samples bordering the top boundary within the current block, and the values ​​of the samples adjacent to the bottom of the current block may be generated by copying the values ​​of the prediction samples bordering the bottom boundary within the current block. Similarly, the values ​​of the samples adjacent to the left of the current block may be generated by copying the values ​​of the prediction samples bordering the left boundary within the current block, and the values ​​of the samples adjacent to the right of the current block may be generated by copying the values ​​of the prediction samples bordering the right boundary within the current block.

[0411] As another example, since the upper adjacent area and the left adjacent area of ​​the current block are already restored, the upper restored sample or the left restored sample can be used to filter the prediction sample located at the upper boundary or the left boundary of the current block.

[0412] On the other hand, the right and bottom adjacent regions of the current block have not yet been decrypted. Accordingly, as described in Figure 34, padding can be used to generate samples adjacent to the right or bottom of the current block.

[0413] That is, as explained through Figure 34, samples adjacent to the upper boundary or left boundary of the current block can be generated through padding, while samples adjacent to the right boundary or lower boundary of the current block can be generated through padding, while samples adjacent to the upper boundary or left boundary of the current block can be generated through padding, while samples adjacent to the lower boundary or right boundary of the current block can be generated through padding.

[0414] As another example, the values ​​of unavailable samples surrounding the current block can be derived based on samples surrounding the reference block. For example, since the upper and left adjacent regions of the current block have been previously restored, filtering of predicted samples within the current block can be performed using the upper and left restored samples of the current block.

[0415] On the other hand, the right adjacent region and bottom adjacent region of the current block are not yet decoded. On the other hand, the right adjacent region and bottom adjacent region of the reference block indicated by the motion vector of the current block may already be decoded. Considering this, the value of the right adjacent sample or bottom adjacent sample of the reference block can be copied and set as the value of the right adjacent sample or bottom adjacent sample of the current block.

[0416] As another example, a prediction block of the current block can be obtained using multiple reference blocks and multiple reference templates. Specifically, for each of the multiple reference blocks, a weight parameter is derived based on each reference template. The reference block can be modified based on the derived weight parameter. Thereafter, a prediction block of the current block can be obtained by calculating an average or weighted sum of the modified reference blocks. For convenience of explanation, the number of reference blocks is assumed to be two.

[0417] FIG. 35 is a diagram for explaining an example of obtaining a prediction block of a current block using multiple reference blocks.

[0418] A first weight parameter for the first reference block can be derived using a first reference template adjacent to the first reference block, and a second weight parameter for the second reference block can be derived using a second reference template adjacent to the second reference block.

[0419] Based on the first weight parameter, the first reference block can be modified, and based on the second weight parameter, the second reference block can be modified.

[0420] Thereafter, based on the average or weighted sum operation of the modified first reference block and the modified second reference block, the prediction block of the current block can be obtained. Mathematical expression 8 expresses the above process in a formula.

[0421]

[0422] In the above mathematical expression 8, P'[i, j] represents a prediction sample at the position [i, j]. P[k, i, j] represents a sample at the position [i, j] within the kth reference block. For example, a value of k of 0 may represent a sample within the first reference block, and a value of k of 1 may represent a sample within the second reference block. In addition, w[k] represents a weight parameter for reference block k, and w[n] represents a parameter applied to B. B is a variable induced by the bit depth, and may be induced, for example, according to mathematical expression 6.

[0423] M represents a reference block. That is, [i, j]∈M can represent the coordinates of a sample within the reference block.

[0424] The weight parameter w[k] for the reference block may include multiple filter coefficients. For example, the multiple filter coefficients w[0] to w[n] may be derived based on the following mathematical expression (9).

[0425]

[0426] In the above mathematical expression 9, N represents a template. That is, [i, j]∈N represents the position of a sample within the template. In addition, Y[i, j] represents the sample at position [i, j] within the current template, and X[k, i, j] represents the sample at position [i, j] within the kth reference template. In mathematical expression 9, parameters w[0] to w[n] that make E minimal are derived. At this time, regression analysis can be used.

[0427] The plurality of reference blocks can be derived based on at least one of a merge list (e.g., a motion information merge list or a block vector merge list), a vector prediction list (e.g., a motion vector prediction list or a block vector prediction list), or template matching.

[0428] For example, among the plurality of reference blocks, a first reference block may be derived based on one of a merge list, a vector prediction list, and a template matching, and a second reference block may be derived based on another of the merge list, the vector prediction list, and the template matching.

[0429] Alternatively, multiple merge candidates can be selected from the merge list to determine multiple reference blocks. For example, a first reference block can be determined using the first merge candidate in the merge list, and a second reference block can be determined using the second merge candidate in the merge list. To specify the first and second merge candidates, two sets of index information can be encoded and signaled.

[0430] Alternatively, among the plurality of reference blocks, the first reference block may be determined through inter prediction, and the second reference block may be determined through intra block copying. In this case, the first weight parameter for the first reference block may be obtained based on the current template within the current picture and the reference template within the reference picture, while the second weight parameter for the second reference block may be obtained based on the current template within the current picture and the reference template within the current picture.

[0431] In the embodiments described above, the reference block can be derived based on at least one of a merge list (e.g., a motion information merge list or a block vector merge list), a vector prediction list (e.g., a motion vector prediction list or a block vector prediction list), or template matching.

[0432] Alternatively, an embodiment of correcting a reference block may be applied only when the reference block (or block vector or motion vector) is derived through template matching.

[0433] Alternatively, information indicating whether to modify the reference block may be encoded and signaled when deriving the predicted block. The information may be a 1-bit flag.

[0434] Meanwhile, whether to encode / decode the information may be determined based on at least one of the size / shape of the current block and the mode used to derive the reference block. For example, if the size of the current block (e.g., at least one of the width, height, or number of included samples) is less than a threshold, or if the mode used to derive the reference block is a template matching mode, the information may be encoded / decoded.

[0435] Meanwhile, in the above-described embodiments, the template is depicted as including both the upper region of the block and the left region of the block. In this case, the configuration of the template can be adaptively determined based on the size or shape of the current block.

[0436] For example, if the current block is a square shape with equal width and height, the template may include both the top area of ​​the block and the left area of ​​the block. On the other hand, if the current block is a rectangle with different widths and heights, the template may include only the top area of ​​the block or only the left area of ​​the block. For example, if the current block is a rectangle with a width greater than its height, the template may include only the top area of ​​the block. On the other hand, if the current block is a rectangle with a height greater than its width, the template may include only the left area of ​​the block.

[0437] In Fig. 31, the number of lines constituting the template (i.e., thickness) is illustrated as 4. The thickness of the template may be predefined in the encoder and decoder. Alternatively, the thickness of the template may be adaptively determined depending on the size of the current block. For example, if the size of the current block is equal to or greater than a threshold, the thickness of the template may be set to N. On the other hand, if the size of the current block is smaller than the threshold, the thickness of the template may be set to a value smaller than N (e.g., N-1, N-2, or N-3). N may be a natural number such as 2, 3, or 4.

[0438] The thickness of the upper area and the thickness of the left area may have different values.

[0439] In an encoder and decoder, multiple template candidates may be predefined, each having at least one different shape or thickness. In this case, an index indicating one of the multiple template candidates may be encoded and signaled. Depending on the template candidate selected by the index, at least one of the shape or size of the template may be determined.

[0440]

[0441] Applying the embodiments described above, focusing on the decoding or encoding process, to the encoding or decoding process is within the scope of the present disclosure. Changing the embodiments described above, in a given order, to a different order is also within the scope of the present disclosure.

[0442] Although the above-described disclosure is described based on a series of steps or a flowchart, this does not limit the chronological order of the invention, and may be performed simultaneously or in a different order as needed. In addition, each component (e.g., unit, module, etc.) constituting the block diagram in the above-described disclosure may be implemented as a hardware device or software, or multiple components may be combined to be implemented as a single hardware device or software. For example, the hardware device may include at least one of a processor for performing calculations, a memory for storing data, a transmitter for transmitting data, and a receiver for receiving data.

[0443] The above-described disclosure may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination.

[0444] In addition, according to the present disclosure, a computer-readable recording medium can be provided that stores a bitstream generated by the above-described encoding method. The bitstream can be transmitted by an encoding device, and a decoding device can receive the bitstream and decode an image.

[0445] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memories. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present disclosure, and vice versa.

[0446] The present disclosure may be applied to a computing or electronic device capable of encoding / decoding a video signal.

Claims

1. A step of deriving motion information or block vector for the current block; A step of deriving a reference block for the current block based on the motion information or the block vector; A step of deriving weight parameters based on a current template around the current block and a reference template around the reference block; and An image decoding method, comprising a step of generating a prediction block for the current block based on the weight parameter.

2. In paragraph 1, The above weight parameters include filter coefficients, A method for decoding an image, wherein the filter coefficients are characterized in that, when a filter is applied to a reference sample in the reference template, the difference between the reference sample and the current template is minimized.

3. In paragraph 1, An image decoding method, characterized in that the prediction sample of the current block is derived by modifying a sample in the reference block based on the weight parameter.

4. In paragraph 3, An image decoding method, characterized in that the modified sample within the reference block is obtained by filtering the sample using adjacent samples adjacent to the sample.

5. In paragraph 4, If the adjacent sample adjacent to the above sample is not available, An image decoding method, characterized in that the value of the adjacent sample is set to be the same as the value of the sample.

6. In paragraph 4, If the adjacent sample adjacent to the above sample is located outside the above reference block, A method for decoding an image, characterized in that the value is set to be the same as that of the restored sample at the corresponding location.

7. In paragraph 4, An image decoding method, characterized in that the filter applied to the above sample is a cross-shaped filter having the same width and height.

8. In paragraph 4, An image decoding method, characterized in that a filter type applied to the sample is adaptively determined according to the location of the sample.

9. A step of deriving motion information or block vector for the current block; A step of deriving a reference block for the current block based on the motion information or the block vector; A step of deriving weight parameters based on a current template around the current block and a reference template around the reference block; and A video encoding method, comprising a step of generating a prediction block for the current block based on the weight parameter.

10. In paragraph 9, The above weight parameters include filter coefficients, A video encoding method, characterized in that the filter coefficients minimize the difference between the reference sample in the current template and the reference sample in the reference template when the filter is applied to the reference sample in the current template.

11. In paragraph 9, A video encoding method, characterized in that the prediction sample of the current block is derived by modifying a sample in the reference block based on the weight parameter.

12. In paragraph 11, A video encoding method, characterized in that the modified sample within the reference block is obtained by filtering the sample using adjacent samples adjacent to the sample.

13. In paragraph 12, If the adjacent sample adjacent to the above sample is not available, A video encoding method, characterized in that the value of the adjacent sample is set to be the same as the value of the sample.

14. In paragraph 12, If the adjacent sample adjacent to the above sample is located outside the above reference block, A method for encoding an image, characterized in that the value is set to be the same as that of the restored sample at the corresponding location.

15. A step of deriving motion information or block vector for the current block; A step of deriving a reference block for the current block based on the motion information or the block vector; A step of deriving weight parameters based on a current template around the current block and a reference template around the reference block; and A computer-readable recording medium storing a bitstream generated by a video encoding method, the method comprising the step of generating a prediction block for the current block based on the weight parameter.

Citation Information

Patent Citations

  • Finger surfing

    KR1020220131484A

  • Connection assembly between shower head and hose

    KR102542701B1

  • Prediction systems and methods for video coding based on filtering nearest neighboring pixels

    US20190182482A1

  • Prediction method using current picture referencing mode, and video decoding device therefor

    US20220329855A1