Encoding method, decoding method, computer-readable storage medium, and transmission method

By integrating multiple motion vectors into the motion vector predictor candidate list, the method addresses inefficiencies in inter prediction modes, resulting in improved compression and transmission of high-resolution, high-quality images.

WO2025170344A1PCT designated stage Publication Date: 2025-08-14LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/001792
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2025-02-06
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing image compression technologies struggle to efficiently handle high-resolution, high-quality images, particularly in inter prediction modes, due to limitations in utilizing various motion vectors, leading to suboptimal compression and transmission performance.

Method used

Incorporating multiple motion vectors into the motion vector predictor candidate list during inter prediction, enhancing the construction of prediction blocks by including base and additional motion vectors, which improves inter-screen prediction performance.

Benefits of technology

This approach significantly enhances the efficiency of image compression and transmission by improving inter-screen prediction performance, particularly for high-resolution and high-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025001792_14082025_PF_FP_ABST
    Figure KR2025001792_14082025_PF_FP_ABST
Patent Text Reader

Abstract

A decoding method according to an aspect of the present disclosure comprises the steps of: acquiring image information from a bitstream; on the basis of the acquired image information, determining a prediction mode applied to the current block as a prediction mode that uses a motion vector predictor (MVP); configuring an MVP candidate list including MVP candidates for the current block; and generating a prediction block for the current block on the basis of at least one MVP candidate within the MVP candidate list, wherein the step of configuring the MVP candidate list comprises including an MVP candidate having a basic motion vector and an additional motion vector in the MVP candidate list.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding method, decoding method, computer-readable storage medium and transmission method The present disclosure relates to a method for encoding / decoding image information, a computer-readable storage medium for storing image information, and a method for transmitting image information. Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed. There are various technologies for image compression, such as inter prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and entropy coding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance, and these technologies can be used to effectively compress and transmit or store image data. Accordingly, a highly efficient image compression technology is required to effectively transmit, store, and play high-resolution, high-quality image information. The present disclosure provides a method for including a candidate including multiple motion vectors in a list of motion vector predictor candidates in a process of constructing a list of motion vector predictor candidates in an inter prediction mode. The present disclosure provides a method for improving inter-screen prediction performance by utilizing various motion vectors. A decoding method according to one aspect of the present disclosure comprises the steps of: obtaining image information from a bitstream; determining a prediction mode to be applied to a current block based on the obtained image information as a prediction mode using a motion vector predictor (MVP); constructing an MVP candidate list including MVP candidates for the current block; and generating a prediction block for the current block based on at least one MVP candidate in the MVP candidate list; wherein the step of constructing the MVP candidate list includes including an MVP candidate having a base motion vector and an additional motion vector in the MVP candidate list. An encoding method according to one aspect of the present disclosure includes the steps of: determining a prediction mode applied to a current block as a prediction mode using a motion vector predictor (MVP); constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on a final MVP candidate in the MVP candidate list; and encoding image information including information about the prediction mode; wherein the step of constructing the MVP candidate list includes including an MVP candidate having a base motion vector and an additional motion vector in the MVP candidate list. In a storage medium storing a bitstream according to one aspect of the present disclosure, an encoding method for generating the bitstream includes the steps of: determining a prediction mode applied to a current block as a prediction mode using a motion vector predictor (MVP); constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on a final MVP candidate in the MVP candidate list; and encoding image information including information about the prediction mode; wherein the step of constructing the MVP candidate list includes including an MVP candidate having a base motion vector and an additional motion vector in the MVP candidate list. A transmission method according to one aspect of the present disclosure comprises: a step of obtaining a bitstream for the image, wherein the bitstream is generated based on a step of determining a prediction mode applied to a current block as a prediction mode using a motion vector predictor (MVP); a step of constructing an MVP candidate list including MVP candidates for the current block; a step of generating a prediction block for the current block based on a final MVP candidate in the MVP candidate list; and a step of encoding image information including information about the prediction mode; and a step of transmitting the data including the bitstream; wherein the step of constructing the MVP candidate list includes including an MVP candidate having a base motion vector and an additional motion vector in the MVP candidate list. According to the present disclosure, in the process of constructing a motion vector predictor candidate list in an inter prediction mode, inter-screen prediction performance can be improved by utilizing various motion vectors by including candidates including multiple motion vectors in the list. The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below. FIG. 1 illustrates a video / image coding system according to the present disclosure. FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed. FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed. FIG. 4 illustrates an example of a video / image decoding method to which an embodiment of the present disclosure can be applied. FIG. 5 illustrates an example of a video / image encoding method to which an embodiment of the present disclosure can be applied. Figure 6 is a diagram showing an example of a search area used in intra template matching. FIG. 7 and FIG. 8 illustrate examples of inter prediction-based video / image encoding methods to which embodiments of the present disclosure can be applied. FIGS. 9 and 10 illustrate examples of inter-prediction based video / image decoding methods to which embodiments of the present disclosure can be applied. FIG. 11 exemplarily illustrates an inter prediction procedure to which an embodiment of the present disclosure can be applied. Figure 12 is a diagram showing examples of blocks used to construct a merge candidate list. Figure 13 is a diagram showing four movements that can be expressed in the affine movement model. Figure 14 is a diagram showing an example of a control point motion vector used in affine motion prediction. Figure 15 is an example showing inheritance of control point motion vectors. Figure 16 is a drawing showing an example of surrounding blocks for the current block. Figure 17 illustrates the process of SbTMVP. Figure 18 illustrates an example of GPM segments grouped at the same angle. Figure 19 illustrates the left and upper neighboring blocks used in CIIP weight derivation. Figure 20 illustrates available IPM candidates for GPM including inter and intra prediction. Figure 21 is a diagram illustrating a method for deriving a motion vector from template matching. Figure 22 is a diagram showing a case where multiple reference blocks are used in one embodiment. FIG. 23 is a flowchart illustrating an example of a method for including an MVP candidate including multiple motion vectors in an MVP candidate list in a decoding method or an encoding method according to one embodiment. FIG. 24 is a diagram illustrating an example of a method for constructing a list of MVP candidates in one embodiment. FIG. 25 is a diagram illustrating an example of a method for generating a candidate having multiple motion vectors according to one embodiment. FIG. 26 is a diagram illustrating another example of a method for generating a candidate having multiple motion vectors according to one embodiment. Figure 27 shows the distance between the current picture and reference pictures. Figure 28 is a drawing showing an example of the position of a reference block indicated by the motion vector of the current block. Figure 29 is a drawing showing an example of an offset applicable to one embodiment. Figure 30 is a diagram showing an example of a process of correcting a position by applying an offset within a corresponding block and deriving a motion vector at each corrected position. FIG. 31 is a diagram illustrating an example of a content streaming system to which an embodiment according to the present disclosure can be applied. The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Throughout the description of each drawing, similar reference numerals have been used to designate similar components. While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein. When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening. The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the versatile video coding (VVC) standard. In addition, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or the next generation of video / image coding standards (e.g., H.267 or H.268). This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other. In this specification, video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs that has a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively according to the CTU raster scan, while tiles within a picture may be arranged consecutively according to the tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture, which may be exclusively contained within a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture. A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows. In this specification, “A or B” can mean “only A,” “only B,” or “both A and B.” In other words, “A or B” in this specification can be interpreted as “A and / or B.” For example, “A, B or C” in this specification can mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C." In this specification, “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” Additionally, in this specification, the expressions “at least one of A or B” or “at least one of A and / or B” may be interpreted identically to “at least one of A and B.” Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.” Additionally, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously. FIG. 1 illustrates a video / image coding system according to the present disclosure. Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting device. The receiving device may include a receiving device, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component. A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data. An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream. The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The storage medium can be a computer-readable storage medium and can store data non-transitory. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device. The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device. The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit. FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed. Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component. The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units (PUs). For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure. For example, a single coding unit may be split into multiple coding units with deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient. The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. A sample can be used as a term corresponding to a pixel or pel of a picture (or image). The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231). The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information related to prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream. The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples, i.e., the reference samples, may be located in the neighborhood of the current block or may be located a certain distance away from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of the DC mode or the planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail in the prediction direction. However, this is merely an example, and a greater or lesser number of directional modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks. The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. Temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and reference pictures including temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference. The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal. The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square size and the same size, or can be applied to a block of a non-square variable size. The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately from quantized transform coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240). The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process. The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream. The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency. The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks in the current picture and transfer them to the intra prediction unit (222). Image information output in the form of a bitstream from the encoding device (200) can be transmitted to the decoding device (300). FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed. Image information transmitted in the form of a bitstream from the encoding device (200) can be received by the decoding device (300). Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component. When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit of decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device. The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded and obtained from the bitstream through the decoding procedure. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310). Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331). The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients. In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array). The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra-prediction or inter-prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter-prediction mode. The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be signaled and included in the video / image information. The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks. The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block. The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the restoration block. The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process. The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of blocks in an already reconstructed picture. The stored motion information can be transmitted to the inter prediction unit (332) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra prediction unit (331). In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively. FIG. 4 illustrates an example of a video / image decoding method to which an embodiment of the present disclosure can be applied. In image / video coding, the pictures that make up an image / video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, and based on this, not only forward prediction but also backward prediction can be performed during inter prediction. In FIG. 4, S400 may be performed in the entropy decoding unit (310) of the aforementioned decoding device (300), S410 may be performed in the prediction unit (330), S420 may be performed in the residual processing unit (320), S430 may be performed in the addition unit (340), and S440 may be performed in the filtering unit (350). S400 may include a decoding procedure according to the present disclosure, S410 may include an inter / intra prediction procedure according to the present disclosure, S420 may include a residual processing procedure according to the present disclosure, S430 may include a block / picture restoration procedure according to the present disclosure, and S440 may include an in-loop filtering procedure according to the present disclosure. Referring to FIG. 4, the decoding device obtains image / video information from a bitstream (S400), performs prediction based on the obtained image / video information (S410), and restores a picture through residual processing (S420, inverse quantization for quantized transform coefficients, inverse transformation) (S430). A modified restored picture can be generated by applying an in-loop filtering procedure (S440) to a restored picture generated through the above restoration procedure, and the modified restored picture can be output as a decoded picture and can be stored in a buffer or memory of a decoding device to be used as a reference picture in an inter prediction procedure when decoding a next picture. In some cases, the in-loop filtering procedure can be omitted, in which case the restored picture can be output as a decoded picture and can be stored in a buffer or memory of a decoding device to be used as a reference picture in an inter prediction procedure when decoding a next picture. The in-loop filtering procedure (S440) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bi-lateral filter procedure, and some or all of them may be omitted. In addition, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bi-lateral filter procedure may be sequentially applied, or all of them may be sequentially applied. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture. Or, for example, the ALF procedure may be performed after the deblocking filtering procedure is applied to the restored picture. This may also be performed in an encoding device. FIG. 5 illustrates an example of a video / image encoding method to which an embodiment of the present disclosure can be applied. In FIG. 5, the prediction step (S500) may be performed in the prediction unit (220) of the encoding device (200) described above, residual processing (S510) based on the prediction result may be performed in the residual processing unit (230), and the step (S520) of encoding image information including prediction information and residual information may be performed in the entropy encoding unit (240). S500 may include an inter / intra prediction procedure according to the present disclosure, S510 may include a residual processing procedure according to the present disclosure, and S520 may include an encoding procedure according to the present disclosure. The encoding procedure may optionally include a procedure for encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, as well as a procedure for generating a restored picture for the current picture and a procedure for applying in-loop filtering to the restored picture. The encoding device (200) can derive (corrected) residual samples from the quantized transform coefficients through the inverse quantization unit (234) and the inverse transformation unit (235), and can generate a restored picture based on the prediction samples and (corrected) residual samples, which are outputs of S500. The restored picture generated in this way can be the same as the restored picture generated by the decoding device (300) described above. A modified restored picture can be generated through an in-loop filtering procedure for the restored picture, which can be stored in a buffer or memory, and, as in the case of the decoding device, can be used as a reference picture in the inter prediction procedure when encoding a subsequent picture. As described above, some or all of the in-loop filtering procedure may be omitted in some cases. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) may be encoded by the entropy encoding unit (240) and output in the form of a bitstream, and the decoding device (300) may perform the in-loop filtering procedure in the same manner as the encoding device based on the filtering-related information. Through this in-loop filtering procedure, noise occurring during image / video coding, such as blocking artifacts and ringing artifacts, can be reduced, and subjective / objective image quality can be improved. In addition, by performing the in-loop filtering procedure in both the encoding device (200) and the decoding device (300), the same prediction results can be derived from the encoding device (200) and the decoding device (300), thereby increasing the reliability of picture coding and reducing the amount of data that must be transmitted for picture coding. As described above, the picture restoration procedure can be performed not only in the decoding device (300) but also in the encoding device (200). A restoration block can be generated based on intra-prediction / inter-prediction for each block, and a restoration picture including the restoration blocks can be generated. If the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based only on intra-prediction. On the other hand, if the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra-prediction or inter-prediction. In this case, inter-prediction may be applied to some blocks in the current picture / slice / tile group, and intra-prediction may be applied to some remaining blocks. The color component of a picture may include a luma component and a chroma component, and unless explicitly limited in the present disclosure, embodiments according to the present disclosure may be applied to the luma component and the chroma component. Meanwhile, when intra prediction is performed, the prediction unit (220, 330) of the encoding device (200) / decoding device (300) can derive a reference sample according to the intra prediction mode of the current block among the surrounding samples of the current block, and can generate a prediction sample of the current block based on the reference sample. For example, (i) the prediction sample can be derived based on the average or interpolation of neighboring reference samples of the current block, and (ii) the prediction sample can be derived based on a reference sample existing in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. The case of (i) can be called a non-directional mode or a non-angular mode, and the case of (ii) can be called a directional mode or an angular mode. Additionally, linear interpolation intra prediction (LIP) may be applied to perform intra prediction on the current block by linearly interpolating prediction sample values generated based on the intra prediction mode of the current block. Additionally, a temporary prediction sample of the current block may be derived based on filtered peripheral reference samples, and a prediction sample of the current block may be derived by weighting at least one reference sample derived according to an intra prediction mode among existing peripheral reference samples, i.e., unfiltered peripheral reference samples, and the temporary prediction sample. Such prediction may be referred to as Position Dependent Intra Prediction Combination (PDPC). In addition, intra prediction encoding can be performed by selecting a reference sample line with the highest prediction accuracy among the surrounding multiple reference sample lines of the current block, deriving a prediction sample using the reference sample located in the prediction direction of the selected line, and then instructing (signaling) the used reference sample line to the decoding device. This case can be referred to as multi-reference line intra prediction (MRL) or MRL-based intra prediction. Additionally, the current block can be divided into vertical or horizontal subpartitions, and intra prediction can be performed based on the same intra prediction mode, while peripheral reference samples can be derived and utilized for each subpartition. In other words, in this case, the intra prediction mode for the current block is applied equally to the subpartitions, but peripheral reference samples can be derived and utilized for each subpartition, thereby improving intra prediction performance in some cases. This prediction method can be called intra subpartitions (ISP) or ISP-based intra prediction. Additionally, if the prediction direction based on the prediction sample points between surrounding reference samples, i.e., if the prediction direction points to a fractional sample location, the value of the prediction sample can be derived through interpolation of multiple reference samples located around the prediction direction (around the fractional sample location). The MPM list for deriving the intra prediction mode described above may be configured differently depending on the intra prediction type. Alternatively, the MPM list may be configured in common regardless of the intra prediction type. Spatial Geometric Partitioning Mode (SGPM) SGPM is an intra mode similar to GPM's inter-coding tool, where two prediction parts are generated through the intra prediction process. In this mode, a candidate list is created for each entry, containing one partition and two intra prediction modes. The partition mode and three intra prediction modes are used to form a combination. The candidate list length can be set to 16, and the selected candidate index can be signaled. The candidate list is reordered using templates, where the SAD between the template's prediction and reconstruction is used for alignment. The template size can be fixed to 1. For each partition mode, an IPM list for each part is derived using the same intra-inter GPM list derivation. The size of the IPM list can be set to 3. In the list, the TIMD derivation mode can be replaced by two derivation modes in the horizontal and vertical directions. SGPM mode can be applied with limited block sizes as follows: 4<=width<=64, 4<=height<=64, width <height*8, height<width*8, width*height> =32 Adaptive blending is also used in spatial GPM, and the blending depth τ can be derived as follows: If min(width, height)==4, 1 / 2 τ is selected else if min(width, height)==8, τ is selected else if min(width, height)==16, 2 τ is selected else if min(width, height)==32, 4 τ is selected else, 8 τ is selected Intra Block Copy (IBC) IBC is a method that can significantly improve the coding efficiency of screen content materials. Since IBC mode is implemented as a block-level coding mode, block matching (BM) can be performed in the encoder to find an optimal block vector (or motion vector) for each CU. Here, the block vector can be used to indicate the displacement from the current block to a reference block already reconstructed within the current block. The luma block vector of an IBD-coded CU can be of integer resolution (or precision). The chroma block vector can be rounded to integer resolution. When combined with AMVR, IBC mode can switch between 1-pel (pixel) and 4-pel (pixel) motion vector resolutions. An IBD-coded CU can be treated as a third prediction mode, rather than an intra- or inter-prediction mode. IBC mode can be applied to CUs with both a width and a height of 64 luma samples or less. On the encoder side, hash-based motion estimation can be performed for IBC. The encoder can perform a block-by-block (BD) check for blocks whose width and height are no greater than 16 luma samples. For non-merge mode, block vector search can be performed first using a hash-based search. If the hash search does not return a valid candidate, a block-matching-based local search can be performed. In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block can be extended to all allowed block sizes. The hash key calculation for each location in the current picture is based on 4X4 sub-blocks. For larger current blocks, if all the hash keys of the 4X4 sub-blocks match the hash keys of the corresponding reference locations, the hash key can be determined to match that of the reference block. When the hash keys of multiple prediction blocks are found to match those of the current block, the block vector cost of each matched reference is calculated and the minimum cost can be selected. In block matching searches, the search range can be set to cover both previous and current CTUs. At the CU level, the IBC mode is signaled by a flag, which can be signaled as IBC AMVP mode or IBC Skip / Merge mode as follows. - IBC Skip / Merge Mode: A merge candidate index can be used to indicate which block vectors from a list of neighboring candidate IBC coded blocks are used to predict the current block. The merge list can include spatial, HMVP, and pairwise candidates. - IBC AMVP mode: Block vector differences can be coded in the same way as motion vector differences. The block vector prediction method can use two candidates as predictors: one from the left neighbor and one from the upper neighbor (if IBC coded). If neither neighbor is available, the default block can be used as the predictor. A flag can be signaled to indicate the block vector predictor index. Intra TMP (Intra Template Matching Prediction) Intra-template matching prediction (IntraTMP) is a special intra-prediction mode that copies the optimal prediction block where the L-shaped template matches the current template within the reconstructed portion of the current frame. Within a predefined search range, the encoder searches the reconstructed region of the current frame for the template most similar to the current template and uses that block as the prediction block. The encoder then signals the use of this mode, and the decoder performs the same prediction operation. Figure 6 is a diagram showing an example of a search area used in intra template matching. The prediction signal is generated by matching the L-shaped, top-only, or left-only causal neighbors of the current block with other blocks within the predefined search regions of Fig. 6. As illustrated in Fig. 6, there can be a total of six predefined search regions (i.e., R1 to R6), which include not only some of the reconstructed samples within the current CTU located above, left, bottom-left, and top-right of the current block, but also samples reconstructed from the upper CTU and the left CTU: The sum of absolute differences (SAD) is used as a cost function. A given search order is utilized across six search regions (i.e., R4, R5, R6, R1, R2, R3). Within each region, the decoder generates a list of up to 19 template-matching block vector candidates, sorted in ascending order by template cost (SAD). The supported modes are: 1. Single Predictor: A single predictor is selected from the candidate list. 2. Fusion of Multiple Predictors: Multiple predictors are fused to derive the final prediction block. The fusion weights can be calculated based on the template matching cost of each predictor, or a weight derivation method based on a Wiener filter can be used. 3. Sub-pixel Precision: When using a single predictor, it supports 1 / 2 pixel, 1 / 4 pixel, and 3 / 4 pixel precision, and provides 8 directions each. 4. Linear Filter Model: Applies a linear filter learned between the reference template and the current template to the reference block. This mode can be applied to a single predictor that does not use subpixel precision. To ensure a fixed number of SAD comparisons per pixel, the sizes of all search regions (SearchRange_w, SearchRange_h) are set to be proportional to the block sizes (BlkW, BlkH). That is: SearchRange_w = min(64,a * BlkW) SearchRange_h = min(64,a * BlkH) Here, 'a' is a constant that controls the trade-off between gain and complexity, and can be set to a = 5. To accelerate the template matching process, the search range of each search area can be subsampled by a factor of three. After finding the optimal match, a refinement process is performed. This refinement is achieved through a second template matching search around the optimal match in the reduced range. Intra-template matching can be enabled in CUs with a width and height of 64 or less. The maximum CU size for intra-template matching is configurable. Intra template matching prediction mode can be signaled at the CU level via a dedicated flag when DIMD is not used for the current CU. Meanwhile, when inter prediction is applied, the prediction unit of the encoding device / decoding device can perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction can refer to a prediction derived in a manner dependent on data elements (e.g., sample values, or motion information) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and / or a reference picture index. The motion information may further include information on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the above reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, a motion information candidate list may be constructed based on the neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference. The above motion information may include L0 motion information and / or L1 motion information depending on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be called an L0 prediction, prediction based on an L1 motion vector may be called an L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be called a bi-prediction (Bi). Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures preceding the current picture in output order as reference pictures, and the reference picture list L1 may include pictures succeeding the current picture in output order. The preceding pictures may be called forward (reference) pictures, and the succeeding pictures may be called backward (reference) pictures. The above reference picture list L0 may further include pictures subsequent to the current picture in output order as reference pictures. In this case, the previous pictures may be indexed first and the subsequent pictures may be indexed next within the reference picture list L0. The above reference picture list L1 may further include pictures previous to the current picture in output order as reference pictures. In this case, the subsequent pictures may be indexed first and the previous pictures may be indexed next within the reference picture list L1. Here, the output order may correspond to a POC (picture order count) order. FIG. 7 and FIG. 8 illustrate examples of inter prediction-based video / image encoding methods to which embodiments of the present disclosure can be applied. Referring to FIG. 7, the encoding device (200) can perform inter prediction for the current block (S600). The encoding device can derive the inter prediction mode and motion information of the current block, and generate prediction samples of the current block. Here, the inter prediction mode determination, motion information derivation, and prediction sample generation procedures may be performed simultaneously, or one procedure may be performed before the other. For example, as illustrated in FIG. 8, the inter prediction unit (221) of the encoding device (200) may include a prediction mode determination unit (221a), a motion information derivation unit (221b), and a prediction sample derivation unit (221c), and the prediction mode determination unit (221a) may determine the prediction mode for the current block, the motion information derivation unit (221b) may derive motion information of the current block, and the prediction sample derivation unit (221c) may derive prediction samples of the current block. For example, the inter prediction unit of the encoding device can search for a block similar to the current block within a certain area (search area) of reference pictures through motion estimation, and derive a reference block whose difference from the current block is minimal or below a certain standard. Based on this, a reference picture index indicating a reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine a mode to be applied to the current block among various prediction modes. The encoding device can compare RD costs for the various prediction modes and determine an optimal prediction mode for the current block. For example, when skip mode or merge mode is applied to the current block, the encoding device may configure a merge candidate list described below, and derive a reference block among the reference blocks indicated by the merge candidates included in the merge candidate list, wherein the difference between samples from the current block, i.e., the difference in SAD or SATD, is minimal or below a certain standard. In this case, a merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding device. Motion information of the current block may be derived using motion information of the selected merge candidate. As another example, when the (A)MVP mode is applied to the current block, the encoding device may configure an (A)MVP candidate list described below, and use the motion vector of a motion vector predictor (mvp) candidate selected from among motion vector predictor (mvp) candidates included in the (A)MVP candidate list as the motion vector predictor of the current block. In this case, for example, a motion vector pointing to a reference block derived by the above-described motion estimation may be used as the motion vector of the current block, and a motion vector predictor candidate having a motion vector with the smallest difference from the motion vector of the current block among the motion vector predictor candidates may be the selected motion vector predictor candidate. A Motion Vector Difference (MVD), which is the difference obtained by subtracting the motion vector predictor from the motion vector of the current block, may be derived. In this case, information about the MVD may be signaled to the decoding device. In addition, when the (A)MVP mode is applied, the value of the reference picture index may be configured with reference picture index information and signaled separately to the decoding device. The encoding device can derive residual samples based on predicted samples (S610). The encoding device can derive residual samples by comparing the original samples of the current block with the predicted samples. An encoding device can encode image information including prediction information and residual information (S620). The encoding device can output the encoded image information in the form of a bitstream. The prediction information is information about prediction and may include prediction mode information (e.g., skip flag, merge flag, or merge index, etc.) and / or motion information. The motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index), which is information for deriving a motion vector. In addition, the motion information may include information about the above-described MVD and / or reference picture index information. In addition, the motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about residual samples. The residual information may include information about quantized transform coefficients for the residual samples. The output bitstream can be stored on a (digital) storage medium and transmitted to a decoding device, or can be transmitted to a decoding device via a network. Meanwhile, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to derive the same prediction result as that performed by the decoding device from the encoding device, thereby improving coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the reconstructed picture. FIGS. 9 and 10 illustrate examples of inter-prediction based video / image decoding methods to which embodiments of the present disclosure can be applied. A video / image decoding procedure based on inter prediction may roughly include, for example: Referring to FIGS. 9 and 10, the decoding device (300) can perform an operation corresponding to the operation performed in the encoding device (200). The decoding device can perform a prediction on the current block based on the received prediction information and derive prediction samples. Specifically, the decoding device can determine a prediction mode for the current block based on the received prediction information (S700). The prediction mode determination unit (332a) of the decoding device (300) can determine which inter prediction mode is applied to the current block based on the prediction mode information in the prediction information. For example, based on the merge flag, it can be determined whether the current block is subject to merge mode or (A)MVP mode. Alternatively, one of various inter prediction mode candidates can be selected based on the mode index. The inter prediction mode candidates can include skip mode, merge mode, and / or (A)MVP mode, or can include various inter prediction modes described below. The decoding device can derive motion information of the current block based on the determined inter prediction mode (S710). For example, when skip mode or merge mode is applied to the current block, the motion information derivation unit (332b) of the decoding device (300) can construct a merge candidate list described below and select one merge candidate from among the merge candidates included in the merge candidate list. This selection can be performed based on the above-described selection information (merge index). Motion information of the current block can be derived using motion information of the selected merge candidate. Motion information of the selected merge candidate can be used as motion information of the current block. As another example, when the (A)MVP mode is applied to the current block, the decoding device may construct an (A)MVP candidate list described below, and use the motion vector of an MVP candidate selected from among the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. This selection may be performed based on the selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block may be derived based on information about the MVD, and the motion vector of the current block may be derived based on the MVP and MVD of the current block. In addition, the reference picture index of the current block may be derived based on reference picture index information. A picture indicated by a reference picture index within the reference picture list for the current block may be derived as a reference picture referenced for inter prediction of the current block. Meanwhile, the motion information of the current block may be derived without constructing a candidate list, in which case the motion information of the current block may be derived according to the procedure initiated in the prediction mode. In this case, the candidate list construction described above may be omitted. The decoding device can generate prediction samples for the current block based on the motion information of the current block (S720). In this case, the prediction sample derivation unit (332c) of the decoding device (300) can derive a reference picture based on the reference picture index of the current block, and derive prediction samples of the current block using samples of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as described below, a prediction sample filtering procedure may be further performed on all or part of the prediction samples of the current block, depending on the case. In other words, the inter prediction unit (332) of the decoding device (300) may include a prediction mode determination unit (332a), a motion information derivation unit (332b), and a prediction sample derivation unit (332c), and may determine a prediction mode for the current block based on the prediction mode information received from the prediction mode determination unit (332a), derive motion information (motion vector and / or reference picture index, etc.) of the current block based on the information about motion information received from the motion information derivation unit (332b), and derive or generate prediction samples of the current block from the prediction sample derivation unit (332c). The decoding device generates residual samples for the current block based on the received residual information (S730). The decoding device (300) generates restoration samples for the current block based on the prediction samples and residual samples, and can generate a restoration picture based on these (S740). As described above, in-loop filtering procedures, etc. may be further applied to the restoration picture. FIG. 11 exemplarily illustrates an inter prediction procedure to which an embodiment of the present disclosure can be applied. Referring to FIG. 11, the inter prediction procedure (S600) as described above may include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction performance (prediction sample generation) step based on the derived motion information. The inter prediction procedure may be performed in an encoding device and a decoding device as described above. In this document, a coding device may include an encoding device and / or a decoding device. Referring to FIG. 11, the coding device determines an inter prediction mode for a current block (S800). Various inter prediction modes can be used to predict the current block within a picture. For example, various modes such as merge mode, skip mode, MVP (Motion Vector Prediction) mode, affine mode, sub-block merge mode, and MMVD (Merge with MVD) mode can be used. Decoder side Motion Vector Refinement (DMVR) mode, Adaptive Motion Vector Resolution (AMVR) mode, bi-prediction with CU-level Weight (BCW), bi-directional optical flow (BDOF), etc. can be used as auxiliary modes. In addition, according to an embodiment, the above-described inter prediction mode can include a multi-hypethesis prediction (MHP) mode. The multi-hypethesis prediction mode represents a method of performing prediction by weighting an additional prediction block generated based on additional motion information to an inter prediction block. The multi-hypethesis prediction mode will be described in detail below. In the present disclosure, the affine mode may be referred to as the affine motion prediction mode. In addition, the MVP mode may be referred to as the Advanced Motion Vector Prediction (AMVP) mode. In the present disclosure, motion information candidates derived from some modes and / or some modes may be included as one of the motion information-related candidates of other modes. For example, an HMVP candidate may be added as a merge candidate of the merge / skip mode, or may be added as a motion vector predictor candidate of the AMVP mode. When an HMVP candidate is used as a motion information candidate of the merge mode or the skip mode, the HMVP candidate may be referred to as an HMVP merge candidate. Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information can include index information indicating one of multiple candidate modes. Alternatively, the inter-prediction mode can be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether skip mode is applied, a merge flag may be signaled to indicate whether merge mode is applied when skip mode is not applied, and MVP mode may be indicated to be applied when merge mode is not applied, or additional flags may be signaled for additional distinction. Affine mode may be signaled as an independent mode, or as a mode dependent on merge mode or MVP mode. For example, an affine mode may include an affine merge mode and an affine MVP mode. The coding device can derive motion information for the current block (S810). The motion information can be derived based on the inter-prediction mode determined in the aforementioned step. The coding device can perform inter-prediction using the motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, an encoding device can use an original block within an original picture for a current block to search for a similar reference block with a high correlation within a predetermined search range within the reference picture in fractional pixel units, thereby deriving motion information. The similarity of blocks can be derived based on the difference in phase-based sample values. For example, the similarity of blocks can be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search range. The derived motion information can be signaled to a decoding device in various ways based on an inter prediction mode. The coding device can perform inter prediction based on motion information for the current block to generate prediction samples (S820). The current block containing the prediction samples may be referred to as a prediction block. Meanwhile, information indicating whether the above-described List0 (L0) prediction, List1 (L1) prediction, or bi-prediction is used for the current block (current coding unit) can be signaled. This information may be called motion prediction direction information, inter-prediction direction information, or inter-prediction indication information, and may be configured / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether the above-described List0 (L0) prediction, List1 (L1) prediction, or bi-prediction is used for the current block (current coding unit). In this document, for the convenience of explanation, the inter-prediction type (L0 prediction, L1 prediction, or BI prediction) indicated by the inter_pred_idc syntax element may be represented as motion prediction direction. L0 prediction may be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, depending on the value of the inter_pred_idc syntax element, the prediction type can be indicated as in Table 1 below. [Table 1] As described above, a picture may include one or more slices. A slice may have one of the following types: intra (I) slice, predictive (P) slice, and bi-predictive (B) slice. The slice type may be indicated based on slice type information. For blocks within an I slice, inter prediction is not used for prediction, and only intra prediction can be used. Of course, even in this case, the original sample values can be coded and signaled without prediction. For blocks within a P slice, either intra prediction or inter prediction can be used, and when inter prediction is used, only uni prediction can be used. On the other hand, for blocks within a B slice, either intra prediction or inter prediction can be used, and when inter prediction is used, up to bi prediction can be used. L0 and L1 may include reference pictures encoded / decoded before the current picture. For example, L0 may include reference pictures that are before and / or after the current picture in POC order, and L1 may include reference pictures that are after and / or before the current picture in POC order. In this case, L0 may be assigned a relatively lower reference picture index to reference pictures that are before the current picture in POC order, and L1 may be assigned a relatively lower reference picture index to reference pictures that are after the current picture in POC order. For B slices, bi-prediction may be applied, and in this case, either uni-directional bi-prediction or bi-directional bi-prediction may be applied. Bi-directional bi-prediction may be called true bi-prediction. Inter prediction can be performed using motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can search for similar reference blocks with high correlation within a predetermined search range within the reference picture using the original block within the original picture for the current block, in fractional pixel units, and thereby derive motion information. The similarity between blocks can be derived based on the difference in phase-based sample values. For example, the similarity between blocks can be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search range. The derived motion information can be signaled to the decoding device in various ways based on the inter prediction mode. Merge mode and skip mode When merge mode is applied, the motion information of the current prediction block is not directly transmitted, but the motion information of the surrounding prediction blocks is used to derive the motion information of the current prediction block. Therefore, the motion information of the current prediction block can be indicated by transmitting flag information indicating that merge mode is used and a merge index indicating which surrounding prediction block was used. The merge mode may also be called regular merge mode. To perform merge mode, the encoder can search for merge candidate blocks to derive motion information of the current prediction block. For example, up to five merge candidate blocks can be used, but the present invention is not limited thereto. In addition, the maximum number of merge candidate blocks can be transmitted in the slice header or tile group header, but the present invention is not limited thereto. After finding the merge candidate blocks, the encoder can generate a merge candidate list, and select the merge candidate block with the lowest cost among them as the final merge candidate block. Figure 12 is a diagram showing examples of blocks used to construct a merge candidate list. The present invention provides various embodiments for merge candidate blocks constituting a merge candidate list. The merge candidate list may utilize, for example, five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate may be utilized. As a specific example, in the case of spatial merge candidates, the blocks illustrated in Fig. 12 may be utilized as spatial merge candidates. Hereinafter, the spatial merge candidates or the spatial MVP candidates described below may be referred to as SMVPs, and the temporal merge candidates or the temporal MVP candidates described below may be referred to as TMVPs. The list of merge candidates for the current block can be constructed based on, for example, the following procedure: - Insert spatial merge candidates derived by exploring spatial surrounding blocks into the merge candidate list. - Insert the temporal merge candidates derived by exploring the temporal surrounding blocks into the merge candidate list. - Compare the number of current merge candidates with the maximum number of merge candidates. - If the number of current merge candidates is less than the maximum number of merge candidates, insert additional merge candidates into the merge candidate list. Specifically, the above-described procedure is explained in detail, the coding device (encoding device / decoding device) searches the spatial neighboring blocks of the current block and inserts the derived spatial merge candidates into the merge candidate list. For example, the spatial neighboring blocks may include the lower left corner neighboring blocks, the left neighboring blocks, the upper right corner neighboring blocks, the upper neighboring blocks, and the upper left corner neighboring blocks of the current block. However, this is merely an example, and in addition to the above-described spatial neighboring blocks, additional neighboring blocks such as the right neighboring blocks, the lower neighboring blocks, and the lower right neighboring blocks may be further used as the spatial neighboring blocks. The coding device may search the spatial neighboring blocks based on priorities to detect available blocks and derive motion information of the detected blocks as spatial merge candidates. For example, the encoder and decoder may search the five blocks illustrated in FIG. 12 in the order of A1, B1, B0, A0, and B2, and sequentially index the available candidates to form a merge candidate list. The coding device searches the temporal neighboring blocks of the current block and inserts the derived temporal merge candidates into the merge candidate list. The temporal neighboring blocks may be located on a reference picture that is different from the current picture where the current block is located. The reference picture where the temporal neighboring blocks are located may be called a collocated picture or col picture. The temporal neighboring blocks may be searched in the order of the lower right corner neighboring blocks and the lower right center block of the co-located block of the current block in the col picture. Meanwhile, when motion data compression is applied, specific motion information can be stored as representative motion information for each storage unit in a col picture. In this case, there is no need to store motion information for all blocks within a certain storage unit, and thus, a motion data compression effect can be obtained. In this case, a certain storage unit may be predetermined, for example, a 16x16 sample unit or an 8x8 sample unit, or size information for a certain storage unit may be signaled from the encoder to the decoder. When motion data compression is applied, the motion information of a temporal neighboring block may be replaced with representative motion information of a certain storage unit where the temporal neighboring block is located. That is, in terms of implementation, a temporal merge candidate may be derived based on the motion information of a prediction block that covers a position that is arithmetically shifted to the right by a certain value based on the coordinates of the temporal neighboring block (upper left sample position), rather than a prediction block located at the coordinates of the temporal neighboring block. For example, if the certain storage unit is 2 n x2 nIn the case of sample units, if the coordinates of the temporal surrounding block are (xTnb, yTnb), then the modified position is ((xTnb>>n)<<n), (yTnb> >n)< <n))에 위치하는 예측 블록의 움직임 정보가 시간적 머지 후보를 위하여 사용될 수 있다. 구체적으로 예를 들어, 일정 저장 단위가 16x16 샘플 단위인 경우, 시간적 주변 블록의 좌표가 (xTnb, yTnb)라 하면, 수정된 위치인 ((xTnb> >4)<<4), (yTnb>>4)<<4)) motion information of the prediction block located at can be used for the temporal merge candidate. Or, for example, if the storage unit is an 8x8 sample unit and the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the modified position ((xTnb>>3)<<3), (yTnb>>3)<<3)) can be used for the temporal merge candidate. The coding device can compare the current number of merge candidates with the maximum number of merge candidates. The maximum number of merge candidates can be predefined or signaled from the encoder to the decoder. For example, the encoder can generate information regarding the maximum number of merge candidates, encode it, and transmit it to the decoder in bitstream form. Once the maximum number of merge candidates is reached, the subsequent candidate addition process may not proceed. If the comparison result shows that the number of current merge candidates is less than the maximum number of merge candidates, the coding device inserts an additional merge candidate into the merge candidate list. The additional merge candidate may include, for example, at least one of the following: history-based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate (if the slice / tile group type of the current slice / tile group is type B), and / or zero-vector merge candidate. If the comparison result shows that the number of current merge candidates is not less than the maximum number of merge candidates, the coding device can terminate the construction of the merge candidate list. In this case, the encoder can select the optimal merge candidate among the merge candidates constituting the merge candidate list based on the RD (rate-distortion) cost and signal the selection information (e.g., merge index) indicating the selected merge candidate to the decoder. The decoder can select the optimal merge candidate based on the merge candidate list and the selection information. As described above, the motion information of the selected merge candidate can be used as the motion information of the current block, and prediction samples of the current block can be derived based on the motion information of the current block. The encoder can derive residual samples of the current block based on the prediction samples, and can encode residual information about the residual samples and transmit them to the decoder. As described above, the decoder can generate reconstructed samples based on the residual samples and prediction samples derived based on the transmitted residual information, and generate a reconstructed picture based on the reconstructed samples. When skip mode is applied, motion information for the current block can be derived in the same way as when merge mode is applied. However, when skip mode is applied, the residual signal for the corresponding block is omitted, so the predicted samples can be used directly as restored samples. History-based merge candidate derivation A history-based MVP (HMVP) merge candidate can be added to the merge list after spatial and temporal merge candidates. In this method, motion information of previously encoded blocks is stored in a table and used as the MVP of the current coding unit (CU). A table containing multiple HMVP candidates is maintained during encoding and decoding processes. The table is initialized (emptied) when a new coding tree unit (CTU) row begins. Whenever a non-subblock inter-encoded CU exists, the associated motion information is added to the last entry of the table as a new HMVP candidate. In VVC, the size S of the HMVP table is set to 5, indicating that up to five history-based MVP (HMVP) candidates can be added to the table. When inserting a new move candidate into the table, a limited first-in, first-out (FIFO) rule is applied, and a duplicate check is performed first to determine if there are identical HMVPs in the table. If an identical HMVP is found, it is removed from the table, and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the merge candidate list construction process. The most recent HMVP candidates in the table are sequentially examined and inserted into the merge candidate list after the TMVP candidates. HMVP candidates are also checked for overlap with spatial and temporal merge candidates. To reduce the number of redundant check operations, the following simplifications are introduced: The number of HMVP candidates used to generate the merge list is set to M if (N ≤ 4), otherwise to (8 - N), where N represents the number of existing candidates in the merge list and M represents the number of available HMVP candidates in the table. When the total number of possible merge candidates reaches the maximum number of allowed merge candidates minus 1, the process of constructing the merge candidate list using HMVP ends. Pair-wise average merge candidates derivation In this document, the pair-wise average merge candidate may be referred to as the pair-wise average candidate or pair-wise candidate. The pair-wise average candidate is generated by averaging predefined pair-wise candidates within the existing merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}. Here, the numbers represent merge indices within the merge candidate list. The average motion vector is calculated individually for each reference list. If both motion vectors are available in a single reference list, the average is calculated even if they point to different reference pictures. If only one motion vector is available, that motion vector is used directly. If no motion vectors are available, the list remains invalid. If the merge list is not filled after adding pairwise average merge candidates, a zero MVP is inserted at the end of the list until the maximum number of merge candidates is reached. Merge mode with MVD (MMVD) This section describes the MMVD mode mentioned above. In addition to Merge mode, which directly uses implicitly derived motion information for the prediction samples of the current block, Merge mode (MMVD) can also be used. Since similar motion information derivation methods are used in Skip mode and Merge mode, MMVD can also be applied to Skip mode. After signaling the Skip flag and Merge flag, MMVD flag information (e.g., mmvd_flag) can be signaled to indicate whether MMVD mode is to be used for the current block. In MMVD mode, after a merge candidate is selected, the candidate can be further refined based on the signaled MVD information. When MMVD is applied to the current block (i.e., when mmvd_flag is 1), additional information about MMVD can be signaled. The additional information can include a merge candidate flag (e.g., mmvd_merge_flag) indicating whether the first or second candidate in the merge candidate list is used with motion vector difference, a distance index indicating the motion magnitude (e.g., mmvd_distance_idx), and a direction index indicating the motion direction (e.g., mmvd_direction_idx). In MMVD mode, one of the first two candidates in the merge list can be selected to be used as the MV reference. The merge candidate flag is signaled to indicate which candidate is used. The distance index specifies motion size information and represents a predefined offset from the starting point. The offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is shown in Table 2 below. [Table 2] Here, if slice_fpel_mmvd_enabled_flag is 1, it indicates that the merge mode with motion vector difference uses integer sample precision in the current slice. If slice_fpel_mmvd_enabled_flag is 0, it indicates that the merge mode with motion vector difference can use fractional sample precision in the current slice. The slice_fpel_mmvd_enabled_flag syntax element can be signaled via the slice header or included in the slice header. The direction index indicates the direction of the MVD with respect to the starting point. The direction index can indicate one of four directions, as shown in Table 3 below. The meaning of the MVD symbol may vary depending on the information of the starting MV. If the starting MV is a non-predicted MV or both lists are bidirectional MVs pointing to the same side of the current picture (i.e., the POCs of both references are both greater than or less than the POC of the current picture), the sign in Table 3 below can indicate the sign of the MV offset added to the starting MV. If the starting MVs are bidirectional predictive MVs and the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), the sign in Table 3 below indicates the sign of the MV offset added to the MV element of List 0 of the starting MV, and the sign of List 1 MV has the opposite value. [Table 3] The two elements of the merge plus MVD offset MmvdOffset[x0][y0] can be derived as follows: - MmvdOffset[ x0 ][ y0 ]

[0000] = ( MmvdDistance[ x0 ][ y0 ] << 2 ) * MmvdSign[ x0 ][ y0 ][0] - MmvdOffset[ x0 ][ y0 ]

[0001] = ( MmvdDistance[ x0 ][ y0 ] << 2 ) * MmvdSign[ x0 ][ y0 ][1] MVP (Motion Vector Prediction) The MVP (Motion Vector Prediction) mode may also be referred to as the AMVP (advanced motion vector prediction) mode. When the MVP mode is applied, a motion vector predictor (mvp) candidate list can be generated using the motion vectors of the reconstructed spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks (or Col blocks). That is, the motion vectors of the reconstructed spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks can be used as motion vector predictor candidates. When paired prediction is applied, an mvp candidate list for deriving L0 motion information and an mvp candidate list for deriving L1 motion information can be generated and used separately. The above-described prediction information (or information regarding prediction) may include selection information (e.g., MVP flag or MVP index) indicating an optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the list. In this case, the prediction unit of the decoding device may use the selection information to select a motion vector predictor of the current block from among the motion vector predictor candidates included in the motion vector candidate list. The prediction unit of the encoding device can obtain a motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can encode and output it in the form of a bitstream. That is, the MVD can be obtained as a value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit of the decoding device can obtain the motion vector difference included in the information regarding the prediction, and derive the motion vector of the current block through the addition of the motion vector difference and the motion vector predictor. The prediction unit of the decoding device can obtain or derive a reference picture index indicating a reference picture, etc. from the information regarding the prediction. For example, a motion vector predictor candidate list can be configured as follows: - Search for spatial candidate blocks for motion vector prediction and insert them into the prediction candidate list. - Check if the number of spatial candidate blocks is less than 2 - If the number of spatial candidate blocks is less than 2, search for temporal candidate blocks and insert them into the prediction candidate list. - If no temporal candidate block is available, use zero motion vector. - If the number of spatial candidate blocks is not less than 2, the construction of the motion vector predictor candidate list is terminated. Meanwhile, when the MVP mode is applied, the reference picture index can be explicitly signaled. In this case, the reference picture index for L0 prediction (refidxL0) and the reference picture index for L1 prediction (refidxL1) can be signaled separately. For example, when the MVP mode is applied and bi-prediction (BI prediction) is applied, both information about refidxL0 and information about refidxL1 can be signaled. Motion Vector Difference (MVD) coding When the MVP mode is applied, information about the MVD derived from the encoding device as described above can be signaled or encoded and transmitted to the decoding device. The information about the MVD can include, for example, information indicating x and y components for the MVD absolute value and sign. In this case, information indicating whether the MVD absolute value is greater than 0 and greater than 1, and the MVD remainder can be signaled in stages. For example, information indicating whether the MVD absolute value is greater than 1 can be signaled only when the value of the flag information indicating whether the MVD absolute value is greater than 0 is 1. For example, information about an MVD can be encoded in an encoding device and signaled to a decoding device using the following syntax: [Table 4] For example, MVD[compIdx] can be derived based on abs_mvd_greater0_flag[compIdx] 2 * mvd_sign_flag[compIdx]). Here, compIdx (or cpIdx) represents the index of each component and can have the value 0 or 1. A compIdx value of 0 can represent the x component, and a compIdx value of 1 can represent the y component. However, this is just an example, and values for each component can be represented using a coordinate system other than the x, y coordinate system. Meanwhile, MVD for L0 prediction (MVDL0) and MVD for L1 prediction (MVDL1) may be signaled separately, and information about MVD may include information about MVDL0 and / or information about MVDL1. For example, if MVP mode and BI prediction are applied to the current block, information about MVDLO and information about MVDL1 may both be signaled. Symmetric MVD Meanwhile, when BI prediction is applied, symmetric MVD may be used considering coding efficiency. In this case, signaling of some of the motion information may be omitted. For example, when symmetric MVD is applied to the current block, information about refidxL0, information about refidxL1, and information about MVDL1 may not be signaled from the encoding device to the decoding device, but may be derived internally. For example, when MVP mode and BI prediction are applied to the current block, flag information indicating whether symmetric MVD is applied (e.g., symmetric MVD flag information or sym_mvd_flag syntax element) may be signaled, and when the value of the flag information is 1, the decoding device may determine that symmetric MVD is applied to the current block. When the symmetric MVD mode is applied (i.e., when the value of the symmetric MVD flag information is 1), information about mvp_l0_flag, mvp_l1_flag, and MVDL0 may be explicitly signaled, and signaling of information about refidxL0, information about refidxL1, and information about MVDL1 may be omitted and derived internally as described above. For example, refidxL0 may be derived as an index pointing to a previous reference picture that is closest to the current picture in POC order within reference picture list 0 (which may be referred to as list 0 or L0). refidxL1 may be derived as an index pointing to a subsequent reference picture that is closest to the current picture in POC order within reference picture list 1 (which may be referred to as list 1 or L1). Or, for example, both refidxL0 and refidxL1 may be derived as 0. Or, for example, the refidxL0 and refidxL1 may be derived as the minimum indexes having the same POC difference in relation to the current picture. As a specific example, when [POC of the current picture] - [POC of the first reference picture indicated by refidxL0] is referred to as the first POC difference, and [POC of the second reference picture indicated by refidxL1] is referred to as the second POC difference, only when the first POC difference and the second POC difference are the same, the value of refidxL0 pointing to the first reference picture may be derived as the value of refidxL0 of the current block, and the value of refidxL1 pointing to the second reference picture may be derived as the value of refidxL1 of the current block.Also, for example, if there are multiple sets in which the first POC difference and the second POC difference are the same, refidxL0 and refidxL1 of the set with the minimum difference can be derived as refidxL0 and refidxL1 of the current block. MVDL1 can be derived from -MVDL0. For example, the final MV for the current block can be derived as shown in Equation 1 below. [Formula 1] Affine Prediction Existing video coding systems use only a single motion vector (using a translation motion model) to represent the motion of a coded block. While this method can optimally represent motion at the block level, it does not accurately represent the optimal motion for each pixel. Therefore, determining the optimal motion vector at the pixel level can improve encoding efficiency. To achieve this, we describe an affine motion prediction method that uses an affine motion model to encode the data. Figure 13 is a diagram showing four movements that can be expressed in the affine movement model. Affine motion prediction methods can express motion vectors at each pixel unit of a block using two, three, or four motion vectors. The affine motion model can express four types of motion, as illustrated in Fig. 13. The affine motion model that expresses three types of motion (translation, scale, and rotation) among the motions that the affine motion model can express is called a similarity (or simplified) affine motion model, and the proposed methods are described below based on the similarity affine motion model. However, the disclosed embodiments are not limited to the motion model. Figure 14 is a diagram showing an example of a control point motion vector used in affine motion prediction. As illustrated in Figure 14, affine motion prediction can determine the motion vector of the pixel position included in the block using two or more control point motion vectors (CPMV). At this time, the set of motion vectors is called an affine motion vector field (MVF) and can be determined by the equations below. For the 4-parameter affine motion model, the motion vector at the sample location (x,y) of the block can be derived by Equation 2 below. [Formula 2] For the 6-parameter affine motion model, the motion vector at the sample location (x,y) of the block can be derived by Equation 3 below. [Formula 3] Here is the CPMV of the CP at the top-left corner of the coding block. is the CPMV of the CP at the top-right corner location, is the CPMV of the CP at the bottom-left corner location. And W corresponds to the width of the current block, H corresponds to the height of the current block, is the motion vector at position {x, y}. In the encoding / decoding process, the affine MVF can be determined in units of pixels or pre-defined subblocks. When determined in units of pixels, a motion vector is obtained based on each pixel value, and when determined in units of subblocks, the motion vector of the corresponding block is obtained based on the pixel value at the center of the subblock (the lower right side of the center, i.e., the lower right sample among the four central samples). For example, as illustrated in Fig. 14, it is possible for the affine MVF to be determined in units of 4*4 subblocks. However, this is merely an example, and the size of the subblock applied to affine prediction can be varied. When affine prediction is available, the motion models applicable to the current block can include three models: a translational motion model, a 4-parameter affine motion model, and a 6-parameter affine motion model. Here, the translational motion model can represent a model in which existing block-level motion vectors are used, the 4-parameter affine motion model can represent a model in which two CPMVs are used, and the 6-parameter affine motion model can represent a model in which three CPMVs are used. Affine motion prediction may include affine MVP (or affine inter) mode and affine merge. In affine motion prediction, the motion vectors of the current block can be derived on a sample-by-sample or sub-block-by-subblock basis. Affine merge In affine merge mode, CPMV can be determined based on the affine motion model of neighboring blocks coded using affine motion prediction. Neighboring blocks affine-coded in search order can be used in affine merge mode. If one or more neighboring blocks are coded using affine motion prediction, the current block can be coded using affine merge. That is, when the affine merge mode is applied, the CPMVs of the current block can be derived using the CPMVs of the surrounding blocks. In this case, the CPMVs of the surrounding blocks can be used as the CPMVs of the current block as they are, or the CPMVs of the surrounding blocks can be modified based on the size of the surrounding blocks and the size of the current block and then used as the CPMVs of the current block. Meanwhile, in the case of affine merge where MV is derived in units of subblocks, it can be called subblock merge mode, and this can be indicated based on the merge subblock flag (merge_subblock_flag (value 1)). In this case, the affine merging candidate list described later can also be called a subblock merging candidate list. In this case, the subblock merging candidate list can further include a candidate derived by SbTMVP. In this case, the candidate derived by sbTMVP can be used as the candidate for index 0 of the subblock merging candidate list. In other words, the candidate derived by sbTMVP can be positioned before the inherited affine candidates and constructed affine candidates described later in the subblock merging candidate list. When affine merge mode is applied, an affine merge candidate list may be constructed to derive CPMVs for the current block. The affine merge candidate list may include, for example, at least one of the following candidates: 1) Inherited affine candidates 2) Constructed affine candidates 3) Zero MVs candidates Here, the inherited affine candidate is a candidate derived based on the CPMVs of the surrounding blocks when the surrounding blocks are coded in affine mode, the constructed affine candidate is a candidate derived by constructing CPMVs based on the MVs of the corresponding CP surrounding blocks for each CPMV unit, and the zero MVs candidate can represent a candidate composed of CPMVs whose value is 0. For example, the zero MVs candidate can be optionally inserted into the candidate list when the number of current candidates is less than the number of maximum candidates. Figure 15 is an example showing inheritance of control point motion vectors. Two inherited affine candidates can be derived from the affine motion model of the surrounding blocks: one from the left surrounding CUs, and the other from the upper CUs. Referring back to Figure 12 described above, for the left predictor, the scanning order is A0 -> A1, and for the upper predictor, the scanning order is B0 -> B1 -> B2. Only the first inherited candidate on each side can be selected, and no pruning check is performed between the two inherited candidates. Once the surrounding affine CUs are identified, the control point motion vectors are used to derive CPMVP candidates from the affine merge list of the current CU. Referring to Figure 15, when the surrounding lower-left block A is coded in affine mode, the motion vectors v2, v3, and v4 of the upper-left corner, upper-right corner, and lower-left corner of the CU containing block A are obtained. When block A is coded with a 4-parameter affine model, two CPMVs of the current CU are calculated according to v2 and v3. When block A is coded with a 6-parameter affine model, three CPMVs of the current CU are calculated according to v2, v3, and v4. Figure 16 is a drawing showing an example of surrounding blocks for the current block. The constructed affine candidate refers to a candidate constructed by combining neighbor translational motion information of each control point. Motion information for the control point can be derived from the specified spatial and temporal surroundings, as illustrated in Fig. 16. CPMV k (k=1, 2, 3, 4) represents the kth control point. For CPMV1, blocks B2 -> B3 -> A2 are checked in that order, and the MV of the first available block is used. For CPMV2, blocks B1 -> B0 are checked in that order, and for CPMV3, blocks A1 -> A0 are checked in that order. TMVP can be used as CPMV4 if available. After acquiring the MVs for the four control points, affine merge candidates can be constructed based on motion information. The following combinations of control point MVs can be used sequentially for construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3} A combination of three CPMVs forms a six-parameter affine merge candidate, and a combination of two CPMVs forms a four-parameter affine merge candidate. To avoid the motion scaling procedure, if the reference indices of the control points are different, the corresponding combination of control point MVs is not used. Subblock-based temporal motion vector prediction (SbTMVP) Subblock-based temporal motion vector prediction (SbTMVP) can be used. Similar to Temporal Motion Vector Prediction (TMVP) in HEVC, SbTMVP uses the motion fields of collocated pictures to improve motion vector prediction and merge mode for CUs of the current picture. The same collocated pictures used in TMVP are also used in SbTMVP. SbTMVP differs from TMVP in two main aspects: 1. TMVP predicts motion at the CU level, whereas SbTMVP predicts motion at the sub-CU level. 2. While TMVP obtains the motion vector of a collocated block from a collocated picture (a collocated block is a block located at the bottom right or center (bottom-right center) of the current CU), SbTMVP obtains temporal motion information from a collocated picture after applying a motion shift. Here, the motion shift is obtained from the motion vector of one of the spatially adjacent blocks of the current CU. Figure 17 illustrates the process of SbTMVP. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, the spatially adjacent block A1, illustrated in (a) of Figure 17, is examined. If A1 has a motion vector that uses the collocated picture as a reference picture, this motion vector (which may be referred to as a temporal motion vector (tempVM)) is selected as the motion shift to be applied. If such a motion is not identified, the motion shift is set to (0,0). In the second step, the motion shift identified in the first step is applied (i.e., added to the coordinates of the current block) to obtain motion information (motion vectors and reference indices) at the sub-CU level from the collocated picture, as illustrated in (b) of Fig. 17. In the example of (b) of Fig. 17, it is assumed that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of the corresponding block (the minimum motion grid covering the center sample) in the collocated picture is used to derive the motion information of the sub-CU. The center sample (the bottom-right center sample) can correspond to the bottom-right sample among the four center samples in the sub-CU if the horizontal and vertical lengths of the sub-block are even. After the motion information of the collocated sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC. In the second step, the motion shift identified in the first step is applied (i.e., added to the coordinates of the current block) to obtain motion information (motion vectors and reference indices) at the sub-CU level from the collocated picture, as illustrated in (b) of Fig. 17. In the example of (b) of Fig. 17, it is assumed that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of the corresponding block (the minimum motion grid covering the center sample) in the collocated picture is used to derive the motion information of the sub-CU. The center sample (the bottom right center sample) may correspond to the bottom right sample among the four center samples in the sub-CU when the horizontal and vertical lengths of the sub-block are even. After the motion information of the collocated sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC. At this time, temporal motion scaling may be applied, which aligns the reference picture of the temporal motion vector with the reference picture of the current CU. Geometric partitioning mode (GPM) GPM can be supported for inter prediction. GPM is a type of merge mode that can be signaled using a CU-level flag. Other merge modes can include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. Each possible CU size is wxh = 2. m x 2 n A total of 64 partitions can be supported for , where {36} ∋ m,n, and 8x64 and 64x8 are excluded. When this mode is used, a CU can be divided into two parts by a geometrically positioned straight line. The location of the division line can be mathematically derived from the angle and offset parameters of a specific partition. Each part of the geometric partition of the CU can be inter-predicted using its own motion. Only a single prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. A uni-prediction motion constraint can be applied to ensure that only two motion-compensated predictions are required for each CU, similar to conventional bi-prediction. Figure 18 illustrates an example of GPM segments grouped at the same angle. A single predicted motion for each partition can be derived using a process as illustrated in Fig. 18. When GPM is used for the current block, a geometric partition index indicating the partition mode of the geometric partition (angle and offset) and two merge indices (one for each partition) can be additionally signaled. The maximum number of GPM candidate sizes can be explicitly signaled in the SPS, and the syntax binarization of the GPM merge indices can be specified. After predicting each part of the geometric partition, the sample values along the geometric partitioning edge can be adjusted using adaptive weighted blending processing, such as blending along the geometric partitioning edge described below. This is a prediction signal for the entire CU, and the transform and quantization processes can be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using GPM can be stored in the motion field storage for the GPM geometric partitioning mode. Combined inter and intra prediction (CIIP) Combined inter and intra prediction (CIIP) may be applied to the current block. An additional flag (e.g., ciip_flag) may be signaled to indicate whether the CIIP mode applies to the current CU. For example, when the CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., the product of the CU width and the CU height is greater than or equal to 64), and both the CU width and the CU height are less than 128 luma samples, the additional flag may be signaled to indicate whether the CIIP mode applies to the current CU. CIIP prediction combines inter-prediction signals and intra-prediction signals. The inter-prediction signal P_inter in CIIP mode can be derived using the same inter-prediction process applied in regular merge mode, and the intra-prediction signal P_intra can be derived according to the regular intra-prediction process in planar mode. The intra- and inter-prediction signals can then be combined using a weighted average. Figure 19 illustrates neighboring blocks used in CIIP weight derivation. Here, the weight values can be calculated as follows depending on the coding modes of the left and upper neighboring blocks: - If the upper neighbor is available and intra-coded, isIntraTop is set to 1, otherwise isIntraTop is set to 0. - If the left neighbor is available and intra-coded, isIntraLeft is set to 1, otherwise isIntraLeft is set to 0. - If (isIntraTop + isIntraLeft) is 2, wt is set to 3 - Otherwise, if (isIntraTop + isIntraLeft) is 1, wt is set to 2 - Otherwise, wt is set to 1 The CIIP prediction can be constructed as shown in Equation 4 below. [Formula 4] GPM including inter and intra prediction In a GPM including inter and intra prediction, final predicted samples can be generated by weighting the inter-predicted samples and intra-predicted samples for each GPM-separated region. The inter-predicted samples are derived from the inter-GPM, while the intra-predicted samples can be derived from an intra-prediction mode (IPM) candidate list and an index signaled from an encoding device. For example, the IPM candidate list size can be predefined as 3. Figure 20 illustrates available IPM candidates for GPM including inter and intra prediction. Available IPM candidates can be parallel angular mode (Parallel mode), perpendicular angular mode (Perpendicular mode), and planar mode with respect to the GPM block boundary, as illustrated in (a), (b), and (c) of FIG. 20. Furthermore, as illustrated in (d) of FIG. 20, GPMs including inter and intra prediction can be restricted to reduce signaling overhead for IPMs and prevent an increase in the intra prediction circuit size of the hardware decoder. In addition, direct motion vector and IPM storage in the GPM blending region can be introduced to further improve coding performance. In DIMD and adjacent mode-based IPM derivation, parallel modes can be registered first. Therefore, if there are no identical IPM candidates in the list, up to two IPM candidates derived from the decoder-side intra-mode derivation (DIMD) method and / or neighboring blocks can be registered. In deriving the neighbor mode, there are up to five locations for the available neighbor blocks, but this may be limited by the GPM block boundary angles already used in GPM using template matching (GPM-TM). Table 5 shows the locations of available neighboring blocks for deriving IPM candidates according to the GPM block boundary angle. A and L can represent the upper and left sides of the predicted block. [Table 5] GPM-Intra can be combined with GPM-MMVD (GPM with merge with motion vector difference). To further improve coding performance, TIMD can be used to generate IPM candidates for GPM-Intra. Parallel modes can be registered first, followed by TIMD, DIMD, and IPM candidates from neighboring blocks. Template matching (TM) Template Matching (TM) is a decoder-side motion vector (MV) derivation method that finds the optimal match between a template of the current coding unit (CU) within the current picture (i.e., the upper and / or left adjacent blocks of the current CU) and a block within a reference picture (i.e., a block of the same size as the template) to improve the motion information of the current coding unit (CU). Figure 21 is a diagram illustrating a method for deriving motion vectors from template matching. For example, as illustrated in Figure 21, a more appropriate motion vector is searched for around the initial motion of the current CU within the [-8, +8] pel search range. Furthermore, the search step size may be determined according to the AMVR mode, and TM may be applied continuously with the bidirectional matching process in merge mode. In AMVP mode, the MVP candidate is determined based on the template matching error to select the candidate with the smallest difference between the template of the current block and the template of the reference block, and then TM is performed only on the MVP candidate for MV improvement. TM improves the MVP candidate using an iterative diamond search within the pel search range [-8, +8], starting from full-pel MVD precision (or 4-pel precision in 4-pel AMVR mode). The AMVP candidate can be further improved by a cross search using the full-pel MVD precision (or 4-pel precision in 4-pel AMVR mode), and then sequentially improved to half-pel and quarter-pel precision according to Table 6 specified in AMVR mode. [Table 6] This search process ensures that AMVP candidates maintain the same MV precision specified by the AMVR mode even after the TM process. During the search process, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than a threshold equal to the area of the block, the search process ends. In merge mode, a similar search method is applied to merge candidates specified by the merge index. As shown in Table 6 above, TM can be performed up to 1 / 8 pel MVD precision, or steps exceeding the half pel MVD precision can be skipped. This depends on whether the alternative interpolation filter used when AMVR is in half pel mode is used based on the merged motion information. Additionally, when TM mode is enabled, template matching can operate as an independent process or as an additional MV enhancement process between block-based and sub-block-based bidirectional matching (BM) methods. This depends on whether BM can be applied by satisfying the activation conditions. Multi-Hypothesis Prediction (MHP) In multi-home inter prediction mode, one or more additional motion-compensated prediction signals are signaled in addition to the existing bidirectional prediction signal, and the resulting overall prediction signal can be obtained by sample-wise weighted superposition. Bidirectional prediction signal p bi And using the first additional inter prediction signal / assumption h3, the resulting prediction signal p3 can be obtained as follows. [Formula 5] The weight factor α can be specified by a new syntax element add_hyp_weight_idx according to the mapping shown in Table 7 below. [Table 7] Similarly, one or more additional prediction signals can be used. The resulting overall prediction signal can be iteratively accumulated with each additional signal, as in Equation 6 below. [Formula 6] The resulting overall prediction signal is the final p n (i.e. p with the largest index n n ) can be obtained. For example, up to two additional prediction signals can be used (i.e., n is limited to 2). The motion parameters of each additional prediction hypothesis can be explicitly signaled by specifying the reference index, the motion vector predictor index, and the motion vector difference, or implicitly signaled by specifying the merge index. A separate multi-hypothesis merge flag can distinguish these two signaling modes. Example The disclosed embodiments provide a method for enhancing compression performance by utilizing various motion vectors in an inter-frame prediction process. Specifically, by allowing a motion vector predictor (MVP) containing multiple motion vectors, the accuracy of the predicted block can be improved through a weighted sum of the prediction blocks represented by each motion vector. Furthermore, the disclosed embodiments allow for predictions beyond unidirectional and bidirectional predictions. In one embodiment, during the process of constructing MVP candidates, a candidate having multiple motion vectors may be included in the candidate list. The candidate having multiple motion vectors may represent a candidate that includes additional motion vectors in addition to unidirectional or bidirectional motion vectors (base motion vectors). The number of additional motion vectors may be predefined or signaled. The candidate having the above multiple motion vectors can be applied to various modes such as AMVP mode and AMVP-merge mode as well as merge mode. A candidate having the above multiple motion vectors may be included after the HMVP candidate, and may also be included in the list by replacing another MVP candidate (e.g., a pairwise average candidate) in the list. When the above multiple motion vectors are included in the MVP candidate list, overlap with other candidates can be checked. In the above overlap check process, the number of motion vectors, the values of the included motion vectors, or weight values that can be applied to the reference block indicated by each motion vector can be used. A candidate for constructing the above multiple motion vectors may be referred to as a target block, and a candidate included in the MVP candidate list may be used as the target block. The target block may have an independent MVP candidate list when in a mode that allows multiple motion vectors. The candidate having the above multiple motion vectors may include a target candidate and a motion vector derived using a motion vector included in a block indicated by the motion vector of the candidate. Example 1 In one embodiment, a method for including candidates with multiple motion vectors in the process of constructing MVP candidates in inter-screen prediction mode is described. In inter-screen prediction mode, motion vector predictors are used in various modes, such as AMVP mode and merge mode. Since higher predictor accuracy contributes to improved compression performance, constructing a diverse set of candidates is effective. According to one embodiment, the MVP candidate having multiple motion vectors can be applied to not only the AMVP mode or the merge mode, but also the affine mode, the AMVP-merge mode, the SMVD (Symmetric MVD), the sbTMVP, the MMVD mode, the GPM mode, the GPM-Inter / Intra mode, the CIIP mode, the TM mode, etc. However, the modes / tools listed above are merely examples to which the MVP candidate having multiple motion vectors according to the disclosed embodiment can be applied, and among the modes or tools not listed above, if there is a mode or tool that uses the MVP candidate, the MVP candidate having multiple motion vectors according to the disclosed embodiment can be applied. As mentioned above, the number of multiple motion vectors included in an MVP candidate can be determined based on the number of additional motion vectors defined in advance. For example, referring to the syntax in Table 8 below, information indicating whether multiple motion vectors are supported or allowed in SPS (e.g., sps_multiple_predictor_enabled_flag) and information indicating the maximum number of motion vectors that can be added when multiple motion vectors are allowed, i.e., when the flag is '1' (e.g., sps_max_num_additional_predictor_minus1) can be signaled. This means that 'sps_max_num_additional_predictor_minus1 + 1' additional motion information can be included in addition to the unidirectional or bidirectional motion information that an existing block can have. [Table 8] It is obvious that the locations of the information indicating whether multiple motion vectors are allowed and the information indicating the maximum number of motion vectors that can be added can be included in other upper parameter sets than SPS, such as VPS, PPS, APS, PH, SH, etc. For example, if the flag indicating whether multiple motion vectors are allowed is located in PPS, whether multiple motion vectors are allowed can be determined for each picture rather than for each sequence. In addition, the information indicating whether multiple motion vectors are allowed and the information indicating the maximum number of motion vectors that can be added can be located at different levels so that the maximum number of multiple motion vectors can be specified differently depending on the application unit. For example, as in the example of Table 8 above, sps_multiple_predictor_enabled_flag, which indicates whether multiple motion vectors are allowed, exists in SPS, and sps_max_num_additional_predictor, which indicates the maximum number of multiple motion vectors, can be signaled at different levels such as picture, slice, CTU, and CU, so that different maximum numbers of multiple motion vectors can be variably specified for each unit where sps_max_num_additional_predictor is located. Additionally, the maximum number of multiple motion vectors can be set according to a predefined value without being separately signaled. For example, as shown in the syntax of Table 9 below, when the value of information indicating whether multiple motion vectors are allowed (e.g., sps_multiple_predictor_enabled_flag) is 1, the maximum number of motion vectors that can be added can be predefined as a specific integer value without separate signaling. [Table 9] For example, a specific value of 1 or 2 may be used as the maximum number of motion vectors to which an existing block may have additional motion information, which may include one or two additional motion information pieces in addition to the unidirectional or bidirectional motion information that the existing block may have. FIG. 22 is a diagram illustrating a case in which multiple reference blocks are used in one embodiment. In one embodiment, a unidirectional or bidirectional motion vector may be defined as a regular motion vector, and a reference block obtained using the regular motion vector may be defined as a regular reference block. In addition, a motion vector added here may be defined as an additional motion vector, and a reference block obtained using the motion vector may be defined as an additional reference block. Multiple motion vectors may be defined as including a regular motion vector and an additional motion vector, and multiple reference blocks may be defined as including a regular reference block and an additional reference block. Alternatively, a regular reference block, a regular reference block, or an additional reference block may be referred to as a regular prediction block, a regular prediction block, or an additional prediction block. In addition, in some cases, a regular reference block or a regular prediction block may refer to an additional reference block or an additional prediction block, and a regular motion vector may refer to an additional motion vector. Referring to FIG. 22, when multiple motion vectors are allowed for the current block C, bidirectional basic reference blocks P0 and P1 can be acquired, and additional reference blocks P2 and P3 can be acquired based on additional motion vectors. For example, a weighted sum can be applied to the basic reference blocks and the additional reference blocks to generate a final prediction block. For convenience of explanation, it is assumed that two reference blocks are acquired for each prediction direction, but of course, the number of reference blocks acquired for each direction may change. FIG. 23 is a flowchart illustrating an example of a method for including an MVP candidate including multiple motion vectors in an MVP candidate list, according to a decoding method or encoding method according to one embodiment. The decoding method or encoding method of FIG. 23 may be performed by the aforementioned decoding device (300) or encoding device (200). Referring to FIG. 23, a decoding method or an encoding method according to one embodiment may include the steps of determining a prediction mode applied to a current block as a prediction mode using MVP (S900), constructing an MVP candidate list including MVP candidates having multiple motion vectors (S910), and generating a prediction block based on the final MVP candidate (S920). In the case of a decoding method, image information obtained from a bitstream may include information about prediction, and based on whether the prediction mode indicated by the information about prediction is a prediction mode using MVP, for example, the AMVP mode or merge mode described above, the prediction mode applied to the current block may be determined as a prediction mode using MVP (S900). Image information obtained from a bitstream may include information indicating whether multiple motion vectors are supported or allowed (e.g., sps_multiple_predictor_enabled_flag). Based on the information indicating that multiple motion vectors are allowed (e.g., the value of sps_multiple_predictor_enabled_flag is 1), an MVP candidate list including MVP candidates having multiple motion vectors may be constructed (S910). A prediction block is generated based on the final MVP candidate in the MVP candidate list (S920). The final MVP candidate refers to an MVP candidate used for prediction of the current block among the MVP candidates in the MVP candidate list. The encoding device (200) can encode information (e.g., index information) indicating the final MVP candidate and transmit it to the decoding device (300) in the form of a bitstream. The decoding device (300) can obtain the corresponding information from the transmitted bitstream and derive the final MVP candidate. If the final MVP candidate has multiple motion vectors, a prediction block can be generated using an additional reference block together with a basic reference block. For example, a prediction block can be generated by applying weights to the basic reference block and the additional reference block. Meanwhile, it goes without saying that the description of the above-described decoding method and encoding method can be applied to one embodiment as well. For example, a decoding method according to one embodiment can generate a reconstructed block based on residual information (which can be omitted depending on the prediction mode) included in the generated prediction block and image information after generating the prediction block. An encoding method according to one embodiment can generate residual information (which can be omitted depending on the prediction mode) based on the generated prediction block, and encode image information including information about the prediction and the residual information, and transmit it in the form of a bitstream. FIG. 24 is a diagram illustrating an example of a method for constructing an MVP candidate list in the disclosed embodiment. An example of Fig. 24 illustrates the process of constructing an MVP candidate list in the basic merge mode. Referring to Fig. 24, spatial neighboring candidates (S911), temporal candidates (S912), non-adjacent spatial neighboring candidates (S913), HMVP candidates (S914), candidates with multiple motion vectors (S915), and pairwise candidates (S916) may be included in the MVP candidate list. However, the types, order, and number of MVP candidates illustrated in FIG. 24 are merely examples applicable to one embodiment, and the types, order, or number may be changed, or existing MVP candidates may be replaced with MVP candidates having multiple motion vectors. For example, an MVP candidate including multiple motion vectors may be included in the MVP candidate list, replacing a pairwise average candidate. Meanwhile, after the MVP candidate list is constructed (S910), a reordering based on template matching costs may be performed on each candidate within the MVP candidate list. In the process of calculating the template matching cost for reordering candidates, the cost of a unidirectional reference block may be calculated based on the difference between the adjacent sample of the current block and the adjacent sample of the reference block, and the cost of a bidirectional reference block may be calculated based on the difference between the adjacent sample of the current block and the adjacent sample of each reference block. For candidates containing multiple motion vectors, the cost can be calculated using adjacent samples of the current block and adjacent samples of reference blocks indicated by each motion vector, or the complexity of the calculation can be reduced by using only some of the reference blocks. In this case, if the final prediction block is determined by a weighted sum for each reference block, the cost can be calculated by applying the weight for each reference block to the adjacent samples of each reference block. If only some of the reference blocks are used, the cost can be calculated by applying a corrected form of weight. For example, if only some of the reference blocks are used for cost calculation, a method can be applied to increase the weight for the reference blocks and decrease the weights of the reference blocks not used for cost calculation. However, this is only one example, and it is of course true that various other methods of applying a corrected form of weight can also be applied. In addition, in the process of constructing a list of MVP candidates, a method of constructing multiple candidates including multiple motion vectors, and then reordering and selecting only some of the candidates based on template matching cost can be applied. The process of calculating the template matching cost can be applied to the same cost calculation method as the example described above. In addition, since a candidate including multiple motion vectors means that it includes at least two reference blocks, reordering and selecting only some of the candidates can be performed by calculating the bilateral matching cost using each prediction block. The cost can be calculated based on the difference value between each reference block, and it can be calculated for multiple reference blocks, or it can be calculated only for the selected reference blocks by selecting some reference blocks. Meanwhile, in the process of constructing the MVP candidate list, a redundancy check between the MVP candidate with multiple motion vectors and other candidates can be performed as follows: - Duplication can be determined based on the number of reference blocks for each candidate. - If each candidate to be compared has the same number of reference blocks, the presence or absence of overlap can be determined using the reference index and motion vector representing the motion information of each candidate. - If there are multiple candidates with multiple motion vectors in the MVP candidate list, and a weighted sum between reference blocks obtained from each motion vector can be applied, and the weights between the reference blocks of each candidate are different, they can be considered as non-overlapping candidates. When constructing multiple candidates with multiple motion vectors, the motion vectors contained in each candidate can be restricted to include at least one motion vector different from the motion vectors contained in the preceding candidate with multiple motion vectors in the list. This allows for determining overlap based solely on motion information between candidates, without a weighted overlap check. Example 2 The present embodiment relates to various signaling methods when applying the method of embodiment 1 to merge mode and AMVP mode. The MVP candidate having the above multiple motion vectors can be applied to each tool (Sub-block based MERGE mode, CIIP mode, GPM mode, etc.) that can be combined with the existing merge mode and the merge mode, and it is possible to include a candidate having multiple motion vectors in the process of constructing the MVP candidate list of each mode. It is also possible to implement another merge mode within the basic merge mode. For example, as a merge mode with multiple motion vectors, a flag indicating the mode (e.g., multi_pred_merge_flag) can be signaled. When the value of the flag is '1', the MVP candidate list can be composed of candidates with multiple motion vectors, and a merge index indicating a specific candidate within the MVP candidate list can be signaled. When the process of reordering and selecting only some candidates using the template-based cost or the bilateral-based cost described above is performed, the merge index can be omitted or indicate a candidate within the list that includes only some candidates. Candidates with multiple motion vectors can also be applied to AMVP mode. In AMVP mode, MVP candidate lists can be constructed for each direction, so it is possible to include candidates with multiple motion vectors in each MVP candidate list in the L0 and L1 directions. Figure 25 is a diagram showing an example of an MVP candidate list configured for each direction in AMVP mode. When allowing additional motion vectors in AMVP mode, the number of motion vectors that can be had in each direction can be limited to 2 to 4 motion vectors in the end. In the example of Fig. 25, Case 1 represents a case in which a total of 4 MVPs are obtained through a combination of candidate CAND[2] with multiple motion vectors in the L0 direction MVP candidate list and candidate CAND[2] with multiple motion vectors in the L1 direction MVP candidate list. Case 2 represents a case in which a total of 3 MVPs are obtained through a combination of candidate CAND[0] with a single motion vector in the L0 direction MVP candidate list and candidate CAND[2] with multiple motion vectors in the L1 direction MVP candidate list. The application method in the above AMVP mode can be changed as follows in consideration of encoding / decoding efficiency. Specifically, additional motion vectors for the AMVP mode can be applied only to blocks to which bidirectional prediction is applied. In addition, a limited number of multiple motion vectors can be allowed by configuring a candidate having multiple motion vectors only in the process of configuring MVP candidates in a specific direction (L0 or L1). That is, by allowing n (n is a natural number greater than or equal to 1) additional motion vectors only in a specific prediction direction, the total number of additional motion vectors can be limited to n. Meanwhile, it is also possible to construct a candidate having multiple motion vectors without signaling whether or not multiple motion vectors are allowed, and it is also possible to signal a flag indicating the mode to allow multiple motion vectors. In the latter case, if the mode flag is signaled, and the flag has a value of '1', an MVP candidate having multiple motion vectors can be included in the MVP candidate list. The flag may mean that a candidate having multiple motion vectors consisting of the number of allowed motion vectors among the MVP candidates can be constructed, or it may mean that an MVP candidate list consisting only of candidates having multiple motion vectors can be constructed. Alternatively, it may mean the number of multiple motion vectors that can be included in the MVP candidate list. In addition, mvp_index or mvp_flag, which points to a specific candidate in the MVP candidate list, may be signaled, and when the process of reordering and selecting only some candidates using the template-based cost or bilateral-based cost is performed, mvp_index or mvp_flag may be omitted or may point to a candidate in the list that only includes some candidates. Meanwhile, in AMVP mode, if the final MVP candidate, i.e., the MVP candidate used for predicting the current block, is a candidate with multiple motion vectors, MVD information corresponding to each motion vector can be signaled. Similarly to the conventional AMVP mode, the final MV can be calculated using the derived MVP information and the signaled MVD information. At this time, the signaling method for the MVD information can be applied in various ways, such as indicating the magnitude of the vector or deriving it from information based on distance and direction by indexing the MVD similarly to the MMVD. Furthermore, for an MVP with multiple motion vectors, the MVD information can be reduced in signaling information by signaling only some information, limited to the basic (regular) motion vector, rather than the additional motion vector. In other words, the additional motion vector predictor included in the MVP candidate can be utilized as a motion vector as is, without the MVD. Alternatively, when the final MVP candidate has multiple motion vectors, variations are possible, such as transmitting MVD information for some motion vectors and not transmitting MVD information for some motion vectors, or not transmitting the entire MVD information. The method described in the disclosed embodiment can be similarly applied to AMVP-merge. Additional motion vectors can be included in each direction of the AMVP-merge mode, and it is possible to allow additional motion vectors in a modified way, such as applying them only to each of the AMVP mode or merge mode directions. When including additional motion vectors in the AMVP mode direction, the candidate indicated by mvp_index or mvp_flag can include multiple motion vectors. In this case, modifications such as additionally signaling MVD information or not signaling the additional motion vectors are possible, and when including additional motion vectors in the merge mode direction, the candidate indicated by the merge index can include multiple motion vectors. Example 3 This embodiment describes a method for generating an MVP candidate with multiple motion vectors. According to one embodiment, a candidate with multiple motion vectors can be generated by combining other candidates included in the MVP candidate list. The MVP candidate list may be composed of motion vectors of spatially and temporally adjacent blocks, history-based MVPs (HMVPs), pairwise average candidates, etc., and in one embodiment, such candidates are referred to as target candidates or target blocks. A candidate having multiple motion vectors may be generated using the motion vectors included in the target candidates, and in one embodiment, the motion vector used to generate the multiple motion vectors is referred to as a target motion vector. In addition, not only the candidates included in the MVP candidate list, but also the candidates in another list configured to generate multiple motion vectors may be considered as target candidates. The names “target candidate” and “target motion vector” are given for convenience of explanation, and even if they are referred to by different names, if they are included in the MVP candidate list and are used to generate a candidate having multiple motion vectors, they may be included in the range of the target candidates and target motion vectors in one embodiment. As mentioned above, as a single MVP candidate, multiple motion vectors can be included between the HMVP and pairwise average candidates in the MVP candidate list. Alternatively, they can be applied as a replacement for existing candidates, such as the pairwise average candidate. When the maximum number of MVP candidates is N, if the number of candidates included in the candidate list is less than N, a candidate including multiple motion vectors can be considered, and if the number of candidates included in the list is less than M (e.g., 2), this process can be omitted. In this case, N and M are integer values greater than 0. FIG. 26 is a diagram illustrating another example of a method for generating a candidate having multiple motion vectors according to one embodiment. As mentioned above, a candidate with multiple motion vectors can be generated by combining other candidates included in the MVP candidate list. Referring to the example of Fig. 26, when a list is composed of at most N MVP candidates, other candidates included in the list, CAND[0]...CAND[M-1](M <N) 중 일부를 조합하여 다중 움직임 벡터를 갖는 후보를 생성할 수 있다. According to the example, the first target candidate CAND[0] and the second target candidate CAND[1] on the MVP candidate list can be combined to generate a candidate CAND[M]: CAND[0] + CAND[1] having multiple motion vectors, the first target candidate CAND[0] and the third target candidate CAND[2] can be combined to generate a candidate CAND[M+1]: CAND[0] + CAND[2] having multiple motion vectors, and the second target candidate CAND[1] and the third target candidate CAND[2] can be combined to generate a candidate CAND[M+2]: CAND[1] + CAND[2] having multiple motion vectors. In the example, the case of generating three candidates having multiple motion vectors is described, but the number of candidates having multiple motion vectors generated by combining target candidates may be different from this, and the number of target candidates used in the combination, the order of target candidates in the list, the order of candidates having multiple motion vectors, etc. may also be implemented differently from the example of Fig. 25. The target candidate used to generate multiple motion vectors can be an MVP candidate for unidirectional prediction or an MVP candidate for bidirectional prediction. Depending on the allowable number of additional motion information for a candidate with multiple motion vectors, all or part of the motion vectors of the target candidate can be used. For example, if the target candidates for unidirectional prediction, or the target candidates for bidirectional prediction, and the candidates for unidirectional prediction satisfy the allowable number of additional motion vectors, all motion information can be included in the candidate with multiple motion vectors without any additional conditions. Additionally, multiple motion vectors can be constructed in the following order. In this example, to generate candidates with multiple motion vectors, we assume that there are target candidates A and B for bidirectional prediction. In this case, candidates A and B can be rearranged within the candidate list based on template-based costs or bidirectional-based costs. - The motion vectors of candidate A in the L0 and L1 directions are included first, and the motion vectors of candidate B in the L0 or L1 direction can be additionally included. - The motion vectors of candidate B in the L0 and L1 directions are included first, and the motion vectors of candidate A in the L0 or L1 direction can be additionally included. - The motion vector in the L0 direction of candidate A and the motion vector in the L1 direction of candidate B are first included, and then the motion vector in the L1 direction of candidate A or the motion vector in the L0 direction of candidate B may be additionally included. - The motion vector in the L1 direction of candidate A and the motion vector in the L0 direction of candidate B are first included, and then the motion vector in the L0 direction of candidate A or the motion vector in the L1 direction of candidate B may be additionally included. - The motion vector in the L0 direction of candidate A and the motion vector in the L0 direction of candidate B are first included, and then the motion vector in the L1 direction of candidate A or the motion vector in the L1 direction of candidate B may be additionally included. - The motion vector in the L1 direction of candidate A and the motion vector in the L1 direction of candidate B are first included, and then the motion vector in the L0 direction of candidate A or the motion vector in the L0 direction of candidate B may be additionally included. Additionally, the MVP candidate list with the added candidate having the above multiple motion vectors can undergo an additional reordering process using a predefined cost (e.g., template-based cost, bilateral-based cost, etc.). The above-described embodiment described a case where a list of MVP candidates is reordered based on a predefined cost, and then candidates with multiple motion vectors are generated based on the reordered MVP candidates. However, to reduce encoding and decoding complexity, it is also possible to perform reordering based on a predefined cost after generating multiple motion vectors and adding them to the MVP candidate list, rather than performing reordering on the MVP candidate list before generating candidates with multiple motion vectors. In addition, when only some of the motion vectors need to be selected, such as when a candidate having other multiple motion vectors is used to generate a candidate having multiple motion vectors, or when the number of motion vectors of the target candidate exceeds the allowable number of additional motion vectors, the methods listed above can be considered, but the additional motion vectors can be determined by allowing only some of the motion vectors of the candidates to be combined. As a specific example, in addition to the motion information primarily included in the methods listed above, additional motion information can be added by using the distance between the reference picture of each candidate to be compared and the current picture, and including the motion information of candidates with a close distance to the current picture as additional motion information. Alternatively, when a weighted sum is applied between reference blocks represented by each motion vector, the motion vector representing the reference block to which a large weight is applied can be included as an additional motion vector. Alternatively, the encoding / decoding process can be simplified by including a motion vector in a fixed direction. As in the example described above, it is possible to generate multiple candidates having multiple motion vectors and include them in the MVP candidate list. For example, the number of candidates having multiple motion vectors that can be included in the MVP candidate list can be predetermined. Alternatively, in order to consider candidates that can be included in the MVP candidate list after the candidate having multiple motion vectors, multiple multiple motion vectors can be generated until the number of MVP candidates including candidates having multiple motion vectors reaches N-2. Here, 2 is just one example, and it is also possible to change it to N-1, N-3, etc., considering the locations of multiple reference blocks and apply it. Example 4 This embodiment describes another method for generating an MVP candidate with multiple motion vectors. According to the disclosed embodiment, a candidate with multiple motion vectors can be generated by combining the motion vectors of other candidates included in the MVP candidate list and motion vectors derived using the same. As described above, the MVP candidate list may be composed of motion vectors of spatially and temporally adjacent blocks, history-based MVPs (HMVPs), pairwise average candidates, etc., and in the disclosed embodiment, such candidates are referred to as target candidates. A candidate having multiple motion vectors may be generated using the motion vectors included in the target candidates, and in the disclosed embodiment, the motion vector used to generate the multiple motion vectors is referred to as a target motion vector. In addition, not only the candidates included in the MVP candidate list, but also the candidates in another list configured to generate multiple motion vectors may be considered as target candidates. The names “target candidate” and “target motion vector” are given for the convenience of explanation, and even if they are referred to by different names, as long as they are included in the MVP candidate list and used to generate a candidate having multiple motion vectors, they may be included in the scope of the target candidates and target motion vectors in the disclosed embodiment. For example, as a single MVP candidate, multiple motion vectors can be included between HMVP and pairwise average candidates in the MVP candidate list. Alternatively, they can be applied as a replacement for existing candidates such as pairwise average candidates. When the maximum number of MVP candidates is N, if the number of candidates included in the candidate list is less than N, a candidate including multiple motion vectors can be considered, and if the number of candidates included in the list is less than M (e.g., 2), this process can be omitted. In this case, N and M are integer values greater than 0. As a concrete example, when any candidate in the list is CAND[i] and the block at the position indicated by the motion vector in the L0 or L1 direction of CAND[i] is called a corresponding block (reference block), if the corresponding block is an inter mode (if the corresponding block is an inter block), an additional motion vector can be derived. That is, if the inter mode is applied to the corresponding block, the motion vector and reference picture of the corresponding block can be used as additional motion information. Therefore, when the motion vector in the L0 or L1 direction of CAND[i] is MV[X], (X=0..1), and the motion vector of the corresponding block indicated by the motion vector is CX_MV[Y], (X=0..1, Y=0..1), the additional motion vector can be calculated as MV[X] + CX_MV[Y]. In this case, CX_MV[Y] represents the motion vector in the LY direction of the corresponding block in the LX direction of CAND[i]. Therefore, a candidate having multiple motion information can be composed of the following motion vectors, or can be composed of some of the motion vectors listed below. - MV[0], - MV[1], - MV[0] + C0_MV[0], - MV[0] + C0_MV[1], - MV[1] + C1_MV[0] - MV[1] + C1_MV[1] That is, an MVP candidate having multiple motion vectors can be constructed using the motion vectors MV[0], MV[1] of the target block, and MV[0] + C0_MV[0], MV[0] + C0_MV[1], MV[1] + C1_MV[0], MV[1] + C1_MV[1] derived using the motion vectors MV[0] and MV[1] of the target block and the bidirectional motion vectors of the corresponding block. However, this is just one example applicable to one embodiment, and various combinations can be generated depending on which direction of the motion vector of the target block and which direction of the motion vector of the corresponding block are taken. At this time, the reference picture indicated by the additional motion information may be the reference picture indicated by the motion information of the corresponding block. In addition, the motion vector of the corresponding block can be scaled through the distance between the picture of the target block and the reference picture and the distance between the picture of the corresponding block and the picture indicated by the motion vector of the corresponding block. Meanwhile, the above-described method is not limited to when the corresponding block is in inter mode, and the same method can be applied to obtain additional motion information using a block vector when the corresponding block is in IBC mode. That is, various additional motion vectors can be generated by combining a motion vector and a block vector. Specifically, when the motion vector of the target block is MV[X] (X = 0..1) and the block vector when the corresponding block is in IBC mode is CX_BV[Y] (X=0..1, Y=0..1), the additional motion vector can be calculated as MV[X] + CX_BV[Y]. At this time, a candidate having multiple motion information can be composed of the following motion vectors, or can be composed of some of the motion vectors listed below. When supporting a unidirectional block vector in the IBC mode, it is possible to calculate an additional motion vector using a single CX_BV. - MV[0], - MV[1], - MV[0] + C0_BV[0], - MV[0] + C0_BV[1], - MV[1] + C1_BV[0], - MV[1] + C1_BV[1] That is, an MVP candidate having multiple motion vectors can be constructed using the motion vectors MV[0], MV[1] of the target block, and MV[0] + C0_BV[0], MV[0] + C0_BV[1], MV[1] + C1_BV[0], MV[1] + C1_BV[1] derived using the bidirectional block vectors of the motion vector of the target block and the corresponding block. However, this is just one example applicable to one embodiment, and various combinations can be generated depending on which direction of the motion vector of the target block is taken and which direction of the block vector of the corresponding block is taken. Since the block vector is used for prediction within the screen, the reference picture in the above example can be a reference picture indicated by the motion vector of the target block. If the corresponding block is an intra-mode block (i.e., if the intra-mode is applied to the corresponding block), the motion vector can be obtained from the corresponding block indicated by the motion vector in a different direction, or a candidate having multiple motion vectors can be derived from another candidate included in the MVP candidate list without using the target block. Alternatively, even if the intra-mode is applied to the corresponding block, if the block vector of the corresponding block is available, the block vector of the corresponding block can be used to obtain additional motion information. That is, as described above, various additional motion vectors can be generated by combining the motion vector and the block vector. Specifically, when the motion vector of the target block is MV[X] (X = 0..1) and the block vector when the corresponding block is in intra mode is CX_BV[Y] (X=0..1, Y=0..1), the additional motion vector can be calculated as MV[X] + CX_BV[Y]. A candidate with multiple motion information can be composed of the following motion vectors, or can be composed of some of the motion vectors listed below. If the corresponding block is in IBC mode and the IBC mode supports unidirectional block vectors, it is possible to calculate the additional motion vector using a single CX_BV. - MV[0], - MV[1], - MV[0] + C0_BV[0], - MV[0] + C0_BV[1], - MV[1] + C1_BV[0], - MV[1] + C1_BV[1] This is an example applicable to one embodiment, and various combinations can be generated depending on which direction of the motion vector of the target block is taken and which direction of the motion vector of the corresponding block is taken. Since the block vector is used for prediction within the screen, the reference picture in the above example can be the reference picture indicated by the motion vector of the target block. In the above example, in intra mode, a block vector is available if the corresponding block contains a block vector. For example, block vectors may be available in the following modes: - Intra TMP: This is a mode that creates a prediction block by finding a block with a small error through template matching between the template of the current block and the template of the reference block in the current picture. - SGPM-IBC: SGPM is an intra-prediction technique that performs different predictions on each region divided by partitioning. Each region is subject to a combined intra / intra prediction, and SGPM-IBC can apply intra / IBC, IBC / intra, or IBC / IBC predictions. The above modes are examples of intra modes in which block vectors are available. In addition to the examples above, block vectors may be used and stored in intra modes through combinations of intra TMP and other intra modes, IBC mode and other intra modes, etc. In other words, any intra mode that uses and stores block vectors may be included in the scope of the disclosed embodiments even if it is not one of the modes listed above. As described above, an MVP candidate having multiple motion vectors can be generated by considering the motion vectors of the target block and the corresponding blocks indicated by the motion vectors of the target block, but the specific method may be applied differently depending on the allowable number of additional motion vectors. For example, when bidirectional prediction is applied to the target block CAND[i], there are motion vectors in the L0 or L1 direction, and when bidirectional prediction is applied to all corresponding blocks in each direction, a total of six motion vectors can be obtained. Depending on the allowable number of additional motion vectors (the allowable number of additional reference blocks), some of the motion vectors can be selected to form a candidate having multiple motion vectors. More specifically, a candidate having multiple motion information can be generated in the following manner: - If a target candidate A exists and bidirectional prediction is applied to it, the motion vectors of the target candidate A in the L0 and L1 directions are first included, and when the corresponding block indicated by the motion vectors in the L0 and / or L1 directions is an inter mode, an intra mode including a block vector, or an IBC mode, the motion vector or block vector of the corresponding block can be used when deriving an additional motion vector. - If the corresponding block indicated by the motion vector in the LX (X = 0 or 1) direction is not an inter mode, an intra mode containing a block vector, or an IBC mode, another reference block can be derived and an additional motion vector can be generated by repeating the same process using the motion vector in the L (1-X) direction. - First, a motion vector in the LX (X = 0 or 1) direction is searched, and when the corresponding block indicated by the motion vector in the corresponding direction is an inter mode, an intra mode including a block vector, or an IBC mode, the motion vector or block vector of the corresponding block can be used to derive an additional motion vector. First, when the corresponding block in the searched direction allows an additional motion vector, such as a unidirectional prediction, it is possible to determine whether additional motion vectors are allowed for the corresponding blocks in the remaining directions. - When bidirectional prediction is applied to a target candidate, the distance between the current picture and the reference pictures in the L0 and L1 directions of the target candidate can be considered to determine the direction to search first. For example, the direction with the shortest distance can be preferentially selected. Figure 27 shows the distance between the current picture and the reference pictures. That is, by comparing d0 and d1 in Figure 27, the motion vector in the direction with the smaller value can be prioritized. - The motion vectors in the L0 and L1 directions included in the target candidate are included in the MVP candidate having multiple motion vectors, and as an additional motion vector, the motion vector in the direction with a short distance can be considered by considering the distance between the current picture and the reference picture including the additional reference block derived from the corresponding block in each of the L0 and L1 directions. That is, by comparing D0 and D1 of Fig. 27, the motion vector in the direction with a small value can be given priority. - Motion vectors in the L0 and L1 directions included in the target candidate are included in the MVP candidate having multiple motion vectors, and by considering the distance to the reference picture including the reference picture in each of the L0 and L1 directions and the additional reference block derived from the corresponding block in each of the L0 and L1 directions as additional motion vectors, a motion vector in a direction with a short distance can be given priority. That is, by comparing (D0 - d0) and (D1 - d1) of Fig. 27, a motion vector in a direction with a small value can be given priority. - If the reference blocks in each direction of L0 and L1 included in the target candidate are different from each other, such as an inter-block and an IBC block or an intra-mode including a block vector, the motion vector of the inter-block may be given priority. Alternatively, the order may be defined as the motion vector of the inter-block, the block vector of the IBC block, and the block vector of the INTRA block. - When the target candidate has a single reference picture list in which the reference pictures in each direction of L0 and L1 are reordered based on a specific cost, a motion vector pointing to a reference picture located at the top of the list, i.e., having a small cost, can be given priority. - Motion information of blocks indicated by motion vectors in the L0 and L1 directions included in the target candidate can be obtained one by one. For example, if all / some bidirectional predictions are applied to blocks indicated by motion vectors in the L0 and L1 directions, one of the motion vectors can be selected and used as an additional reference vector. - When bidirectional prediction is applied to a corresponding block indicated by a motion vector in the LX direction of a target candidate, the motion vector in the LX direction of the corresponding block can be used when calculating an additional motion vector. - The motion vector in the L0 direction included in the corresponding block can be used when calculating an additional motion vector. - Among the motion vectors included in the corresponding block, a motion vector indicating a reference picture with a close distance to the current picture can be used when calculating an additional motion vector. In addition to the motion information primarily included in the methods listed above, additional motion information can be used by calculating the distance between the reference picture of each candidate to be compared and the current picture, and using the motion information of candidates with a close distance to the current picture as an additional motion vector. Alternatively, when a weighted sum is applied between the reference blocks represented by each motion vector, the motion vector pointing to the reference block to which a large weight is applied can be used as an additional motion vector. Alternatively, the encoding / decoding process can be simplified by including a motion vector in a fixed direction. Meanwhile, the number of candidates with multiple motion vectors that can be included in the MVP candidate list can be predetermined. Alternatively, the number of MVP candidates that include candidates with multiple motion vectors can be set to N-2, allowing consideration of candidates that can be included in the MVP list after candidates with multiple motion vectors. Note that 2 is just one example, and it is also possible to change this to N-1, N-3, etc., considering the locations of multiple reference blocks. As in the example described above, it is possible to generate multiple candidates having multiple motion vectors and include them in the MVP candidate list. For example, the number of candidates having multiple motion vectors that can be included in the MVP candidate list can be predetermined. Alternatively, in order to consider candidates that can be included in the MVP candidate list after the candidate having multiple motion vectors, multiple multiple motion vectors can be generated until the number of MVP candidates including candidates having multiple motion vectors reaches N-2. Here, 2 is just one example, and it is also possible to change it to N-1, N-3, etc., considering the locations of multiple reference blocks and apply it. Example 5 The present embodiment describes a method for deriving additional motion information based on the position of a prediction block indicated by a motion vector of a target block CAND[i]. The motion vector included in the current block is stored in units of a fixed size (e.g., 16x16) rather than CU block units after encoding / decoding in units of pictures in order to compress the amount of motion information in units of pictures. Fig. 28 is a diagram showing an example of the position of a reference block indicated by the motion vector of the current block. Referring to Fig. 28, when the motion vector of the current block indicates a specific position within an encoded / decoded reference picture, an example in which the reference block for the current block is included in one motion information storage unit (CASE 1) and an example in which it is included in multiple storage units (CASE 2) are shown. CASE 1 of Fig. 28 shows a case in which the motion vector of CAND[i] points to one storage unit in which motion information within the reference picture is stored. In addition, CASE 2 shows a case in which the motion vector points to a position including multiple storage units within the reference picture. Typically, the position pointed to by the motion vector is set to the (0, 0) position of the block, and the block containing the center sample at the (width / 2, height / 2) position can be used as the corresponding block to use the motion vector of that block. However, if the positions point to different motion vectors, as in CASE 2, this can lead to inaccurate motion vectors. In the disclosed embodiment, the motion vector stored at the position indicated by the motion vector of CAND[i] can be used as an additional motion vector or used to calculate an additional motion vector. Specifically, the motion vector of a block including a specific position within the corresponding block can be used, and various modifications are possible, such as, for example, the motion vector of a block including a sample at the central position (width / 2, height / 2) being used. In addition, the motion vector of a block including a sample at a position shifted by a predefined offset can also be used as an additional motion vector. Figure 29 is a drawing showing an example of an offset applicable to one embodiment. As shown in (a) of Fig. 29, a block including a corrected position can be found by applying one of the four offset candidates based on the center position as in Equation 7 below. [Formula 7] (width / 2, height / 2) (width / 2, height / 2) + (Offset, 0) (width / 2, height / 2) + (-Offset, 0) (width / 2, height / 2) + (0, Offset) (width / 2, height / 2) + (0, -Offset) The offset can have an integer value, and for example, values such as 1, 2, 4, 8, and 16 can be applied. In addition, the motion vector at the corrected position as in Equation 7 can be expressed as C_MV[X][0]..C_MV[X][K-1], etc. Here, X represents the prediction direction and can have a value of 0 or 1, and K can represent the number of corrected position candidates that can have. When the above target block CAND[i] exists, multiple motion vectors can be derived by applying various correction values, and the thus derived motion vectors can be used to construct multiple additional motion vectors as follows. - MV[0], - MV[1], - MV[0] + C_MV[0][0], - MV[0] + C_MV[0][1] Figure 30 is a diagram showing an example of a process of correcting a position by applying an offset within a corresponding block and deriving a motion vector at each corrected position. Referring to FIG. 30, the position can be corrected using an offset within the corresponding block indicated by the motion vector MV[0] pointed to by the target block CAND[i], and two motion vectors, MV[0] + C_MV[0][0] and MV[0] + C_MV[0][1], can be derived at each corrected position. Alternatively, when the target block CAND[i] exists, it is possible to generate multiple candidates with additional motion vectors by applying various offsets. Table 10 below shows an example of generating two candidates with additional motion vectors. Referring to Table 10, two different MVP candidates can be generated by using the motion vectors MV_i[0], MV_i[1] of the target block CAND[i], but using the motion vectors C_MV[0][0] and C_MV[0][1] of the corresponding block to which the offset is applied. [Table 10] Alternatively, when a target block CAND[i] exists and another target block CAND[j], (i != j), various MVP candidates can be generated by applying the motion vectors C_MV[0][0] and C_MV[0][1] of the corresponding blocks with different offsets. Table 11 below shows an example of generating two candidates with additional motion vectors using the motion vectors MV_i[X] and MV_j[X] of different target blocks. In this case, X has a value of 0 or 1 and represents the x-axis or the y-axis. [Table 11] That is, it is possible to generate multiple candidates by correcting the position using an offset within the corresponding block indicated by the motion vector as well as the motion vector MV[0] included in the target block CAND[i]. Alternatively, it is also possible to generate multiple candidates by deriving from one or more predetermined sample positions within the corresponding block. For example, the motion vector of a block including samples at the positions of each vertex, i.e., (0, 0), (width-1, 0), (0, height-1), (width-1, height-1) or some positions, may be used as an additional motion vector or used to calculate an additional motion vector. Alternatively, variations such as selecting a representative motion vector covered by the largest number of samples and using it as a candidate are possible. In addition, it goes without saying that an overlap check may be performed between candidates so that various candidates can be considered when generating an additional motion vector. Although the embodiments have been described separately for the convenience of explanation so far, a combination of two or more embodiments is possible, and changes required by a combination of embodiments may also be included in the scope of the disclosed invention or the disclosed embodiments. FIG. 31 is a diagram illustrating an example of a content streaming system to which an embodiment according to the present disclosure can be applied. Referring to FIG. 31, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device. The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted. The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream. The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system. The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time. Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc. Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner. The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method. Embodiments according to the present disclosure can be used to encode / decode images.

Claims

1. A step of obtaining image information from a bitstream; A step of determining a prediction mode to be applied to the current block based on the acquired image information as a prediction mode using a motion vector predictor (MVP); A step of constructing an MVP candidate list including MVP candidates for the current block; and A step of generating a prediction block for the current block based on at least one MVP candidate in the MVP candidate list; The steps for forming the above MVP candidate list are: A decoding method comprising including an MVP candidate having a base motion vector and an additional motion vector in the MVP candidate list.

2. In paragraph 1, The image information obtained above is, A decoding method comprising at least one of information related to whether the additional motion vector is allowed or information related to the maximum number of the additional motion vectors.

3. In paragraph 1, Further comprising a step of reordering the MVP candidate list based on the calculated cost for the MVP candidates; The steps for performing the above rearrangement are: A decoding method comprising applying weights to adjacent samples of a reference block indicated by the basic motion vector and adjacent samples of a reference block indicated by the additional motion vector to calculate a cost for an MVP candidate having the basic motion vector and the additional motion vector.

4. In paragraph 1, The steps for forming the above MVP candidate list are: A decoding method comprising checking for redundancy between an MVP candidate having the base motion vector and an additional motion vector and another MVP candidate in the MVP candidate list.

5. In paragraph 1, The steps for forming the above MVP candidate list are: A decoding method for including an MVP candidate having the base motion vector and the additional motion vector in one of an MVP candidate list in the L0 direction and an MVP candidate list in the L1 direction, based on the prediction mode applied to the current block being an AMVP (Advanced MVP) mode.

6. In paragraph 1, The image information obtained above is, A decoding method, based on the fact that the final MVP candidate among the above MVP candidates is an MVP candidate having the basic motion vector and the additional motion vector, including motion vector differential information for at least one of the basic motion vector and the additional motion vector.

7. A step of determining the prediction mode applied to the current block as a prediction mode using a motion vector predictor (MVP); A step of constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on the final MVP candidate in the MVP candidate list; and A step of encoding image information including information about the above prediction mode; The steps for forming the above MVP candidate list are: An encoding method comprising including an MVP candidate having a base motion vector and an additional motion vector in the MVP candidate list.

8. In paragraph 7, The above encoded video information is, An encoding method comprising at least one of information related to whether the additional motion vector is allowed or information related to the maximum number of the additional motion vectors.

9. In paragraph 7, Further comprising a step of reordering the MVP candidate list based on the calculated cost for the MVP candidates; The steps for performing the above rearrangement are: An encoding method comprising applying weights to adjacent samples of a reference block indicated by the basic motion vector and adjacent samples of a reference block indicated by the additional motion vector to calculate a cost for an MVP candidate having the basic motion vector and the additional motion vector.

10. In paragraph 7, The steps for forming the above MVP candidate list are: An encoding method comprising checking for redundancy between an MVP candidate having the base motion vector and an additional motion vector and another MVP candidate in the MVP candidate list.

11. In paragraph 7, The steps for forming the above MVP candidate list are: An encoding method for including an MVP candidate having the base motion vector and the additional motion vector in one of an MVP candidate list in the L0 direction and an MVP candidate list in the L1 direction, based on the prediction mode applied to the current block being an AMVP (Advanced MVP) mode.

12. In paragraph 7, The step of encoding the above image information is: An encoding method comprising encoding motion vector differential information for at least one of the basic motion vector and the additional motion vector, based on the final MVP candidate being an MVP candidate having the basic motion vector and the additional motion vector.

13. In a computer-readable storage medium storing a bitstream generated by an encoding method, The above encoding method is, A step of determining a prediction mode applied to the current block as a prediction mode using a motion vector predictor (MVP); A step of constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on the final MVP candidate in the MVP candidate list; and A step of encoding image information including information about the above prediction mode; The steps for forming the above MVP candidate list are: A storage medium comprising an MVP candidate having a base motion vector and an additional motion vector, and including the MVP candidate list.

14. In the method of transmitting data for video, A step of obtaining a bitstream for the image, wherein the bitstream is generated based on the steps of: determining a prediction mode applied to a current block as a prediction mode using a motion vector predictor (MVP); constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on a final MVP candidate in the MVP candidate list; and encoding image information including information about the prediction mode; and A step of transmitting the data including the bitstream; The steps for forming the above MVP candidate list are: A transmission method including an MVP candidate having a base motion vector and an additional motion vector in the MVP candidate list.

Citation Information

Patent Citations

  • Battery pack and device including the same

    KR1020230124486A

  • Image decoding method and device based on inter prediction in an image coding system

    KR102387363B1

  • Generating system for using small hydropower

    KR102818755B1

  • KR20210124270A

  • KR20240004159A