Encoding method, decoding method, computer-readable storage medium, and transmission method

CN122700512APending Publication Date: 2026-09-04LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580013802.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2025-02-06
Publication Date
2026-09-04

AI Technical Summary

Benefits of technology

[0014] According to this disclosure, in the process of constructing a candidate list of motion vector predictors in inter-frame prediction mode, by including candidates with multiple motion vectors in the list, it is possible to improve inter-frame prediction performance by using various motion vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122700512A_ABST
    Figure CN122700512A_ABST
Patent Text Reader

Abstract

According to one aspect of the present disclosure, a decoding method includes the steps of obtaining picture information from a bitstream, determining a prediction mode applied to a current block as a prediction mode using a motion vector predictor (MVP) based on the obtained picture information, configuring an MVP candidate list including MVP candidates for the current block, and generating a prediction block of the current block based on at least one MVP candidate within the MVP candidate list, wherein the step of configuring the MVP candidate list includes including an MVP candidate having a base motion vector and an additional motion vector in the MVP candidate list.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method for encoding / decoding image information, a computer-readable storage medium for storing image information, and a method for transmitting image information. Background Technology

[0002] Recently, the demand for high-resolution and high-quality images, such as HD (high-definition) and UHD (ultra-high-definition) images, has been increasing in various application fields, and therefore, efficient image compression technology is being discussed.

[0003] There are various techniques, such as inter-frame prediction techniques that use video compression to predict the pixel values ​​included in the current image from images before or after the current image, intra-frame prediction techniques that use pixel information in the current image to predict the pixel values ​​included in the current image, and entropy coding techniques that assign short symbols to values ​​that occur frequently and long symbols to values ​​that occur infrequently. These image compression techniques can be used to effectively compress image data and transmit or store it.

[0004] Therefore, there is a need for efficient image compression technology to effectively transmit, store, and reproduce information from high-resolution and high-quality images. Summary of the Invention

[0005] [Technical Issues]

[0006] This disclosure provides a method for including candidates with multiple motion vectors in a list during the process of constructing a candidate list of motion vector predictors in an inter-frame prediction mode.

[0007] This disclosure provides a method for improving inter-frame prediction performance by using various motion vectors.

[0008] [Technical Solution]

[0009] According to one aspect of this disclosure, a decoding method includes: acquiring image information from a bitstream; determining a prediction mode applicable to the current block as a prediction mode using a motion vector predictor (MVP) based on the acquired image information; constructing an MVP candidate list including MVP candidates for the current block; and generating a prediction block for the current block based on at least one MVP candidate in the MVP candidate list, wherein constructing the MVP candidate list includes: including MVP candidates with regular motion vectors and additional motion vectors in the MVP candidate list.

[0010] According to one aspect of this disclosure, an encoding method includes: determining a prediction mode applied to a current block as a prediction mode using a motion vector predictor (MVP); constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on the final MVP candidates in the MVP candidate list; and encoding image information including information about the prediction mode, wherein constructing the MVP candidate list includes: including MVP candidates with regular motion vectors and additional motion vectors in the MVP candidate list.

[0011] According to one aspect of this disclosure, in a storage medium storing a bitstream, an encoding method for generating a bitstream includes: determining a prediction mode applied to the current block as a prediction mode using a motion vector predictor (MVP); constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on the final MVP candidates in the MVP candidate list; and encoding image information including information about the prediction mode, wherein constructing the MVP candidate list includes: including MVP candidates with regular motion vectors and additional motion vectors in the MVP candidate list.

[0012] According to one aspect of this disclosure, a transmission method includes: acquiring a bitstream for an image, wherein the bitstream is generated based on the following operations: determining a prediction mode applied to a current block as a prediction mode using a motion vector predictor (MVP); constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on the final MVP candidates in the MVP candidate list; and encoding image information including information about the prediction mode; and transmitting data including the bitstream, wherein constructing the MVP candidate list includes: including MVP candidates with regular motion vectors and additional motion vectors in the MVP candidate list.

[0013] [Beneficial Effects]

[0014] According to this disclosure, in the process of constructing a candidate list of motion vector predictors in inter-frame prediction mode, by including candidates with multiple motion vectors in the list, it is possible to improve inter-frame prediction performance by using various motion vectors.

[0015] The effects available from this disclosure are not limited to those described above, and other effects not mentioned will be readily apparent to those skilled in the art from the following description. Attached Figure Description

[0016] Figure 1 The illustration shows a video / image coding system according to this disclosure.

[0017] Figure 2A schematic block diagram is shown of an encoding apparatus in which embodiments of the present disclosure can be applied and in which video / image signal encoding is performed.

[0018] Figure 3 A schematic block diagram of a decoding apparatus in which embodiments of the present disclosure can be applied and in which video / image signal decoding is performed is shown.

[0019] Figure 4 Examples of video / image decoding methods to which embodiments of the present disclosure can be applied are shown.

[0020] Figure 5 Examples of video / image encoding methods to which embodiments of the present disclosure can be applied are shown.

[0021] Figure 6 This is a diagram illustrating an example of the search region used in intra-frame template matching.

[0022] Figure 7 and Figure 8 Examples of video / image coding methods based on inter-frame prediction that can be applied to embodiments of this disclosure are shown.

[0023] Figure 9 and Figure 10 Examples of video / image decoding methods based on inter-frame prediction that can be applied to embodiments of this disclosure are shown.

[0024] Figure 11 An inter-frame prediction process to which embodiments of the present disclosure can be applied is illustrated by way of example.

[0025] Figure 12 This is a diagram showing an example of a block used to build a list of merge candidates.

[0026] Figure 13 This is a diagram illustrating the four types of motion that can be represented in an affine motion model.

[0027] Figure 14 This is a diagram illustrating an example of the control point motion vectors used in affine motion prediction.

[0028] Figure 15 This is an example illustrating the inheritance of the control point motion vector.

[0029] Figure 16 This is a diagram showing an example of neighboring blocks relative to the current block.

[0030] Figure 17 The process of SbTMVP is shown.

[0031] Figure 18 An example is shown of GPM partitions grouped at the same angle.

[0032] Figure 19 An example is shown of the left and top neighbor blocks used for CIIP weight derivation.

[0033] Figure 20 Available IPM candidates for GPM, including inter-frame prediction and intra-frame prediction, are illustrated exemplarily.

[0034] Figure 21 This is a diagram illustrating the method for deriving motion vectors in template matching.

[0035] Figure 22 This is a diagram illustrating the use of multiple reference blocks in one embodiment.

[0036] Figure 23 This is a flowchart illustrating an example of a method in a decoding or encoding method according to one embodiment for including MVP candidates that include multiple motion vectors in an MVP candidate list.

[0037] Figure 24 This is a diagram illustrating an example of a method for constructing an MVP candidate list in one embodiment.

[0038] Figure 25 This is a diagram illustrating an example of a method for generating candidates with multiple motion vectors according to one embodiment.

[0039] Figure 26 This is a diagram illustrating another example of a method for generating candidates with multiple motion vectors according to one embodiment.

[0040] Figure 27 This shows the distance between the current image and the reference image.

[0041] Figure 28 This is a diagram showing an example of the position of a reference block specified by the motion vector of the current block.

[0042] Figure 29 This is a diagram illustrating an example of an offset that may be applicable in one embodiment.

[0043] Figure 30 This is a diagram illustrating an example of the process of correcting the position by applying an offset within the corresponding block and deriving the motion vector at each corrected position.

[0044] Figure 31 An exemplary content streaming system that can be applied according to embodiments of this disclosure is shown. Detailed Implementation

[0045] Because this disclosure can be modified in various ways and has several embodiments, specific embodiments will be illustrated in the accompanying drawings and described in detail in the detailed description. However, this disclosure is not intended to be limited to the specific embodiments and should be understood to include all variations, equivalents, and substitutions included within the spirit and scope of this disclosure. Similar reference numerals are used for similar components in the description of each drawing.

[0046] Terms such as "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. These terms are used only to distinguish one component from other components. For example, without departing from the scope of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. Terms and / or combinations of any one or more of the relevant statement items are included.

[0047] When a component is described as "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to another component, but there may also be another component in between. On the other hand, when a component is described as "directly connected" or "directly linked" to another component, it should be understood that there is no other component in between.

[0048] The terminology used in this application is for describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, it should be understood that terms such as “comprising” or “having” are intended to designate the presence of features, numbers, steps, operations, components, portions, or combinations thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, portions, or combinations thereof.

[0049] This disclosure relates to video / image coding. For example, the methods / implementations disclosed herein can be applied to methods disclosed in the Universal Video Coding (VVC) standard. Additionally, the methods / implementations disclosed herein can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).

[0050] This specification presents various implementations of video / image encoding, and unless otherwise stated, these implementations may be combined with each other to perform the task.

[0051] Here, "video" can refer to a collection of images over time. "Image" generally refers to a unit representing a picture within a specific time period, and a slice / tile is a unit that forms part of an image during encoding. A slice / tile can include at least one Code Tree Unit (CTU). An image can consist of at least one slice / tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of an image. A tile column is a rectangular area of ​​CTUs with the same height as the image and a width assigned by the syntax requirements of the image parameter set. A tile row is a rectangular area of ​​CTUs with the same height assigned by the image parameter set and a width equal to the width of the image. CTUs within a tile can be arranged consecutively according to a CTU raster scan, and tiles within an image can be arranged consecutively according to a tile raster scan. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows that can be exclusively included within a single NAL unit of an image. Simultaneously, an image can be divided into at least two sub-images. A sub-image can be a rectangular area of ​​at least one slice within an image.

[0052] A pixel, cell, or pixel unit can refer to the smallest unit that makes up a picture (or image). Additionally, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.

[0053] A unit can represent a conventional unit in image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., Cb, Cr) blocks. In some cases, units may be used interchangeably with terms such as block or region. In general, an MxN block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.

[0054] Here, "A or B" can refer to "A only", "B only", or "both A and B". In other words, "A or B" can be interpreted as "A and / or B". For example, "A, B or C" can refer to "A only", "B only", "C only", or "any combination of A, B and C".

[0055] The forward slash ( / ) or comma used in this article can refer to "and / or". For example, "A / B" can refer to "A and / or B". Therefore, "A / B" can refer to "A only", "B only", or "both A and B". For example, "A, B, C" can refer to "A, B, or C".

[0056] Here, "at least one of A and B" can refer to "only A", "only B" or "both A and B". Furthermore, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0057] Additionally, here, "at least one of A, B, and C" can refer to "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can refer to "at least one of A, B, and C".

[0058] Additionally, the parentheses used in this document can refer to "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction". In other words, "prediction" here is not limited to "intra-frame prediction", and "intra-frame prediction" can be cited as an example of "prediction". Furthermore, even when the indication is "prediction (i.e., intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction".

[0059] Here, a technical feature described individually in a single figure may be implemented individually or simultaneously.

[0060] Figure 1 A video / image encoding system according to this disclosure is shown.

[0061] refer to Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device).

[0062] A source device can transmit encoded video / image information or data to a receiving device in the form of a file or stream via a digital storage medium or network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may consist of separate devices or external components.

[0063] A video source can acquire video / images through processes of capturing, synthesizing, or generating video / images. A video source may include means for capturing video / images and means for generating video / images. Means for capturing video / images may include at least one camera, a video / image archive containing previously captured video / images, etc. Means for generating video / images may include a computer, tablet, smartphone, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc., and in this case, the process of capturing video / images can be replaced by a process of generating related data.

[0064] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.

[0065] The transmitting unit can send encoded video / image information or data, output in bitstream form, to the receiving unit of the receiving device via digital storage media or a network in the form of a file or stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiving unit can receive / extract the bitstream and send it to a decoding device.

[0066] Decoding devices can decode video / images by performing a series of processes, such as dequantization, inverse transform, and prediction, that correspond to the operations of encoding devices.

[0067] The renderer can render decoded video / images. The rendered video / images can be displayed through a display unit.

[0068] Figure 2 A rough block diagram of an encoding apparatus that can be applied to embodiments of the present disclosure and perform encoding of video / image signals is shown.

[0069] refer to Figure 2The encoding device 200 may consist of an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.

[0070] Image partitioner 210 can partition an input image (or picture, frame) input to encoding device 200 into at least one processing unit. As an example, a processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively partitioned from coding tree unit (CTU) or maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure.

[0071] For example, a coding unit can be split into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary structure can be applied later. Alternatively, a binary tree structure can be applied before the quadtree structure. The coding process according to this specification can be performed based on a final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, the largest coding unit can be used directly as the final coding unit, or if necessary, the coding unit can be recursively divided into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, as described later.

[0072] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be partitioned or segmented from the aforementioned final encoding unit, respectively. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.

[0073] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an MxN block can represent a set of transform coefficients or samples consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term to correspond a picture (or image) to pixels or cells.

[0074] Encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to converter 232. In this case, the unit in encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called subtractor 231.

[0075] Predictor 220 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a block of predictions including prediction samples for the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction on a block or CU basis. Predictor 220 can generate various information about the prediction, such as prediction mode information, and send it to entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction can be encoded in entropy encoder 240 and output as a bitstream.

[0076] Intra-predictor 222 can predict the current block by referencing samples within the current image. Depending on the prediction mode, the referenced samples can be located near the current block or at a distance from it. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional modes can include 33 or 65 directional modes. However, this is just an example, and more or fewer directional modes can be used depending on the configuration. Intra-predictor 222 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0077] Inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, for skip mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information of the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors of surrounding blocks are used as motion vector predictors, and motion vector differences are signaled to indicate the motion vector of the current block.

[0078] Predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as a combined intra-frame and inter-frame prediction (CIIP) mode. Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for prediction against a block. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games, etc. IBC essentially performs prediction within the current image, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current image. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. A palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, sample values ​​within the image can be signaled based on information about the palette table and palette index. The prediction signal generated by predictor 220 can be used to generate a reconstructed signal or a residual signal.

[0079] Transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graphic when the relationship information between pixels is expressed as a graphic. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. Furthermore, the transform process can be applied to square pixel blocks of the same size or to non-square blocks of variable size.

[0080] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0081] The entropy encoder 240 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can encode information necessary for video / video image reconstruction (e.g., values ​​of syntax elements, etc.) in addition to the transform coefficients quantized together or individually.

[0082] Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the network abstraction layer (NAL) unit level. The video / image information may further include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Additionally, the video / image information may further include general constraint information. Here, information transmitted from the encoding device / signaled to the decoding device and / or syntax elements can be included in the video / image information. The video / image information can be encoded by the above-described encoding process and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmission units (not shown) for transmission and / or storage units (not shown) for storing signals output from the entropy encoder 240 can be configured as internal / external components of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.

[0083] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using dequantizer 234 and inverse transformer 235. Adder 250 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and can also be used for inter-frame prediction of the next image by filtering, which will be described later. Meanwhile, a luminance mapping with chroma scaling (LMCS) can be applied during image encoding and / or reconstruction.

[0084] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various information about the filtering and send it to entropy encoder 240. The information about the filtering can be encoded in entropy encoder 240 and output as a bitstream.

[0085] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied through it, the encoding device can avoid prediction mismatch in encoding device 200 and decoding device, and can also improve encoding efficiency.

[0086] The DPB of memory 270 can store modified reconstructed images for use as reference images in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of blocks in the pre-reconstructed image. The stored motion information can be sent to inter-frame predictor 221 to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 222.

[0087] Image information output from encoding device 200 in bitstream form can be sent to decoding device 300.

[0088] Figure 3 A rough block diagram of a decoding device that can be implemented using embodiments of the present disclosure and perform decoding of video / image signals is shown.

[0089] Image information sent from encoding device 200 in bitstream form can be received by decoding device 300.

[0090] refer to Figure 3 The decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-frame predictor 331 and an inter-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transform 321.

[0091] According to the implementation, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded image buffer (DPB) and can be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.

[0092] When the input includes a bitstream containing video / image information, the decoding device 300 can respond to... Figure 2The process of processing video / image information in an encoding device reconstructs the image. For example, the decoding device 300 can derive units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can perform decoding by using processing units applied in the encoding device. Therefore, the decoding processing unit can be an encoding unit, and the encoding unit can be segmented from encoding tree units or maximally encoded units according to a quadtree structure, binary tree structure, and / or ternary tree structure. At least one transform unit can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding device 300 can be played back by a playback device.

[0093] Decoding device 300 can receive data in bitstream form from... Figure 2The signal output by the encoding device and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may further include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. The information sent / received by the signal and / or the syntax elements described later herein can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of surrounding blocks and the block to be decoded, or information about symbols / bins decoded in the previous step, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins based on the determined context model, and generate symbols corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins for the context model used for the next symbol / bin. Among the information decoded in the entropy decoder 310, information about prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values ​​of entropy decoding performed on them in the entropy decoder 310, i.e., the quantized transform coefficients and related parameter information, can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, information about filtering in the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300 or the receiving unit can be a component of the entropy decoder 310.

[0094] Furthermore, the decoding device according to this specification can be referred to as a video / image / picture decoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0095] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can perform dequantization on the quantized transform coefficients and obtain the transform coefficients by using quantization parameters (e.g., quantization step size information).

[0096] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0097] Predictor 320 can perform prediction on the current block and generate a prediction block including prediction samples for the current block. Predictor 320 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from entropy decoder 310, and determine a specific intra-frame / inter-frame prediction mode.

[0098] Predictor 320 can generate prediction signals based on various prediction methods described later. For example, predictor 320 can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as a combined intra-frame and inter-frame prediction (CIIP) mode. Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for block prediction. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be included in the video / image information and transmitted as a signal.

[0099] Intra-predictor 331 can predict the current block by referencing samples within the current image. Depending on the prediction mode, the referenced samples can be located near the current block or at a certain distance away from the current block. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0100] Inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information of neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction may include information indicating the inter-frame prediction mode used for the current block.

[0101] Adder 340 can add the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331) to generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array). When there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstruction block.

[0102] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, output by filtering as described later, or it can be used for inter-frame prediction of the next image. Meanwhile, a luminance map with chroma scaling (LMCS) can be applied during image decoding.

[0103] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and send the modified reconstructed image to memory 360, specifically the DPB of memory 360. Various filtering methods can include deblocking filtering, adaptive sampling offset, adaptive loop filtering, bilateral filtering, etc.

[0104] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame prediction unit 332. Memory 360 can store motion information of blocks derived (or decoded) from motion information in its current image and / or motion information of blocks in the pre-reconstructed image. The stored motion information can be sent to inter-frame predictor 260 as motion information for spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 331.

[0105] The embodiments described here in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can also be applied equally or correspondingly to the filter 350, inter-frame predictor 332 and intra-frame predictor 331 of the decoding device 300.

[0106] Figure 4 Examples of video / image decoding methods to which embodiments of the present disclosure may be applied are provided.

[0107] In video / image coding, the images that make up a video / image can be encoded / decoded according to a series of decoding sequences. The image order corresponding to the output order of the decoded images can be set to be different from the decoding order, and based on this, not only forward prediction but also backward prediction can be performed in inter-frame prediction.

[0108] exist Figure 4 In this process, S400 can be executed by the entropy decoder 310 of the aforementioned decoding device 300, S410 can be executed by the predictor 330, S420 can be executed by the residual processor 320, S430 can be executed by the adder 340, and S440 can be executed by the filter 350. S400 may include the decoding process according to this disclosure, S410 may include the inter-frame / intra-frame prediction process according to this disclosure, S420 may include the residual processing process according to this disclosure, S430 may include the block / image reconstruction process according to this disclosure, and S440 may include the intra-loop filtering process according to this disclosure.

[0109] refer to Figure 4 The decoding device can acquire video / image information from the bitstream (S400), perform prediction based on the acquired video / image information (S410), and reconstruct the image through residual processing (S420, dequantizing and inverse transforming the quantized transform coefficients) (S430).

[0110] Through the reconstruction process, the in-loop filtering process (S440) can be applied to the reconstructed image generated by the reconstruction process to produce a modified reconstructed image. The modified reconstructed image can be output as a decoded image and can also be stored in the buffer or memory of the decoding device, and used as a reference image in the inter-frame prediction process when decoding the next image. In some cases, the in-loop filtering process can be omitted, and in this case, the reconstructed image can be output as a decoded image, and can also be stored in the buffer or memory of the decoding device and used as a reference image in the inter-frame prediction process when decoding subsequent images.

[0111] The in-loop filtering process (S440) may include a deblocking filtering process, a SAO (Sample Adaptive Offset) process, an ALF (Adaptive Loop Filter) process, and / or a bilateral filtering process, and some or all of them may be omitted. Furthermore, one or more of the deblocking filtering process, the SAO process, the ALF process, and the bilateral filtering process may be applied sequentially, or all of them may be applied sequentially. For example, the SAO process may be performed after the deblocking filtering process has been applied to the reconstructed image. Alternatively, for example, the ALF process may be performed after the deblocking filtering process has been applied to the reconstructed image. The same process may be performed in the encoding device.

[0112] Figure 5 Examples of video / image coding methods to which embodiments of the present disclosure may be applied are provided.

[0113] exist Figure 5 In this process, the prediction step (S500) can be performed by the predictor 220 of the encoding device 200, the residual processing based on the prediction result (S510) can be performed by the residual processor 230, and the step of encoding video information including prediction information and residual information (S520) can be performed by the entropy encoder 240. S500 may include the inter-frame / intra-frame prediction process according to this disclosure, S510 may include the residual processing process according to this disclosure, and S520 may include the encoding process according to this disclosure.

[0114] The encoding process may optionally include not only encoding the information used for image reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it as a bitstream, but also generating a reconstructed image for the current image and applying in-loop filtering to the reconstructed image.

[0115] Encoding device 200 can derive (modified) residual samples from quantized transform coefficients using dequantizer 234 and inverse transformer 235, and can generate a reconstructed image based on the predicted sample as the output of S500 and the (modified) residual samples. The reconstructed image generated in this way can be the same as the reconstructed image generated in decoding device 300 described above. The modified reconstructed image can be generated through an in-loop filtering process for reconstructing the image, and the modified reconstructed image can be stored in a buffer or memory and used as a reference image in the inter-frame prediction process when encoding subsequent images, similar to the case of the decoding device.

[0116] As mentioned above, in some cases, some or all of the in-loop filtering process can be omitted. When performing the in-loop filtering process, the filtering-related information (parameters) can be encoded by the entropy encoder 240 and output as a bit stream, and the decoding device 300 can perform the in-loop filtering process in the same manner as the encoding device based on the filtering-related information.

[0117] This in-loop filtering process reduces noise (such as artifacts and ringing) that occurs during image / video encoding and improves subjective / objective image quality. Furthermore, by performing the in-loop filtering process in both the encoding device 200 and the decoding device 300, both devices can derive the same prediction results, thereby improving the reliability of image encoding and reducing the amount of data that needs to be transmitted for image encoding.

[0118] As described above, the image reconstruction process can be performed not only in the decoding device 300 but also in the encoding device 200. Reconstructed blocks can be generated on a per-block basis based on intra-frame prediction / inter-frame prediction, and a reconstructed image including these blocks can be generated. When the current image / slice / tile group is an I-image / slice / tile group, the blocks included in the current image / slice / tile group can be reconstructed based solely on intra-frame prediction. Meanwhile, when the current image / slice / tile group is a P-image / slice / tile group or a B-image / slice / tile group, the blocks included in the current image / slice / tile group can be reconstructed based on either intra-frame prediction or inter-frame prediction. In this case, inter-frame prediction can be applied to some blocks within the current image / slice / tile group, and intra-frame prediction can be applied to other blocks.

[0119] The color components of an image may include luminance components and chrominance components, and unless expressly limited in this disclosure, embodiments of this disclosure may be applied to both luminance and chrominance components.

[0120] Meanwhile, when performing intra-frame prediction, the predictors 220 and 330 of the encoding device 200 and the decoding device 300 can derive reference samples from the neighboring samples of the current block according to the intra-frame prediction mode of the current block, and can generate prediction samples for the current block based on the reference samples.

[0121] For example, (i) the predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can be derived based on the reference sample among the neighboring reference samples of the current block that is located in a specific (predictive) direction of the predicted sample. Case (i) can be called non-directional mode or non-angular mode, and case (ii) can be called directional mode or angular mode.

[0122] In addition, linear interpolation intra-prediction (LIP) can be applied, where intra-prediction is performed for the current block by linearly interpolating the predicted sample values ​​generated by the intra-prediction mode based on the current block.

[0123] Furthermore, the provisional prediction sample for the current block can be derived based on filtered neighbor reference samples, and the prediction sample for the current block can be derived by performing a weighted sum of at least one reference sample derived according to the intra-prediction mode from the existing neighbor reference samples (i.e., unfiltered neighbor reference samples) with the provisional prediction sample. Such a prediction can be called PDPC (Location Dependent Intra-Prediction Combination).

[0124] Furthermore, intra-frame prediction coding can be performed by selecting the reference sample line with the highest prediction accuracy among multiple neighboring reference sample lines in the current block, deriving the prediction sample using reference samples located in the prediction direction on the selected line, and indicating (signaling) the used reference sample line to the decoding device. This approach can be called multi-reference line intra-frame prediction (MRL) or MRL-based intra-frame prediction.

[0125] Furthermore, intra-frame prediction can be performed using the same intra-frame prediction mode by dividing the current block into vertical or horizontal sub-partitions, and neighbor reference samples can be derived and used on a sub-partition basis. That is, in this case, the intra-frame prediction mode used for the current block is applied equivalently to the sub-partitions, but in some cases, deriving and using neighbor reference samples on a sub-partition basis can improve prediction performance. Such a prediction method can be called intra-frame sub-partitioning (ISP) or ISP-based intra-frame prediction.

[0126] Furthermore, when the prediction direction based on the predicted sample points between neighboring reference samples, that is, when the prediction direction points to the fractional sample location, the value of the predicted sample can be derived by interpolating multiple reference samples located around the prediction direction (around the fractional sample location).

[0127] The MPM list used to derive the above intra-prediction modes can be configured differently depending on the intra-prediction type. Alternatively, the MPM list can be configured uniformly regardless of the intra-prediction type.

[0128] Spatial Geometric Partitioning Model (SGPM)

[0129] SGPM is an intra-frame mode of an inter-frame coding tool similar to GPM, generating two prediction parts during intra-frame prediction. In this mode, a candidate list is constructed, where each entry includes a partition split and two intra-frame prediction modes, and a combination is formed using one partition mode and three intra-frame prediction modes. The length of the candidate list can be set to 16, and the selected candidate index can be signaled.

[0130] The candidate list is reordered using a template, where the SAD (Situational Aspect Difference) between the template's prediction and reconstruction is used for sorting. The template size can be fixed at 1.

[0131] For each partition mode, the same intra-frame to inter-frame GPM list derivation is used to derive the IPM list for each partition. The size of the IPM list can be set to 3. In the list, the TIMD derivation mode can be replaced with two derivation modes in the horizontal and vertical directions.

[0132] The SGPM mode can be applied to the following limited block sizes: 4 ≤ width ≤ 64, 4 ≤ height ≤ 64, width < height × 8, height < width × 8, width × height ≥ 32.

[0133] Adaptive blending is also used in spatial GPM, and the blending depth τ can be derived as follows: - If minimum (width, height) = 4, choose 1 / 2τ.

[0134] - Otherwise, if the minimum (width, height) = 8, choose τ.

[0135] Otherwise, if the minimum (width, height) = 16, choose 2τ.

[0136] Otherwise, if the minimum (width, height) = 32, choose 4τ.

[0137] - Otherwise, choose 8τ.

[0138] Intra-block copy (IBC)

[0139] IBC is a tool used in the HEVC extension of SCC. It is well known that the coding efficiency of screen content material is significantly improved. Since IBC mode is implemented as a block-level coding mode, block matching (BM) can be performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector can be used to represent the displacement from the current block to a reference block that has already been reconstructed within the current image. The luma block vector of an IBC-encoded CU can have an integer resolution (or precision). The chroma block vector can be rounded to an integer resolution. Then, the IBC mode combined with AMVR can switch between 1-pixel and 4-pixel motion vector resolutions. IBC-encoded CUs can be treated as a third prediction mode, rather than an intra-frame or inter-frame prediction mode. IBC mode can be applied to CUs whose width and height are both less than or equal to 64 luma samples.

[0140] On the encoder side, hash-based motion estimation can be performed for IBC. The encoder can perform BD checks on blocks with a width and height of no more than 16 luminance samples. For non-merging modes, a block vector search can be performed first using a hash-based search. When the hash search does not return valid candidates, a local search based on block matching can be performed.

[0141] In hash-based searches, hash key matching (32-bit CRC) between the current block and the reference block can be extended to all allowed block sizes. The hash key calculation for all locations in the current image is based on 4×4 sub-blocks. For a current block with a larger size, a hash key match with the reference block's hash key can be determined when all hash keys of the 4×4 sub-blocks match the hash key of the corresponding reference location.

[0142] When multiple predicted blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the minimum cost can be selected.

[0143] In block matching searches, the search scope can be set to cover both the previous CTU and the current CTU. At the CU level, flags are used to signal the IBC mode, and the IBC mode can be signaled as IBC AMVP mode or IBC skip / merge mode as follows.

[0144] - IBC Skip / Merge Mode: Merge candidate indices can be used to indicate which block vectors from the list of neighboring candidate IBC encoded blocks are used to predict the current block. The merge list can include spatial, HMVP, and paired candidates.

[0145] -IBC AMVP Mode: Block vector differences can be encoded in the same way as motion vector differences. The block vector prediction method can use two candidates (in the case of IBC encoding): a predictor from the left neighbor and a predictor from the top neighbor. When neither neighbor is available, a default block is used as the predictor. A signal flag can be used to indicate the block vector predictor index.

[0146] Intra-frame TMP (Intra-frame Template Matching Prediction)

[0147] Intra-template matching prediction (intra-TMP) is a special intra-prediction mode that copies the best prediction block that matches the L-shaped template of the current template in the reconstruction portion of the current frame. Within a predefined search range, the encoder searches for the template most similar to the current template in the reconstruction region of the current frame and uses the corresponding block as the prediction block. Afterward, the encoder signals the use of this mode, and the decoder performs the same prediction operation.

[0148] Figure 6 This is a diagram illustrating an example of the search region used in intra-frame template matching.

[0149] By using the L-shape of the current block, or by selecting only the top or left region, causal neighbors can be linked. Figure 6 It matches another block within a predefined search region to generate a predicted signal. For example... Figure 6 As shown, there can be a total of 6 predefined search regions (i.e., R1~R6), which include: partial reconstruction samples in the current CTU located above, to the left, below the left, and above the right of the current block, as well as reconstruction samples from the CTU above and to the left: The sum of absolute differences (SAD) is used as the cost function.

[0150] The search order of the given six search regions (i.e., R4, R5, R6, R1, R2, R3) is used. Within each region, the decoder generates a list of up to 19 template-matching block vector candidates, ordered in ascending order according to template cost (SAD). Supported modes are as follows: Single predictor: A single predictor is selected from the candidate list.

[0151] Multi-predictor fusion: The final prediction block is derived by fusing several predictors. The fusion weights can be calculated based on the template matching cost of each predictor, or a weight derivation method based on Wiener filters can be used.

[0152] Subpixel precision: When using a single predictor, 1 / 2 pixel, 1 / 4 pixel and 3 / 4 pixel precision are supported, with each precision providing 8 directions.

[0153] Linear filter model: A linear filter learned between the reference template and the current template is applied to the reference block. This mode can be applied to a single predictor without using sub-pixel precision.

[0154] To ensure a fixed number of SAD comparisons per pixel, the sizes of all search regions (SearchRange_w, SearchRange_h) are set to be proportional to the block size (BlkW, BlkH). That is: SearchRange_w = min(64, a BlkW) SearchRange_h = min(64, a BlkH) Here, 'a' is a constant that adjusts the trade-off between gain and complexity, and can be set to a = 5.

[0155] To accelerate the template matching process, the search range of all search regions can be subsampled by a factor of 3. After finding the best match, a refinement process is performed. This refinement process is performed by conducting a second template matching search around the best match with a reduced range.

[0156] Intra-frame template matching can be activated for CUs with a width and height each less than or equal to 64. The maximum CU size for intra-frame template matching is configurable.

[0157] When DIMD is not used for the current CU, the intra-template matching prediction mode can be signaled at the CU level via a dedicated flag.

[0158] Simultaneously, when applying inter-frame prediction, the predictor of the encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction can be represented as a prediction derived in a way that depends on data elements (e.g., sample values ​​or motion information) of images other than the current image. When applying inter-frame prediction to the current block, the prediction block (prediction sample array) of the current block can be derived based on a reference block (reference sample array) specified by motion vectors on a reference image indexed by a reference image.

[0159] To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information for the current block can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and / or reference image indices. Motion information may also include inter-frame prediction type information (L0 prediction, L1 prediction, Bi prediction, etc.). When applying inter-frame prediction, neighboring blocks may include spatially neighboring blocks in the current image and temporally neighboring blocks in the reference image.

[0160] The reference image including the reference block and the reference image including the temporally neighboring block can be the same or different. The temporally neighboring block can be referred to by names such as juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block can also be referred to as the juxtaposed image (colPic). For example, a candidate list of motion information can be constructed based on the neighboring blocks of the current block, and a signal can be sent to indicate which candidate to select (use) to derive the motion vector and / or reference image index of the current block, along with flags or index information.

[0161] Inter-frame prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block can be the same as the motion information of the selected neighboring blocks. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, the motion vectors of the selected neighboring blocks can be used as motion vector predictors, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0162] Motion information may include L0 motion information and / or L1 motion information depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction may be called the L0 motion vector or MVL0, and the motion vector in the L1 direction may be called the L1 motion vector or MVL1. Prediction based on the L0 motion vector may be called L0 prediction, prediction based on the L1 motion vector may be called L1 prediction, and prediction based on both L0 and L1 motion vectors may be called bi-prediction. Here, the L0 motion vector may represent the motion vector associated with the reference image list L0 (L0), and the L1 motion vector may represent the motion vector associated with the reference image list L1 (L1). The reference image list L0 may include images that are earlier in the output order than the current image as reference images, and the reference image list L1 may include images that are later in the output order than the current image. The earlier image may be called the forward (reference) image, and the later image may be called the backward (reference) image.

[0163] The reference image list L0 can also include images that are output later than the current image. In this case, within the reference image list L0, earlier images can be indexed first, followed by later images. The reference image list L1 can also include images that are output earlier than the current image. In this case, within the reference image list L1, later images can be indexed first, followed by earlier images. Here, the output order can correspond to the POC (Picture Order Count) order.

[0164] Figure 7 and Figure 8 Examples of video / image coding methods based on inter-frame prediction that can be applied to embodiments of this disclosure are shown.

[0165] refer to Figure 7 The encoding device 200 can perform inter-frame prediction (S600) on the current block. The encoding device can deduce the inter-frame prediction mode and motion information of the current block, and can generate prediction samples for the current block. Here, the processes of determining the inter-frame prediction mode, deduce motion information, and generate prediction samples can be performed simultaneously, or one process can be performed before the other. For example, as... Figure 8 As shown, the inter-frame predictor 221 of the encoding device 200 may include a prediction mode determiner 221a, a motion information deducer 221b, and a prediction sample deducer 221c, wherein the prediction mode determiner 221a determines the prediction mode of the current block, the motion information deducer 221b derives the motion information of the current block, and the prediction sample deducer 221c derives the prediction samples of the current block.

[0166] For example, the inter-frame predictor of the encoding device can search for blocks similar to the current block within a specific region (search region) of the reference image through motion estimation, and can deduce the reference block whose difference from the current block is minimal, equal to, or less than a specific criterion. Based on this, a reference image index indicating the reference image in which the reference block is located can be derived, and motion vectors can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode to apply to the current block among various prediction modes. The encoding device can compare the RD costs of various prediction modes and determine the optimal prediction mode for the current block.

[0167] For example, when applying a skip mode or merge mode to the current block, the encoding device can construct a merge candidate list, described later, and deduce, among the reference blocks specified by the merge candidates included in the merge candidate list, the reference block whose difference (i.e., the difference in SAD or SATD) with the current block's samples is the smallest, equal to, or less than a specific criterion. In this case, a merge candidate associated with the deduced reference block is selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be deduced using the motion information of the selected merge candidate.

[0168] As another example, when applying (A)MVP mode to the current block, the encoding device can construct an (A)MVP candidate list, described later, and use the motion vector of the selected MVP candidate from the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the motion vector predictor for the current block. In this case, for example, the motion vector of the reference block derived through the motion estimation described above can be used as the motion vector of the current block, and the motion vector predictor candidate with the motion vector having the smallest difference from the motion vector of the current block can be the selected motion vector predictor candidate. MVD (Motion Vector Difference) can be derived, which is the difference obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, information about MVD can be signaled to the decoding device. Furthermore, when applying (A)MVP mode, the value of the reference image index can be configured as reference image index information and signaled separately to the decoding device.

[0169] The encoding device can derive residual samples based on the predicted samples (S610). The encoding device can derive residual samples by comparing the original samples of the current block with the predicted samples.

[0170] The encoding device can encode image information including prediction information and residual information (S620). The encoding device can output the encoded image information in the form of a bitstream. The prediction information is information about prediction and may include prediction mode information (e.g., skip flag, merge flag, merge index, etc.) and / or motion information. The motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index), which is information used to derive motion vectors. In addition, the motion information may include information about the aforementioned MVD and / or reference image index information. Furthermore, the motion information may include information indicating whether L0 prediction, L1 prediction, or bidirectional prediction is applied. The residual information is information about residual samples. The residual information may include information about the quantization transform coefficients of the residual samples.

[0171] The output bitstream can be stored in (digital) storage media and transmitted to the decoding device, or it can be transmitted to the decoding device via a network.

[0172] Simultaneously, as mentioned above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is because the aim is to derive the same prediction result at the encoding device as the prediction result performed at the decoding device, thereby improving encoding efficiency. Therefore, the encoding device can store the reconstructed image (or reconstructed samples, reconstructed blocks) in memory and use it as a reference image for inter-frame prediction. As mentioned above, in-loop filtering processes, etc., can be further applied to the reconstructed image.

[0173] Figure 9 and Figure 10 Examples of video / image decoding methods based on inter-frame prediction that can be applied to embodiments of this disclosure are shown.

[0174] The video / image decoding process based on inter-frame prediction can schematically include, for example, the following.

[0175] refer to Figure 9 and Figure 10 The decoding device 300 can perform operations corresponding to those performed by the encoding device 200. The decoding device can perform predictions on the current block based on the received prediction information and derive prediction samples.

[0176] Specifically, the decoding device can determine the prediction mode of the current block based on the received prediction information (S700). The prediction mode determiner 332a of the decoding device 300 can determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.

[0177] For example, a merge flag can be used to determine whether to apply a merge mode or (A)MVP mode to the current block. Alternatively, an inter-frame prediction mode candidate can be selected from a variety of inter-frame prediction mode candidates based on a mode index. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter-frame prediction modes described later.

[0178] The decoding device can deduce motion information for the current block based on a determined inter-frame prediction mode (S710). For example, when a skip mode or merge mode is applied to the current block, the motion information derivator 332b of the decoding device 300 can construct a merge candidate list, described later, and select a merge candidate from among the merge candidates included in the merge candidate list. Such selection can be performed based on the selection information (merge index) described above. The motion information of the current block can be deduced using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0179] As another example, when applying (A)MVP mode to the current block, the decoding device can construct an (A)MVP candidate list, described later, and can use the motion vector of the selected MVP (Motion Vector Predictor) candidate from the MVP (Motion Vector Predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. Such selection can be performed based on the selection information mentioned above (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the MVP and MVD of the current block. Furthermore, the reference image index of the current block can be derived based on reference image index information. The image specified by the reference image index within the reference image list of the current block can be derived as the reference image referenced for inter-frame prediction of the current block.

[0180] Meanwhile, the motion information of the current block can be derived without constructing a candidate list. In this case, the motion information of the current block can be derived based on the process disclosed in the prediction model. In this case, the aforementioned candidate list construction process can be omitted.

[0181] The decoding device can generate prediction samples for the current block based on the motion information of the current block (S720). In this case, the prediction sample derivator 332c of the decoding device 300 can derive a reference image based on the reference image index of the current block, and derive prediction samples for the current block using samples of the reference block specified by the motion vector of the current block on the reference image. In this case, as described later, depending on the circumstances, a prediction sample filtering process can be further performed on all or part of the prediction samples of the current block.

[0182] In other words, the inter-frame predictor 332 of the decoding device 300 may include a prediction mode determiner 332a, a motion information inferrer 332b, and a prediction sample inferrer 332c, wherein the prediction mode determiner 332a determines the prediction mode of the current block based on the received prediction mode information, the motion information inferrer 332b infers the motion information (motion vector and / or reference image index, etc.) of the current block based on the received information about motion information, and the prediction sample inferrer 332c infers or generates prediction samples of the current block.

[0183] The decoding device generates residual samples for the current block based on the received residual information (S730). The decoding device 300 can generate reconstructed samples for the current block based on the predicted samples and residual samples, and can generate a reconstructed image based on this (S740). Thereafter, as described above, in-loop filtering processes, etc., can be further performed on the reconstructed image.

[0184] Figure 11An inter-frame prediction process to which embodiments of the present disclosure can be applied is illustrated by way of example.

[0185] refer to Figure 11 As described above, the inter-frame prediction process (S600) may include the steps of determining an inter-frame prediction mode, deriving motion information based on the determined prediction mode, and performing prediction (generating prediction samples) based on the derived motion information. The above-described inter-frame prediction process may be performed by an encoding apparatus and a decoding apparatus. In this document, the encoding apparatus may include an encoding apparatus and / or a decoding apparatus.

[0186] refer to Figure 11 The encoding device determines the inter-frame prediction mode for the current block (S800). Various inter-frame prediction modes can be used for the prediction of the current block in the image. For example, modes such as merge mode, skip mode, MVP (Motion Vector Prediction) mode, affine mode, sub-block merge mode, and MMVD (Merge with MVD) mode can be used. DMVR (Decoder-Side Motion Vector Refinement) mode, AMVR (Adaptive Motion Vector Resolution) mode, CU-level Weighted Bi Prediction (BCW), and Bi-Directional Optical Flow (BDOF) can be used as auxiliary modes. Furthermore, according to one embodiment, the above-mentioned inter-frame prediction modes may include a multi-hypothesis prediction (MHP) mode. The multi-hypothesis prediction mode represents a method of performing prediction by applying a weighted sum of additional prediction blocks generated based on additional motion information to the inter-frame prediction block. The multi-hypothesis prediction mode will be described in detail later.

[0187] In this disclosure, the affine pattern can also be referred to as the affine motion prediction pattern. Furthermore, the MVP pattern can also be referred to as the AMVP (Advanced Motion Vector Prediction) pattern. In this disclosure, certain patterns and / or motion information candidates derived from certain patterns can be included as one of the motion information related candidates for another pattern. For example, an HMVP candidate can be added as a merge candidate for a merge / skip pattern, or it can be added as a motion vector predictor candidate for an AMVP pattern. When an HMVP candidate is used as a motion information candidate for a merge or skip pattern, the HMVP candidate can be referred to as an HMVP merge candidate.

[0188] Prediction mode information, indicating the inter-frame prediction mode of the current block, can be signaled from the encoding device to the decoding device. This prediction mode information can be included in the bitstream and received at the decoding device. The prediction mode information may include index information indicating one of several candidate modes. Alternatively, the inter-frame prediction mode can be specified via hierarchical signaling of flag information.

[0189] In this context, the predictive pattern information may include one or more flags. For example, a skip flag may be signaled to indicate whether a skip pattern is applied, and when a skip pattern is not applied, a merge flag may be signaled to indicate whether a merge pattern is applied, and when a merge pattern is not applied, the MVP pattern may be applied, or further flags for further differentiation may be signaled. Affine patterns may be signaled as independent patterns or as patterns dependent on merge or MVP patterns, etc. For example, affine patterns may include affine merge patterns and affine MVP patterns.

[0190] The encoding device can derive motion information for the current block (S810). The motion information can be derived based on the inter-frame prediction mode determined in the above steps. The encoding device can use the motion information of the current block to perform inter-frame prediction. The encoding device can derive optimal motion information for the current block through a motion estimation process.

[0191] For example, the encoding device can search for highly correlated similar reference blocks within a predetermined search range of the reference image, using original blocks in the original image of the current block, on a fractional pixel basis, and can derive motion information from this. Block similarity can be derived based on the difference in sample values ​​based on phase. For example, block similarity can be calculated based on the SAD (Short-Average Difference) between the current block (or its template) and the reference block (or its template). In this case, motion information can be derived based on the reference block with the minimum SAD within the search area. The derived motion information can then be signaled to the decoding device based on inter-frame prediction modes using various methods.

[0192] The encoding device can perform inter-frame prediction based on the motion information of the current block to generate prediction samples (S820). The current block that includes the prediction samples can be called the prediction block.

[0193] Simultaneously, a signal can be sent indicating whether to apply the aforementioned List 0 (L0) prediction, List 1 (L1) prediction, or bidirectional prediction to the current block (current coding unit). Such information can be referred to as motion prediction direction information, inter-frame prediction direction information, or inter-frame prediction indication information, and can be configured / encoded / signaled, for example, in the form of the `inter_pred_idc` syntax element. That is, the `inter_pred_idc` syntax element can specify whether to apply the aforementioned List 0 (L0) prediction, List 1 (L1) prediction, or bidirectional prediction to the current block (current coding unit). In this document, for ease of description, the inter-frame prediction type (L0 prediction, L1 prediction, or BI prediction) specified by the `inter_pred_idc` syntax element can be referred to as the motion prediction direction. L0 prediction can be represented as `pred_L0`, L1 prediction as `pred_L1`, and bidirectional prediction as `pred_BI`. For example, depending on the value of the `inter_pred_idc` syntax element, the prediction type can be specified as shown in Table 1 below.

[0194] [Table 1]

[0195] As described above, an image can include one or more slices. Slices can have one of the following slice types: intra-frame (I) slices, prediction (P) slices, and bidirectional prediction (B) slices. Such slice types can be specified based on slice type information. For blocks within an I slice, inter-frame prediction is not used for prediction, and only intra-frame prediction can be used. Of course, even in this case, the raw sample values ​​can be encoded and signaled without prediction. For blocks within a P slice, either intra-frame or inter-frame prediction can be used, and when inter-frame prediction is used, only unidirectional prediction can be used. Meanwhile, for blocks within a B slice, either intra-frame or inter-frame prediction can be used, and when inter-frame prediction is used, at most bidirectional prediction can be used.

[0196] L0 and L1 can include reference images encoded / decoded earlier than the current image. For example, L0 can include reference images that are earlier and / or later than the current image in the POC order, and L1 can include reference images that are later and / or earlier than the current image in the POC order. In this case, for reference images that are earlier than the current image in the POC order, L0 can be assigned a relatively lower reference image index, and for reference images that are later than the current image in the POC order, L1 can be assigned a relatively lower reference image index. For B-slices, bidirectional prediction can be applied, and in this case, either unidirectional bidirectional prediction or two-way bidirectional prediction can be applied. Two-way bidirectional prediction can be referred to as true bidirectional prediction.

[0197] Inter-frame prediction can be performed using motion information from the current block. The encoding device can derive optimal motion information for the current block through a motion estimation process. For example, the encoding device can search for highly correlated similar reference blocks within a predetermined search range of the reference image, using original blocks in the original image of the current block, on a fractional pixel basis, and can derive motion information from this. Block similarity can be derived based on the difference in phase-based sample values. For example, block similarity can be calculated based on the SAD between the current block (or its template) and a reference block (or its template). In this case, motion information can be derived based on the reference block with the minimum SAD within the search area. The derived motion information can be signaled to the decoding device based on the inter-frame prediction mode according to various methods.

[0198] Merge mode and skip mode

[0199] When the merge mode is applied, the motion information of the current prediction block is not directly transmitted; instead, it is derived from the motion information of neighboring prediction blocks. Therefore, the motion information of the current prediction block can be specified by transmitting a flag indicating that the merge mode was used and a merge index indicating which neighboring prediction block was used. The merge mode can also be called the regular merge mode.

[0200] To execute the merging mode, the encoder can search for candidate blocks to merge, used to derive motion information for the current predicted block. For example, up to five candidate blocks can be used, but the invention is not limited to this. Furthermore, the maximum number of candidate blocks can be transmitted in the slice header or tile group header, but the invention is not limited to this. After finding candidate blocks, the encoder can generate a list of candidate blocks and select the one with the lowest cost from these as the final candidate block to merge.

[0201] Figure 12 This is a diagram showing an example of a block used to build a list of merge candidates.

[0202] This invention provides various embodiments regarding merge candidate blocks that constitute a merge candidate list.

[0203] For example, the merge candidate list can use 5 merge candidate blocks. For instance, it can use 4 spatial merge candidates and 1 temporal merge candidate. As a specific example, for spatial merge candidates, it can use... Figure 12 The blocks shown are spatial merge candidates. In the following text, the spatial merge candidates or spatial MVP candidates described later may be referred to as SMVPs, and the temporal merge candidates or temporal MVP candidates described later may be referred to as TMVPs.

[0204] For example, the list of candidate merges for the current block can be constructed based on the following process: - Insert spatial merge candidates derived by searching for spatial neighboring blocks into the merge candidate list; - Insert time merge candidates derived by searching for time neighboring blocks into the merge candidate list; - Compare the current number of merge candidates with the maximum number of merge candidates; - If the current number of merge candidates is less than the maximum number of merge candidates, then additional merge candidates will be inserted into the merge candidate list.

[0205] To describe the process in detail above, the encoding device (encoding / decoding device) inserts spatial merging candidates derived by searching for spatial neighboring blocks of the current block into the merging candidate list. For example, spatial neighboring blocks may include the neighboring block at the lower left corner, the left neighboring block, the upper right neighboring block, the upper neighboring block, and the upper left neighboring block of the current block. However, this is just an example, and in addition to the aforementioned spatial neighboring blocks, additional neighboring blocks such as the right neighboring block, the lower neighboring block, and the lower right neighboring block can be used as spatial neighboring blocks. The encoding device can search for spatial neighboring blocks based on priority to detect available blocks, and can derive spatial merging candidates from the motion information of the detected blocks. For example, the encoder and decoder can search in the order A1, B1, B0, A0, B2. Figure 12 The five blocks shown can be indexed sequentially to build a list of merged candidates.

[0206] The encoding device inserts temporal merge candidates derived by searching temporally neighboring blocks of the current block into the merge candidate list. Temporally neighboring blocks can be located on a reference picture, which is a different picture from the current picture containing the current block. The reference picture containing the temporally neighboring block can be called the juxtaposition picture or col picture. Temporally neighboring blocks can be searched in the order of the neighboring block at the bottom right corner of the juxtaposition block of the current block on the juxtaposition picture, and then the bottom right center block.

[0207] Meanwhile, when motion data compression is applied, in a collocated picture, specific motion information may be stored as representative motion information for each predetermined storage unit. In this case, there is no need to store motion information of all blocks within the predetermined storage unit, whereby a motion data compression effect can be obtained. In this case, the predetermined storage unit may for example be predetermined as a 16x16 sample unit or an 8x8 sample unit, or information on the size of the predetermined storage unit may be signaled from an encoder to a decoder. When motion data compression is applied, motion information of a temporal neighboring block may be replaced with representative motion information of the predetermined storage unit in which the temporal neighboring block is located. That is, in this case, from an implementation perspective, the temporal merge candidate may be derived based on motion information of a prediction block that covers a position obtained by first arithmetically shifting the coordinate (top-left sample position) of the temporal neighboring block to the right by a specific value and then arithmetically shifting to the left by said specific value, instead of a prediction block located at the coordinate of the temporal neighboring block. For example, when the predetermined storage unit is 2 n x 2 n sample unit, and the coordinate of the temporal neighboring block is (xTnb, yTnb), motion information of the prediction block located at the modified position ((xTnb>>n)<<n, (yTnb>>n)<<n) may be used for the temporal merge candidate. As a specific example, when the predetermined storage unit is a 16x16 sample unit, and the coordinate of the temporal neighboring block is (xTnb, yTnb), motion information of the prediction block located at the modified position ((xTnb>>4)<<4, (yTnb>>4)<<4) may be used for the temporal merge candidate. Alternatively, for example, when the predetermined storage unit is an 8x8 sample unit, and the coordinate of the temporal neighboring block is (xTnb, yTnb), motion information of the prediction block located at the modified position ((xTnb>>3)<<3, (yTnb>>3)<<3) may be used for the temporal merge candidate.

[0208] An encoding device may compare the number of current merge candidates with the maximum number of merge candidates. The maximum number of merge candidates may be predefined, or may be signaled from the encoder to the decoder. For example, the encoder may generate information on the maximum number of merge candidates, encode the same, and transmit it to the decoder in the form of a bitstream. When the maximum number of merge candidates is fully filled, the subsequent candidate adding process may not be continued.

[0209] As a result of the comparison, if the current number of merge candidates is less than the maximum number of merge candidates, the encoding device inserts additional merge candidates into the merge candidate list. These additional merge candidates may include, for example, at least one of the following described later: history-based merge candidates, pairwise average merge candidates, ATMVP, combined bidirectional prediction merge candidates (when the current slice / tile group's slice / tile group type is type B), and / or zero-vector merge candidates.

[0210] As a result of the comparison, if the current number of merge candidates is not less than the maximum number of merge candidates, the encoding device can terminate the construction of the merge candidate list. In this case, the encoder can select the optimal merge candidate from the merge candidates constituting the merge candidate list based on the RD (rate-distortion) cost, and can signal the selection information (e.g., merge index) indicating the selected merge candidate to the decoder. The decoder can then select the optimal merge candidate based on the merge candidate list and the selection information.

[0211] The motion information of the selected merging candidate can be used as the motion information of the current block, and as described above, the predicted sample of the current block can be derived based on the motion information of the current block. The encoder can derive the residual sample of the current block based on the predicted sample, and can encode the residual information about the residual sample and transmit it to the decoder. As described above, the decoder can generate reconstructed samples based on the residual samples derived based on the transmitted residual information and the predicted samples, and can generate a reconstructed image based on this.

[0212] When the skip mode is applied, the motion information of the current block can be derived in the same way as when the merge mode is applied, as described above. However, when the skip mode is applied, the residual signal of the corresponding block is omitted, and therefore the predicted sample can be directly used as the reconstructed sample.

[0213] Historical Merger Candidate Derivation

[0214] Following spatial and temporal merge candidates, history-based MVP (HMVP) merge candidates can be added to the merge list. In this approach, motion information from previous coded blocks is stored in a table and used as the MVP for the current coding unit (CU). A table containing multiple HMVP candidates is maintained during the encoding and decoding process. This table is initialized (cleared) when a new coding tree unit (CTU) row begins. Whenever a non-sub-block inter-coded CU exists, the associated motion information is added as a new HMVP candidate to the last entry in the table.

[0215] In VVC, the size S of the HMVP table is set to 5, indicating that up to 5 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a restricted first-in-first-out (FIFO) rule is applied, and a redundancy check is performed first to determine if the same HMVP already exists in the table. If the same HMVP is found, it is removed from the table, and all subsequent HMVP candidates are moved forward.

[0216] HMVP candidates can be used in the merge candidate list construction process. The latest multiple HMVP candidates in the table are checked sequentially and inserted after the TMVP candidates in the merge candidate list. Redundancy checks are also performed on HMVP candidates for spatial or temporal merge candidates.

[0217] The following simplifications are introduced to reduce the number of redundant check operations: If (N ≤ 4), the number of HMVP candidates used to generate the merge list is set to M; otherwise, it is set to (8 - N), where N represents the number of existing candidates in the merge list and M represents the number of available HMVP candidates in the table.

[0218] Once the total number of available merge candidates reaches the maximum allowed number of merge candidates minus 1, the merge candidate list building process using HMVP is terminated.

[0219] Derivation of Pairwise Average Merge Candidates

[0220] In this document, the pairwise average merge candidate may be referred to as the pairwise average candidate or the pairwise candidate. The pairwise average candidate is generated by averaging predefined pairs of candidates within an existing merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}. Here, the numbers represent the merge indices within the merge candidate list.

[0221] The average motion vector is calculated separately for each reference list. When both motion vectors are available in a reference list, the average is calculated even if they point to different reference images. If only one motion vector is available, that vector is used directly. If no motion vector is available, the corresponding list remains invalid.

[0222] If the merge list is still not full even after adding pairwise average merge candidates, insert a zero MVP at the end of the list until the maximum number of merge candidates is reached.

[0223] Merge pattern with MVD (MMVD)

[0224] The aforementioned MMVD modes are described below. Besides the merging mode that directly uses implicitly derived motion information for the predicted samples of the current block, a merging mode utilizing motion vector differences (MMVD) can also be used. Since the skip mode uses a similar motion information derivation scheme as the merging mode, MMVD can also be applied to the skip mode. After signaling the skip flag and the merging flag, MMVD flag information (e.g., mmvd_flag) can be signaled to specify whether to use the MMVD mode for the current block.

[0225] In MMVD mode, after selecting a merge candidate, that candidate can be further refined based on the MVD information signaled. When applying MMVD to the current block (i.e., when mmvd_flag is 1), additional information for MMVD can be signaled. This additional information may include: a merge candidate flag (e.g., mmvd_merge_flag) indicating whether the first or second candidate in the merge candidate list should be used with the motion vector difference; a distance index (e.g., mmvd_distance_idx) indicating the magnitude of the motion; and a direction index (e.g., mmvd_direction_idx) indicating the direction of the motion. In MMVD mode, one of the first two candidates in the merge list can be selected as the MV basis. The merge candidate flag is signaled to specify the candidate used.

[0226] The distance index specifies motion amplitude information and a predefined offset from the starting point. This offset is added to the horizontal or vertical component of the initial MV. The relationship between the distance index and the predefined offset is shown in Table 2 below.

[0227] [Table 2]

[0228] Here, if slice_fpel_mmvd_enabled_flag is 1, it specifies that the merge mode with motion vector difference uses integer sample precision in the current slice. If slice_fpel_mmvd_enabled_flag is 0, it specifies that the merge mode with motion vector difference can use fractional sample precision in the current slice. The slice_fpel_mmvd_enabled_flag syntax element can be notified via the slice hairline signal or can be included in the slice header.

[0229] The direction index specifies the direction of the MVD relative to the starting point. As shown in Table 3 below, the direction index can indicate one of four directions. The meaning of the MVD symbol may vary depending on the information of the starting MV. When the starting MV is a unidirectional prediction MV, or a bidirectional prediction MV where both lists point to the same side of the current image (i.e., when both reference POCs are greater than or less than the current image's POC), the symbols in Table 3 below can specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV, and the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), the symbols in Table 3 below indicate the sign of the MV offset added to the list 0 MV components of the starting MV, while the signs for the list 1 MVs have the opposite value.

[0230] [Table 3]

[0231] The two components of the merged additional MVD offset MmvdOffset[x0][y0] can be derived as follows: MmvdOffset[x0][y0][0] = (MmvdDistance[x0][y0] << 2) MmvdSign[x0][y0][0] MmvdOffset[x0][y0][1] = (MmvdDistance[x0][y0] << 2) MmvdSign[x0][y0][1]

[0232] MVP (Motion Vector Prediction)

[0233] The MVP (Motion Vector Prediction) mode can also be called the AMVP (Advanced Motion Vector Prediction) mode. When applying the MVP mode, motion vector predictor (MVP) candidate lists can be generated using the motion vectors of reconstructed spatial neighbor blocks and / or motion vectors corresponding to temporal neighbor blocks (or colblocks). That is, the motion vectors of reconstructed spatial neighbor blocks and / or motion vectors corresponding to temporal neighbor blocks can be used as motion vector predictor candidates. When applying bidirectional prediction, MVP candidate lists for deriving L0 motion information and MVP candidate lists for deriving L1 motion information can be generated and used separately.

[0234] The aforementioned prediction information (or information about prediction) may include selection information (e.g., an MVP flag or MVP index) that indicates the optimal motion vector predictor candidate to be selected from among the motion vector predictor candidates included in the list. In this case, the predictor of the decoding device can use this selection information to select the motion vector predictor for the current block from among the motion vector predictor candidates included in the motion vector candidate list.

[0235] The predictor of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can encode it and output it as a bitstream. That is, the MVD can be obtained as the value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this point, the predictor of the decoding device can obtain the motion vector difference included in the information about the prediction, and deduce the motion vector of the current block by adding the motion vector difference to the motion vector predictor. The predictor of the decoding device can obtain or deduce reference image indices, etc., from the information about the prediction. For example, the candidate list of motion vector predictors can be constructed as follows: - Search for spatial candidate blocks for motion vector prediction and insert them into the prediction candidate list. - Check if the number of space candidate blocks is less than 2 - If the number of spatial candidate blocks is less than 2, search for temporal candidate blocks and additionally insert them into the prediction candidate list. - If a time candidate block is unavailable, use the zero motion vector; - If the number of spatial candidate blocks is not less than 2, then terminate the construction of the motion vector predictor candidate list.

[0236] Furthermore, when applying the MVP pattern, the reference image indexes can be explicitly signaled. In this case, the reference image indexes used for L0 prediction (refidxL0) and L1 prediction (refidxL1) can be signaled separately. For example, when applying the MVP pattern and applying bidirectional prediction (BI prediction), information about refidxL0 and information about refidxL1 can be signaled separately.

[0237] MVD (Motion Vector Difference) Coding

[0238] When applying the MVP pattern, as described above, information about the MVD derived at the encoding device can be signaled or encoded and transmitted to the decoding device. This MVD information may include, for example, information indicating the x and y components of the MVD's absolute value and sign. In this case, the following information can be signaled in stages: whether the absolute value of the MVD is greater than 0, whether it is greater than 1, and information indicating the remainder of the MVD. For example, information indicating whether the absolute value of the MVD is greater than 1 can only be signaled if the flag indicating whether the absolute value of the MVD is greater than 0 is 1.

[0239] For example, information about MVD can be configured using the following syntax, encoded at the encoding device, and then signaled to the decoding device.

[0240] [Table 4]

[0241] For example, MVD[compIdx] can be based on abs_mvd_greater0_flag[compIdx] (abs_mvd_minus2[compIdx] + 2) (1 - 2 The value can be derived using `mvd_sign_flag[compIdx]`. Here, `compIdx` (or `cpIdx`) represents the index of each component and can have a value of 0 or 1. A `compIdx` value of 0 represents the x component, and a `compIdx` value of 1 represents the y component. However, this is just an example, and the value of each component can be represented using coordinate systems other than the x and y coordinate systems.

[0242] Simultaneously, the MVD (MVDL0) used for L0 prediction and the MVD (MVDL1) used for L1 prediction can be signaled separately, and the information about the MVD can include information about MVDL0 and / or information about MVDL1. For example, when applying the MVP pattern and BI prediction to the current block, both information about MVDL0 and information about MVDL1 can be signaled.

[0243] Symmetric MVD

[0244] Simultaneously, when applying BI prediction, symmetric MVD can be used to consider coding efficiency. In this case, the signal notification of some motion information can be omitted. For example, when applying symmetric MVD to the current block, information about refidxL0, refidxL1, and MVDL1 can be derived internally instead of being signaled from the encoding device to the decoding device. For example, when applying MVP mode and BI prediction to the current block, a signal indicating whether symmetric MVD is applied can be sent (e.g., symmetric MVD flag information or the sym_mvd_flag syntax element), and when the value of this flag information is 1, the decoding device can determine that symmetric MVD has been applied to the current block.

[0245] When applying the symmetric MVD mode (i.e., when the value of the symmetric MVD flag information is 1), the `mvp_l0_flag`, `mvp_l1_flag`, and information about `MVDL0` can be explicitly signaled. As mentioned above, the signaling of information about `refidxL0`, `refidxL1`, and `MVDL1` can be omitted and derived internally. For example, `refidxL0` can be derived as the index of the preceding reference image in reference image list 0 (which may be called list 0 or L0) that is closest to the current image in the POC order. `refidxL1` can be derived as the index of the following reference image in reference image list 1 (which may be called list 1 or L1) that is closest to the current image in the POC order. Alternatively, for example, both `refidxL0` and `refidxL1` can be derived as 0. Or, for example, `refidxL0` and `refidxL1` can each be derived as the smallest index with the same POC difference relative to the current image. As a specific example, when the difference between [the POC of the current image] and [the POC of the first reference image specified by refidxL0] is called the first POC difference, and [the POC of the second reference image specified by refidxL1] is called the second POC difference, the value of refidxL0 indicating the first reference image can be derived as the value of refidxL0 of the current block only if the first POC difference and the second POC difference are the same, and the value of refidxL1 indicating the second reference image can be derived as the value of refidxL1 of the current block. Furthermore, for example, when there are multiple sets where the first POC difference and the second POC difference are the same, the refidxL0 and refidxL1 of the set with the smallest difference among them can be derived as the refidxL0 and refidxL1 of the current block.

[0246] MVDL1 can be derived as -MVDL0. For example, the final MV of the current block can be derived as shown in Equation 1 below.

[0247] [Formula 1]

[0248] Affine prediction

[0249] Existing video coding systems use only one motion vector (using a translational motion model) to represent the motion of a coded block. This method can represent optimal motion at the block level, but since it is not the optimal motion for every actual pixel, determining the optimal motion vector at the pixel level would improve coding efficiency. Therefore, this paper introduces an affine motion prediction method that uses an affine motion model for coding.

[0250] Figure 13 This is a diagram showing the four types of motion that can be represented in an affine motion model.

[0251] Affine motion prediction methods can use 2, 3, or 4 motion vectors to represent the motion vector of each pixel unit within a block.

[0252] like Figure 13 As shown, an affine motion model can represent four types of motion. Among the motions that can be represented by an affine motion model, an affine motion model representing three types of motion (translation, scaling, and rotation) is called a similar (or simplified) affine motion model, and the proposed method described below is based on a similar affine motion model. However, embodiments of this disclosure are not limited to this motion model.

[0253] Figure 14 This is a diagram illustrating an example of the control point motion vectors used in affine motion prediction.

[0254] like Figure 14 As shown, affine motion prediction can use two or more control point motion vectors (CPMVs) to determine the motion vectors of the pixel positions included in a block. This set of motion vectors is then called the affine motion vector field (MVF) and can be determined by the following formula.

[0255] For a 4-parameter affine motion model, the motion vector at the sample position (x,y) of the block can be derived using Equation 2.

[0256] [Equation 2]

[0257] For a 6-parameter affine motion model, the motion vector at the sample position (x,y) of the block can be derived using Equation 3.

[0258] [Formula 3]

[0259] Here, The CPMV is the CP at the top left corner of the coded block. The CPMV is the CP at the top right corner, and This is the CPMV of the block at the bottom left corner. W corresponds to the width of the current block, and H corresponds to the height of the current block. Let be the motion vector at position {x, y}.

[0260] During encoding / decoding, the affine motion vector field (MVF) can be determined either in pixels or in predefined sub-blocks. When determined in pixels, the motion vector is obtained based on each pixel value. When determined in sub-blocks, the motion vector for the corresponding block is obtained based on the pixel value at the center of that sub-block (the lower right of the center, i.e., the lower right sample among the four center samples). As an example, such as... Figure 14 As shown, the affine MVF can be in 4 The size is determined in units of 4 sub-blocks. However, this is just an example, and the sub-block size applied to affine prediction can be modified in various ways.

[0261] When affine prediction is available, the motion model applicable to the current block can include three models: translational motion model, 4-parameter affine motion model, and 6-parameter affine motion model. Here, the translational motion model can represent a model using existing block element motion vectors, the 4-parameter affine motion model can represent a model using two CPMVs, and the 6-parameter affine motion model can represent a model using three CPMVs.

[0262] Affine motion prediction can include affine MVP (or affine inter-frame) mode and affine merging mode. In affine motion prediction, the motion vector of the current block can be derived on a sample-by-sample or sub-block-by-sub-block basis.

[0263] Affine merging

[0264] In affine merging mode, the CPMV can be determined based on the affine motion model of neighboring blocks encoded according to affine motion prediction. In terms of search order, affine-coded neighboring blocks can be used in affine merging mode. When one or more neighboring blocks are encoded using affine motion prediction, the current block can be encoded using affine merging.

[0265] In other words, when applying the affine merge mode, the CPMV of the neighboring blocks can be used to derive the CPMV of the current block. In this case, the CPMV of the neighboring blocks can be used as the CPMV of the current block, or the CPMV of the neighboring blocks can be modified based on the size of the neighboring blocks and the size of the current block, and then used as the CPMV of the current block.

[0266] Meanwhile, in the case of affine merging where MVs are derived on a sub-block basis, this can be referred to as the sub-block merging pattern, and it can be specified based on the merge sub-block flag (merge_subblock_flag (value 1)). In this case, the affine merge candidate list, described later, can also be referred to as the sub-block merge candidate list. In this case, the sub-block merge candidate list can also include candidates derived by SbTMVP. In this case, the candidate derived by SbTMVP can be used as the candidate at index 0 in the sub-block merge candidate list. In other words, in the sub-block merge candidate list, the candidate derived by SbTMVP can be positioned before the inherited affine candidates and the constructed affine candidates, described later.

[0267] When applying the affine merge pattern, an affine merge candidate list can be constructed to deduce the CPMV of the current block. For example, the affine merge candidate list may include at least one of the following candidates: 1) Affine candidates of inheritance 2) Constructed affine candidates 3) Zero MV candidates Here, the inherited affine candidate is a candidate derived from the CPMV of neighboring blocks when neighboring blocks are encoded in affine mode. The constructed affine candidate is a candidate derived for each CPMV unit by constructing the CPMV based on the MV of the corresponding CP neighboring blocks. A zero MV candidate can represent a candidate consisting of CPMVs with a value of 0. For example, when the number of current candidates is less than the maximum number of candidates, zero MV candidates can be selectively inserted into the candidate list.

[0268] Figure 15 This is an example illustrating the inheritance of the control point motion vector.

[0269] Two inherited affine candidates can be derived from the affine motion model of the neighboring blocks. One of these candidates can be derived from the left neighboring CU, and the other can be derived from the upper CU.

[0270] Please refer to the above again. Figure 12 For the left predictor, the scan order is A0 -> A1; and for the top predictor, the scan order is B0 -> B1 -> B2. Only the first inherited candidate can be selected from each side, and no pruning check is performed between the two inherited candidates.

[0271] When a neighboring affine CU is identified, the control point motion vector is used to derive the CPMVP candidate from the affine merging list of the current CU. (Reference) Figure 15If the adjacent lower-left block A is encoded in affine mode, then obtain the motion vectors v2, v3, and v4 of the CU including block A, specifically the upper-left, upper-right, and lower-left corners. If block A is encoded using a 4-parameter affine model, then calculate the two CPMVs of the current CU based on v2 and v3. If block A is encoded using a 6-parameter affine model, then calculate the three CPMVs of the current CU based on v2, v3, and v4.

[0272] Figure 16 This is a diagram showing an example of neighboring blocks relative to the current block.

[0273] The constructed affine candidate refers to the candidate constructed by combining the translational motion information of the neighbors of each control point. For example... Figure 16 As shown, the motion information of the control point can be derived from the specified spatial and temporal neighbors.

[0274] CPMV k (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, blocks B2 -> B3 -> A2 are checked in sequence, and the motion vector (MV) of the first available block is used. For CPMV2, blocks B1 -> B0 are checked in sequence, and for CPMV3, blocks A1 -> A0 are checked in sequence. When available, TMVP can be used as CPMV4.

[0275] After obtaining the motion signatures (MVs) of the four control points, affine merging candidates can be constructed based on this motion information. The following combinations of control point MVs can be used sequentially for construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3} A combination of three CPMVs constitutes a 6-parameter affine merging candidate, and a combination of two CPMVs constitutes a 4-parameter affine merging candidate. To avoid issues during motion scaling, if the reference indices of the control points are different, the corresponding control point MV combination is not used.

[0276] Sub-block-based temporal motion vector prediction (SbTMVP)

[0277] The Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP) method can be used. Similar to Temporal Motion Vector Prediction (TMVP) in HEVC, SbTMVP utilizes the motion fields of juxtaposed images to improve motion vector prediction and merging patterns of the current image's CUs.

[0278] SbTMVP also uses the same juxtaposed images used in TMVP.

[0279] The differences between SbTMVP and TMVP are in the following two main aspects: 1. TMVP predicts motion at the CU level, while SbTMVP predicts motion at the subCU level.

[0280] 2. TMVP extracts the motion vectors of the juxtaposed blocks (the juxtaposed blocks are those positioned relative to the bottom right corner or center (bottom right center) of the current CU) from the juxtaposed image. SbTMVP first applies motion shift and then extracts temporal motion information from the juxtaposed image. Here, the motion shift is obtained from the motion vector of one of the spatially neighboring blocks of the current CU.

[0281] Figure 17 The process of SbTMVP is shown.

[0282] SbTMVP predicts the motion vectors of sub-CUs within the current CU in two stages. In the first stage, it checks... Figure 17 (a) shows the spatial neighbor block A1. If A1 has a motion vector that uses the juxtaposed image as a reference image, then that motion vector (which may be called the temporal motion vector (tempVM)) is selected as the motion shift to be applied. If no such motion is identified, then the motion shift is set to (0,0).

[0283] In the second stage, such as Figure 17 As shown in (b), in order to obtain sub-CU level motion information (motion vectors and reference indices) from the juxtaposed image, the motion shift identified in the first stage is applied (i.e., added to the coordinates of the current block). Figure 17 In the example of (b), it is assumed that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of that sub-CU is derived using the motion information of the corresponding block (the smallest motion grid covering the center sample) in the juxtaposed image. When the sub-block has an even number of horizontal and vertical dimensions, the center sample (lower right center sample) can correspond to the lower right sample among the four center samples within the sub-block. After identifying the motion information of the juxtaposed sub-CUs, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC.

[0284] In the second stage, such as Figure 17 As shown in (b), in order to obtain sub-CU level motion information (motion vectors and reference indices) from the juxtaposed image, the motion shift identified in the first stage is applied (i.e., added to the coordinates of the current block). Figure 17In example (b), it is assumed that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of the sub-CU is derived using the motion information of the corresponding block (the smallest motion grid covering the center sample) in the juxtaposed image. When the sub-block has an even number of horizontal and vertical dimensions, the center sample (the center sample in the lower right corner) can correspond to the lower right sample among the four center samples within the sub-block.

[0285] After identifying the motion information of the juxtaposed sub-CU, it is converted into the motion vector and reference index of the current sub-CU using a TMVP process similar to HEVC. At this point, temporal motion scaling can be applied to align the reference image of the temporal motion vector with the reference image of the current CU.

[0286] GPM (Geometric Partitioning)

[0287] Inter-frame prediction can support GPM. GPM is a merging mode that uses CU-level flags for signaling. Other merging modes can include regular merging, MMVD, CIIP, and sub-block merging. For each possible CU size, wxh = 2. m x 2 n It can support a total of 64 partitions. Here, m, n ∈ {3...6}, and 8x64 and 64x8 are excluded.

[0288] When using this mode, the CU can be divided into two parts by a geometrically positioned straight line. The position of the dividing line can be mathematically derived using the angle and offset parameters of the specific partition. Each part of the geometric partition of the CU can perform inter-frame prediction using its own motion. Only unidirectional prediction is allowed per partition, meaning each part has one motion vector and one reference index. Unidirectional prediction motion constraints can be applied to ensure that each CU requires only two motion-compensated predictions, which is the same as regular bidirectional prediction.

[0289] Figure 18 An example is shown of GPM partitions grouped at the same angle.

[0290] Unidirectional predicted motion for each partition can be used Figure 18 The process shown is used to derive the result.

[0291] When using GPM on the current block, the geometric partition index indicating the partitioning pattern (angle and offset) and two merge indices (one per partition) can be additionally signaled. The maximum number of GPM candidate sizes is explicitly signaled in SPS, and the syntax binarization of the GPM merge indexes can be specified. After prediction for each part of the geometric partition, a blending process with adaptive weighting can be used to adjust the sample values ​​along the edges of the geometric partition, such as the blending along the edges of the geometric partition described below. This is the prediction signal for the entire CU, and the transformation and quantization process can be applied to the entire CU as with other prediction patterns. Finally, the motion field of the CU predicted using GPM can be stored in the motion field store for the geometric partitioning pattern.

[0292] CIIP (Intra-Frame and Inter-Frame Joint Prediction)

[0293] Intra- and inter-frame joint prediction (CIIP) can also be applied to the current block. Additional flags (e.g., ciip_flag) can be signaled to specify whether CIIP mode is applied to the current CU. For example, when the CU is encoded in merged mode, and the CU contains at least 64 luma samples (i.e., the product of the CU width and CU height is greater than or equal to 64), and both the CU width and CU height are less than 128 luma samples, additional flags can be signaled to specify whether CIIP mode is applied to the current CU.

[0294] CIIP prediction combines inter-frame prediction signals and intra-frame prediction signals. The inter-frame prediction signal P_inter in CIIP mode can be derived using the same inter-frame prediction process applied in regular combining mode, and the intra-frame prediction signal P_intra can be derived using the regular intra-frame prediction process in planar mode. Then, a weighted average can be used to combine the intra-frame and inter-frame prediction signals. Figure 19 An example is shown of the neighboring blocks used in CIIP weight derivation. Here, the weight values ​​can be calculated based on the encoding patterns of the left and top neighboring blocks as follows: - If the adjacent block above is available and intra-coded, isIntraTop is set to 1; otherwise, isIntraTop is set to 0. - If the left neighboring block is available and intra-coded, isIntraLeft is set to 1; otherwise, isIntraLeft is set to 0. - If (isIntraTop + isIntraLeft) is 2, then wt is set to 3; - Otherwise, if (isIntraTop + isIntraLeft) is 1, then wt is set to 2; - Otherwise, wt is set to 1.

[0295] CIIP prediction can be configured as shown in Equation 4 below.

[0296] [Formula 4]

[0297] GPM including inter-frame and intra-frame prediction

[0298] In GPM, which includes inter-frame and intra-frame prediction, the final prediction sample is generated by weighting the inter-frame and intra-frame prediction samples from the regions separated by each GPM. The inter-frame prediction samples are derived from the inter-frame GPM, while the intra-frame prediction samples are derived from the intra-frame prediction mode (IPM) candidate list and index signaled from the coding device. As an example, the IPM candidate list size can be predefined as 3.

[0299] Figure 20 Examples are shown of IPM candidates available in GPM, including inter-frame and intra-frame prediction.

[0300] like Figure 20 As shown in (a), (b), and (c), the available IPM candidates can be parallel angle modes (parallel mode), perpendicular angle modes (perpendicular mode), and planar modes relative to the GPM block boundary. Furthermore, as... Figure 20 As shown in (d), to reduce the signal notification overhead of IPM and prevent an increase in the size of the intra-frame prediction circuitry of the hardware decoder, the GPM, which includes both inter-frame and intra-frame prediction, can be limited. Meanwhile, to further improve coding performance, direct motion vectors and IPM storage can be introduced for the GPM mixing region.

[0301] In DIMD (decoder-side intra-frame mode derivation) and IPM derivation based on neighboring modes, parallel modes can be registered first. Therefore, if there are no identical IPM candidates in the list, up to two IPM candidates can be registered using the DIMD method and / or derived from neighboring blocks.

[0302] In the nearest neighbor pattern derivation, there are at most 5 available neighboring blocks, but this may be limited by the boundary angles of the GPM blocks already used in the GPM with template matching (GPM-TM).

[0303] Table 5 shows the positions of available neighboring blocks used to derive IPM candidates based on the GPM block boundary angles. A and L can represent the top and left sides of the predicted block, respectively.

[0304] [Table 5]

[0305] Intra-frame GPM (GPM-intra) can be used in combination with GPM-MMVD (GPM merged with motion vector difference). To further improve coding performance, TIMD can be used as an IPM candidate for intra-frame GPM. Parallel mode can register first, followed by TIMD, DIMD, and IPM candidates for neighboring blocks.

[0306] Template Matching (TM)

[0307] Template matching (TM) is a decoder-side motion vector (MV) derivation method. In this method, in order to improve the motion information of the current coding unit (CU), the optimal match is found between the template of the current CU in the current image (i.e., the block above and / or to the left of the current CU) and the block in the reference image (i.e., the block with the same size as the template).

[0308] Figure 21 This is a diagram illustrating the method for deriving motion vectors in template matching. For example, as... Figure 21 As shown, a more suitable motion vector is searched around the initial motion of the current CU within a search range of [-8, +8] pixels. Furthermore, the search step size can be determined according to the AMVR mode, and TM can be continuously applied in merging mode along with the bidirectional matching process.

[0309] In AMVP mode, MVP candidates are determined based on template matching error to select the candidate that minimizes the difference between the current block template and the reference block template. Then, TM is performed only for this MVP candidate to refine the MV. TM refines the corresponding MVP candidate using an iterative diamond search within a search range of [-8, +8] pixels, starting at full-pixel MVD precision (or 4-pixel precision in 4-pixel AMVR mode). AMVP candidates can also be further refined using a cross search at full-pixel MVD precision (or 4-pixel precision in 4-pixel AMVR mode), and then sequentially refined at half-pixel and quarter-pixel precision according to Table 6 specified in AMVR mode.

[0310] [Table 6]

[0311] This search process ensures that AMVP candidates maintain the same MV accuracy specified by the AMVR mode even after the TM process. During the search, the search process terminates if the difference between the previous minimum cost and the current minimum cost in a given iteration is less than a threshold equal to the area of ​​the block.

[0312] In merge mode, a similar search scheme is applied to merge candidates specified by the merge index. As shown in Table 6 above, the precision of TM can be up to 1 / 8 pixel MVD precision, or steps beyond half-pixel MVD precision can be skipped. This depends on whether an alternative interpolation filter used by AMVR in half-pixel mode is used based on the merge motion information.

[0313] Furthermore, when TM mode is enabled, template matching can be performed as a standalone process operation, or as an additional MV refinement process operation between block-based and sub-block-based bilateral matching (BM) schemes. This depends on whether BM meets the enabling conditions and whether it can be applied.

[0314] MHP (Multiple Hypothesis Prediction)

[0315] In the multi-hypothesis inter-frame prediction mode, in addition to the existing bidirectional prediction signal, one or more additional motion compensation prediction signals are signaled, and the resulting overall prediction signal can be obtained by weighted summation of samples. Using the bidirectional prediction signal p_bi and the first additional inter-frame prediction signal / hypothesis h3, the resulting prediction signal p3 can be obtained as follows: [Formula 5]

[0316] The weight factor α can be specified by the new syntax element add_hyp_weight_idx, according to the mapping shown in Table 7 below.

[0317] [Table 7]

[0318] Similar to the above, one or more additional prediction signals can be used. The resulting overall prediction signal can be iteratively accumulated together with each additional signal, as shown in Equation 6 below.

[0319] [Formula 6]

[0320] The resulting overall prediction signal can be obtained as the final p_n (i.e., p_n with the largest index n). For example, at most two additional prediction signals can be used (i.e., n is limited to 2).

[0321] The motion parameters of each additional predictive hypothesis can be explicitly signaled by specifying a reference index, a motion vector predictor index, and a motion vector difference, or implicitly signaled by specifying a merging index. A separate multi-hypothesis merging flag distinguishes between these two signaling modes.

[0322] Example

[0323] Embodiments of this disclosure provide a method applicable to inter-frame image prediction processes to improve compression performance by utilizing various motion vectors. Specifically, by allowing a motion vector predictor (MVP) that includes multiple motion vectors, the accuracy of prediction blocks can be improved through a weighted summation between prediction blocks specified by each motion vector. Furthermore, embodiments of this disclosure allow not only unidirectional and bidirectional prediction, but also other types of prediction.

[0324] According to one embodiment, during the construction of MVP candidates, candidates with multiple motion vectors can be included in the candidate list. Candidates with multiple motion vectors can represent candidates that include additional motion vectors in addition to unidirectional or bidirectional motion vectors (regular motion vectors).

[0325] The number of additional motion vectors can be predefined or can be signaled.

[0326] Candidates with multiple motion vectors can be applied not only to merge mode, but also to various modes such as AMVP mode and AMVP merge mode.

[0327] Candidates with multiple motion vectors can be included after HMVP candidates, and can also be included in the list by replacing another MVP candidate in the list (e.g., a pairwise average candidate).

[0328] When the MVP candidate list includes multiple motion vectors, redundancy with other candidates can be checked. During this redundancy check, the number of motion vectors, the values ​​of the included motion vectors, or the weight values ​​applicable to the reference block specified by each motion vector can be used.

[0329] Candidates used to construct multiple motion vectors can be referred to as target blocks, and candidates included in the MVP candidate list can be used as target blocks. When in a mode that allows multiple motion vectors, target blocks can have their own independent MVP candidate list.

[0330] A candidate with multiple motion vectors may include a target candidate, as well as motion vectors derived from motion vectors included in a block specified by the motion vectors of that candidate.

[0331] [Example 1]

[0332] One embodiment describes a method for constructing MVP candidates in an inter-frame prediction mode, including candidates with multiple motion vectors. In inter-frame prediction modes, motion vector predictors are used in various modes such as AMVP mode and merge mode, and constructing various candidates is efficient because higher predictor accuracy helps improve compression performance.

[0333] According to one embodiment, an MVP candidate with multiple motion vectors can be applied not only to AMVP mode or merging mode, but also to modes such as affine mode, AMVP merging mode, SMVD (Symmetric MVD), sbTMVP, MMVD mode, GPM mode, GPM inter-frame / intra-frame mode, CIIP mode, TM mode, etc. However, the modes / tools listed above are merely examples of applications for an MVP candidate with multiple motion vectors according to embodiments of this disclosure, and the MVP candidate with multiple motion vectors according to embodiments of this disclosure can also be applied to modes or tools not listed above, as long as they are modes or tools that use MVP candidates.

[0334] As mentioned above, the number of multiple motion vectors included in an MVP candidate can be determined based on a predefined number of additional motion vectors. For example, referring to the syntax in Table 8 below, information indicating whether multiple motion vectors are supported or enabled in SPS can be signaled (e.g., sps_multiple_predictor_enabled_flag), and when multiple motion vectors are enabled, i.e., when the corresponding flag is '1', information indicating the maximum number of motion vectors that can be added (e.g., sps_max_num_additional_predictor_minus1). This means that in addition to the unidirectional or bidirectional motion information that existing blocks may have, 'sps_max_num_additional_predictor_minus1 + 1' additional motion information can be included.

[0335] [Table 8]

[0336] Clearly, the information indicating whether multiple motion vectors are enabled and the information indicating the maximum number of motion vectors that can be added can be located in higher parameter sets besides SPS, such as VPS, PPS, APS, PH, SH, etc. As an example, if the flag indicating whether multiple motion vectors are enabled is located in PPS, then enabling multiple motion vectors can be determined on a per-image basis, rather than per-sequence basis. Furthermore, the information indicating whether multiple motion vectors are enabled and the information indicating the maximum number of motion vectors that can be added can be located at different levels, allowing the maximum number of multiple motion vectors to be specified differently depending on the application unit. As an example, as in the example in Table 8 above, the `sps_multiple_predictor_enabled_flag` indicating whether multiple motion vectors are enabled can exist in SPS, and the `sps_max_num_additional_predictor` specifying the maximum number of multiple motion vectors can be signaled at another level, such as image, slice, CTU, or CU, so that different maximum numbers of multiple motion vectors can be variably specified for each unit where the `sps_max_num_additional_predictor` is located.

[0337] Furthermore, the maximum number of multiple motion vectors can be set according to a predefined value without requiring separate signaling. As an example, as shown in the syntax of Table 9 below, when the value of the information indicating whether multiple motion vectors are enabled (e.g., sps_multiple_predictor_enabled_flag) is 1, the maximum number of motion vectors that can be added can be predefined as a specific integer value without requiring separate signaling.

[0338] [Table 9]

[0339] For example, a specific value of 1 or 2 can be used as the maximum number of motion vectors that can be added. This means that 1 or 2 additional motion information can be included in the unidirectional or bidirectional motion information that an existing block can have. Figure 22 This is a diagram illustrating the use of multiple reference blocks in one embodiment.

[0340] In one embodiment, a unidirectional or bidirectional motion vector can be defined as a regular motion vector, and a reference block obtained using that regular motion vector can be defined as a regular reference block. Furthermore, motion vectors added to it can be defined as additional motion vectors, and reference blocks obtained using those motion vectors are defined as additional reference blocks. Multiple motion vectors can be defined as including both regular and additional motion vectors, and multiple reference blocks can be defined as including both regular and additional reference blocks. Alternatively, a regular reference block, multiple reference blocks, or additional reference blocks can also be referred to as a regular prediction block, multiple prediction blocks, or additional prediction blocks. Furthermore, in some cases, multiple reference blocks or multiple prediction blocks can refer to additional reference blocks or additional prediction blocks, and multiple motion vectors can refer to additional motion vectors.

[0341] refer to Figure 22 When multiple motion vectors are enabled for the current block C, bidirectional regular reference blocks P0 and P1 can be obtained, and additional reference blocks P2 and P3 can be obtained based on the additional motion vectors. As an example, the final prediction block can be generated by applying a weighted sum to the regular reference blocks and the additional reference blocks.

[0342] For ease of description, we assume that two reference blocks are obtained for each prediction direction, but it is obvious that the number of reference blocks obtained for each direction may vary.

[0343] Figure 23 This is a flowchart illustrating an example of a method in a decoding or encoding method according to one embodiment for including MVP candidates that include multiple motion vectors in an MVP candidate list. Figure 23 The decoding or encoding method can be executed by the aforementioned decoding device 300 or encoding device 200.

[0344] refer to Figure 23 According to one embodiment, the decoding or encoding method may include: determining the prediction mode applied to the current block as a prediction mode using MVP (S900); constructing an MVP candidate list including MVP candidates with multiple motion vectors (S910); and generating a prediction block based on the final MVP candidates (S920).

[0345] In the case of the decoding method, the image information obtained from the bitstream may include information about prediction, and based on the prediction mode specified by the information about prediction, which is a prediction mode using MVP, such as the AMVP mode or the merge mode mentioned above, the prediction mode applied to the current block can be determined as a prediction mode using MVP (S900).

[0346] The image information obtained from the bitstream may include information indicating whether multi-motion vectors are supported or enabled (e.g., sps_multiple_predictor_enabled_flag). Based on this information indicating that multi-motion vectors are enabled (e.g., sps_multiple_predictor_enabled_flag has a value of 1), an MVP candidate list including MVP candidates with multi-motion vectors can be constructed (S910).

[0347] A prediction block is generated based on the final MVP candidate from the MVP candidate list (S920). The final MVP candidate refers to the MVP candidate in the MVP candidate list used for prediction of the current block. The encoding device 200 can encode information indicating the final MVP candidate (e.g., index information) and transmit it to the decoding device 300 in the form of a bitstream. The decoding device 300 can obtain the corresponding information from the transmitted bitstream and deduce the final MVP candidate. When the final MVP candidate has multiple motion vectors, the prediction block can be generated using additional reference blocks and regular reference blocks together. As an example, a prediction block can be generated by applying weights to regular reference blocks and additional reference blocks.

[0348] Furthermore, the above descriptions of the decoding and encoding methods can also be applied to one embodiment. For example, according to one embodiment, after generating a prediction block, the decoding method can generate a reconstruction block based on the generated prediction block and residual information included in the image information (which may be omitted depending on the prediction mode). According to one embodiment, the encoding method can generate residual information (which may be omitted depending on the prediction mode) based on the generated prediction block, and can encode image information including information about the prediction and residual information, and transmit it in the form of a bitstream.

[0349] Figure 24 This is a diagram illustrating an example of a method for constructing an MVP candidate list in an embodiment of this disclosure.

[0350] Figure 24 The example illustrates the process of constructing a list of MVP candidates for a basic merge pattern. (Reference) Figure 24 The MVP candidate list may include spatially neighboring candidates (S911), temporal candidates (S912), non-adjacent spatially neighboring candidates (S913), HMVP candidates (S914), candidates with multiple motion vectors (S915), and paired candidates (S916).

[0351] However, Figure 24The types, order, and number of MVP candidates shown are merely examples applicable to one embodiment; the corresponding types, order, or number can be changed, or MVP candidates with multiple motion vectors can replace existing MVP candidates. As an example, MVP candidates including multiple motion vectors can be included in the MVP candidate list by replacing them with pairwise average candidates.

[0352] Simultaneously, after constructing the MVP candidate list (S910), each candidate in the MVP candidate list can be reordered based on the template matching cost. In calculating the template matching cost used to reorder the candidates, for a unidirectional reference block, the cost can be calculated based on the difference between the neighboring samples of the current block and the neighboring samples of the reference block; and for a bidirectional reference block, the cost can be calculated based on the difference between the neighboring samples of the current block and the neighboring samples of each reference block.

[0353] For candidates involving multiple motion vectors, the cost can be computed using neighboring samples of the current block and neighboring samples of a reference block specified by each motion vector, or computational complexity can be reduced by using only a subset of reference blocks. In this case, when determining the final prediction block through a weighted sum for each reference block, the cost can be computed by applying the weights for each reference block to its neighboring samples, and when using only a subset of reference blocks, the cost can be computed by applying a modified form of weight. As an example, when using only a subset of reference blocks for cost computation, a method can be applied in which the weights of the corresponding reference blocks are increased and the weights of reference blocks not used for cost computation are decreased. However, this is merely an example, and various other methods of applying modified forms of weight can obviously be applied as well.

[0354] Furthermore, in constructing the MVP candidate list, an approach can be applied that involves building multiple candidates including multiple motion vectors, then reordering these candidates based on template matching costs, and selecting only a subset of them. The same cost calculation method as in the example above can be applied to the calculation of template matching costs. Moreover, since a candidate including multiple motion vectors implies that it includes at least two reference blocks, reordering and selecting only a subset of candidates can be performed by calculating the bilateral matching cost for each predicted block. The corresponding cost can be calculated based on the difference between each reference block, and can be calculated for multiple reference blocks, or a subset of reference blocks can be selected and the cost calculated only for the selected reference blocks.

[0355] Meanwhile, during the process of constructing the MVP candidate list, redundancy checks between MVP candidates with multiple motion vectors and other candidates can be performed as follows: - Redundancy can be determined based on the number of reference blocks for each candidate.

[0356] When each candidate to be compared has the same number of reference blocks, the reference index and motion vector representing the motion information of each candidate can be used to determine whether there is redundancy.

[0357] - When there are multiple candidates with multiple motion vectors in the MVP candidate list, and a weighted sum between reference blocks obtained from each motion vector can be applied, and the weights between the reference blocks of each candidate are different, they can be considered as non-redundant candidates.

[0358] When constructing multiple candidates with multiple motion vectors, the motion vectors included in each candidate can be restricted to at least one that is different from the motion vectors included in the previous candidate with multiple motion vectors in the list. In this way, redundancy can be determined solely based on the motion information between candidates, without considering redundancy checks based on weights.

[0359] Example 2

[0360] This embodiment relates to various signaling notification methods when applying the method of Embodiment 1 to merge mode and AMVP mode.

[0361] MVP candidates with multiple motion vectors can be applied to existing merge patterns as well as to every tool that can be combined with merge patterns (sub-block based MERGE pattern, CIIP pattern, GPM pattern, etc.), and candidates with multiple motion vectors can be included in the process of building the MVP candidate list for each pattern.

[0362] Alternatively, it can be implemented as another merging pattern within the basic merging pattern. For example, as a merging pattern with multiple motion vectors, a flag indicating this pattern (e.g., multi_pred_merge_flag) can be signaled. When the value of this flag is '1', the MVP candidate list can consist of candidates with multiple motion vectors, and a merging index indicating a specific candidate in the MVP candidate list can be signaled. When performing the above-described template-based cost or bilateral cost reordering and the process of selecting only a subset of candidates, the merging index can be omitted, or only candidates within a partial candidate list can be specified.

[0363] Candidates with multiple motion vectors can also be applied to the AMVP mode. Since an MVP candidate list can be built for each direction in the AMVP mode, candidates with multiple motion vectors can be included in each of the MVP candidate lists in the L0 and L1 directions.

[0364] Figure 25This is a diagram showing an example of the MVP candidate list built for each direction in the AMVP model.

[0365] When additional motion vectors are enabled in AMVP mode, by limiting the number of motion vectors that can be present in each direction, it can ultimately include 2 to 4 motion vectors. Figure 25 In the examples, Case 1 shows the case where a total of 4 MVPs are obtained by combining the candidate CAND[2] with multiple motion vectors in the L0 direction MVP candidate list with the candidate CAND[2] with multiple motion vectors in the L1 direction MVP candidate list. Case 2 shows the case where a total of 3 MVPs are obtained by combining the candidate CAND[0] with a single motion vector in the L0 direction MVP candidate list with the candidate CAND[2] with multiple motion vectors in the L1 direction MVP candidate list.

[0366] Considering encoding / decoding efficiency, the application method in AMVP mode can be modified as follows. Specifically, the additional motion vectors in AMVP mode can be applied only to blocks that have applied bidirectional prediction. Furthermore, restricted multi-motion vectors can be enabled by constructing candidates with multiple motion vectors only during the MVP candidate construction process for a specific direction (L0 or L1). That is, by allowing only n (where n is a natural number greater than or equal to 1) additional motion vectors for a specific prediction direction, the total number of additional motion vectors can be limited to n.

[0367] Simultaneously, candidates with multiple motion vectors can be constructed without signaling whether multiple motion vectors are enabled, or a flag indicating the mode can be signaled to allow multiple motion vectors. In the latter case, a mode flag is signaled, and when the flag has a value of '1', MVP candidates with multiple motion vectors can be included in the MVP candidate list. This flag can mean that candidates with multiple motion vectors consisting of the number of enabled motion vectors in the MVP candidate list can be constructed, or it can mean that an MVP candidate list consisting only of candidates with multiple motion vectors can be constructed. Alternatively, it can mean the number of multiple motion vectors that can be included in the MVP candidate list. Furthermore, an mvp_index or mvp_flag can be signaled for a specific candidate in the MVP candidate list, and when performing the above-described template-based cost or bilateral cost reordering and partial candidate selection process, mvp_index or mvp_flag can be omitted, or only candidates in the partial candidate list can be specified.

[0368] Simultaneously, in AMVP mode, when the final MVP candidate, i.e., the MVP candidate used for prediction of the current block, is a candidate with multiple motion vectors, MVD information corresponding to each motion vector can be signaled. Furthermore, as in existing AMVP modes, the final MV can be calculated using the derived MVP information and the signaled MVD information. In this case, the signaling method for MVD information can be applied in various ways, such as indicating the magnitude of the vector, or similar to MMVD, indexing the MVD and deriving it from distance- and orientation-based information. Moreover, for MVPs with multiple motion vectors, only a portion of the MVD information can be signaled, limiting it to regular motion vectors rather than additional motion vectors, thereby reducing the amount of signaling information. That is, the additional motion vector predictors included in the MVP candidate can be directly used as motion vectors without MVD. Alternatively, when the final MVP candidate has multiple motion vectors, variations such as transmitting MVD information for a portion of the motion vectors and not transmitting MVD information for a portion of the motion vectors, or not transmitting the entire MVD information, can also be made.

[0369] The method described in this disclosure can also be similarly applied to the AMVP merging mode. Additional motion vectors can be included in each direction of the AMVP merging mode, and the inclusion of additional motion vectors can be modified, such as limiting their application to each of the AMVP mode or merging mode directions. When additional motion vectors are included in the AMVP mode direction, the candidates specified by `mvp_index` or `mvp_flag` can include multiple motion vectors. In this case, variations such as additionally signaling MVD information or not signaling for the additional motion vectors can be made, and when additional motion vectors are included in the merging mode direction, the candidates specified by the merging index can include multiple motion vectors.

[0370] Example 3

[0371] This embodiment describes a method for generating MVP candidates with multiple motion vectors. According to one embodiment, candidates with multiple motion vectors can be generated by combining other candidates included in the MVP candidate list.

[0372] The MVP candidate list may be composed of motion vectors of spatially and temporally neighboring blocks, history-based MVP (HMVP), pairwise average candidates, etc., and in one embodiment, such candidates are referred to as target candidates or target blocks. A candidate with multiple motion vectors can be generated by using motion vectors included in the target candidates, and in one embodiment, the motion vectors used to generate multiple motion vectors will be referred to as target motion vectors. Furthermore, not only candidates included in the MVP candidate list, but also candidates in another list constructed for generating multiple motion vectors may also be regarded as target candidates. Names such as target candidate and target motion vector are given for convenience of description, and even if names other than these are used, as long as they are included in the MVP candidate list for generating a candidate with multiple motion vectors, they can fall within the scope of target candidates and target motion vectors in one embodiment.

[0373] As described above, as an MVP candidate, multiple motion vectors can be included between an HMVP candidate and a pairwise average candidate in the MVP candidate list. Alternatively, it can also be applied by replacing an existing candidate such as a pairwise average candidate. When the maximum number of MVP candidates is N, if the number of candidates included in the candidate list is less than N, a candidate including multiple motion vectors can be considered, and when the number of candidates included in the list is less than M (for example, 2), the corresponding process can be omitted. Here, N and M are integer values greater than 0.

[0374] Figure 26 is a diagram showing another example of a method for generating a candidate with multiple motion vectors according to an embodiment.

[0375] As described above, a candidate with multiple motion vectors can be generated by combining other candidates included in the MVP candidate list. Referring to Figure 26 example, when constructing a list with a maximum of N MVP candidates, a candidate with multiple motion vectors can be generated by combining other partial candidates CAND[0]...CAND[M-1] (M<N) included in the list.

[0376] According to this example, a candidate CAND[M] with multiple motion vectors can be generated by combining the first target candidate CAND[0] and the second target candidate CAND[1] in the MVP candidate list: CAND[0] + CAND[1]. A candidate CAND[M+1] with multiple motion vectors can be generated by combining the first target candidate CAND[0] and the third target candidate CAND[2]: CAND[0] + CAND[2]. A candidate CAND[M+2] with multiple motion vectors can be generated by combining the second target candidate CAND[1] and the third target candidate CAND[2]. This example describes the case of generating three candidates with multiple motion vectors, but the number of candidates with multiple motion vectors generated by combining target candidates may be different. And obviously, the number of target candidates used in the combination, the order of the target candidates in the list, the order of the candidates with multiple motion vectors, etc., may also be different. Figure 25 The examples are implemented differently.

[0377] The target candidate used to generate multiple motion vectors can be an MVP candidate for unidirectional prediction or an MVP candidate for bidirectional prediction. Depending on the amount of additional motion information allowed for a candidate with multiple motion vectors, all or part of the motion vectors of the target candidate can be used. As an example, when the target candidate for unidirectional prediction, or the target candidate for bidirectional prediction, meets the allowed number of additional motion vectors as the candidate for unidirectional prediction, all motion information can be included in the candidate with multiple motion vectors without any additional conditions.

[0378] Furthermore, multiple motion vectors can be constructed in the following order. This example assumes the existence of target candidates A and B for bidirectional prediction, to generate candidates with multiple motion vectors. In this case, candidates A and B can be targets that have been reordered within the candidate list based on template cost or bilateral cost: - It can first include the motion vectors of candidate A in the L0 and L1 directions, and additionally include the motion vectors of candidate B in the L0 or L1 direction.

[0379] - It can first include the motion vectors of candidate B in the L0 and L1 directions, and additionally include the motion vectors of candidate A in the L0 or L1 direction.

[0380] - It can first include the motion vector of candidate A in the L0 direction and the motion vector of candidate B in the L1 direction, and then additionally include the motion vector of candidate A in the L1 direction or the motion vector of candidate B in the L0 direction.

[0381] - It can first include the motion vector of candidate A in the L1 direction and the motion vector of candidate B in the L0 direction, and then additionally include the motion vector of candidate A in the L0 direction or the motion vector of candidate B in the L1 direction.

[0382] - It can first include the motion vector of candidate A in the L0 direction and the motion vector of candidate B in the L0 direction, and then additionally include the motion vector of candidate A in the L1 direction or the motion vector of candidate B in the L1 direction.

[0383] - It can first include the motion vector of candidate A in the L1 direction and the motion vector of candidate B in the L1 direction, and then additionally include the motion vector of candidate A in the L0 direction or the motion vector of candidate B in the L0 direction.

[0384] - In addition, an MVP candidate list with multiple motion vectors has been added, which can undergo an additional reordering process using predefined costs (e.g., template-based costs, bilateral-based costs, etc.).

[0385] In the above embodiments, the following scenario has been described: after the MVP candidate list is reordered based on a predefined cost, candidates with multiple motion vectors are generated based on the reordered MVP candidates. However, to reduce encoding and decoding complexity, it is also possible to generate multiple motion vectors and add them to the MVP candidate list without reordering the MVP candidate list before generating candidates with multiple motion vectors, and then perform reordering based on a predefined cost.

[0386] Furthermore, when using another candidate with multiple motion vectors to generate a candidate with multiple motion vectors, or when only partial motion vectors need to be selected, such as when the number of motion vectors of the target candidate exceeds the allowed number of additional motion vectors, the methods listed above can be considered, but only partial motion vectors of the candidate being combined can be used to determine the additional motion vectors.

[0387] As a specific example, in the methods listed above, in addition to the motion information initially included, additional motion information can be included using the distance between each candidate reference image to be compared and the current image. This allows for the inclusion of motion information from candidates with shorter distances to the current image as additional motion information. Alternatively, when applying a weighted sum between reference blocks specified by each motion vector, motion vectors indicating the reference blocks to which larger weights are applied can be included as additional motion vectors. Alternatively, the encoding / decoding process can be simplified by including motion vectors in a fixed direction.

[0388] As in the example above, multiple candidates with multiple motion vectors can be generated and included in the MVP candidate list. As an example, the number of candidates with multiple motion vectors that can be included in the MVP candidate list can be predetermined. Alternatively, to consider candidates that can be included in the MVP candidate list after those with multiple motion vectors, multiple multiple motion vectors can be generated until the number of MVP candidates including those with multiple motion vectors is filled to N-2. Here, 2 is an example; it can also be changed and applied to N-1, N-3, etc., taking into account the positions of multiple reference blocks.

[0389] Example 4

[0390] This embodiment describes another method for generating MVP candidates with multiple motion vectors. According to embodiments of this disclosure, candidates with multiple motion vectors can be generated by combining the motion vectors of another candidate included in the MVP candidate list with motion vectors derived using that motion vector.

[0391] As described above, the MVP candidate list can consist of motion vectors from spatially and temporally proximate blocks, history-based MVPs (HMVPs), pairwise averaged candidates, etc., and in the embodiments of this disclosure, such candidates are referred to as target candidates. Candidates with multiple motion vectors can be generated using motion vectors included in the target candidates, and in the embodiments of this disclosure, the motion vectors used to generate multiple motion vectors are referred to as target motion vectors. Furthermore, not only candidates included in the MVP candidate list, but also candidates in another list constructed for generating multiple motion vectors can be considered target candidates. The names such as target candidates and target motion vectors are given for ease of description, and even if names other than these are used, as long as they are included in the MVP candidate list for generating candidates with multiple motion vectors, they fall within the scope of target candidates and target motion vectors in the embodiments of this disclosure.

[0392] As an example, as an MVP candidate, multiple motion vectors can be included between the HMVP and pairwise average candidates in the MVP candidate list. Alternatively, this can be applied by replacing existing candidates such as pairwise average candidates. When the maximum number of MVP candidates is N, if the number of candidates included in the candidate list is less than N, candidates with multiple motion vectors can be considered, and when the number of candidates included in the list is less than M (e.g., 2), the corresponding process can be omitted. Here, N and M are integer values ​​greater than 0.

[0393] As a specific example, suppose any candidate in the list is CAND[i], and the block at the position specified by the motion vector of CAND[i] in the L0 or L1 direction is the corresponding block (reference block). If the corresponding block is in inter-frame mode (if the corresponding block is an inter-frame block), the additional motion vector can be derived. That is, when inter-frame mode is applied to the corresponding block, the block's motion vector and the reference image can be used as additional motion information.

[0394] Therefore, assuming the motion vector of CAND[i] in the L0 or L1 direction is MV[X] (X=0..1), and the motion vector of the corresponding block specified by this motion vector is CX_MV[Y] (X=0..1, Y=0..1), then the additional motion vector can be calculated as MV[X] + CX_MV[Y]. Here, CX_MV[Y] represents the motion vector of the corresponding block in the LX direction of CAND[i] in the LY direction. Therefore, a candidate with multiple motion information can be composed of the following motion vectors, or can be composed of a subset of the following listed motion vectors: - MV[0], - MV[1], - MV[0] + C0_MV[0], - MV[0] + C0_MV[1], - MV[1] + C1_MV[0] - MV[1] + C1_MV[1] In other words, an MVP candidate with multiple motion vectors can be constructed using the motion vectors MV[0] and MV[1] of the target block, as well as MV[0] + C0_MV[0], MV[0] + C0_MV[1], MV[1] + C1_MV[0], and MV[1] + C1_MV[1] derived from the motion vectors of the target block and the bidirectional motion vectors of the corresponding block. However, this is just one example that can be applied to one embodiment, and various combinations can be generated depending on which direction of motion vector is taken from the motion vectors of the target block and which direction of motion vector is taken from the motion vectors of the corresponding block. In this case, the reference image specified by the additional motion information can be the reference image specified by the motion information of the corresponding block. Furthermore, the motion vector of the corresponding block can be scaled by the distance between the image of the target block and the reference image, and the distance between the image of the corresponding block and the image specified by the motion vector of the corresponding block.

[0395] Furthermore, the above method is not limited to the case where the corresponding block is in inter-frame mode, and the same method can be applied even when the corresponding block is in IBC mode, using block vectors to obtain additional motion information. That is, various additional motion vectors can be generated by combining motion vectors and block vectors. Specifically, assuming the motion vector of the target block is MV[X] (X = 0..1), and the block vector of the corresponding block in IBC mode is CX_BV[Y] (X = 0..1, Y = 0..1), then the additional motion vector can be calculated as MV[X] + CX_BV[Y]. In this case, candidates with multiple motion information can be composed of the following motion vectors, or can be composed of some of the motion vectors listed below. When IBC mode supports unidirectional block vectors, a single CX_BV can be used to calculate the additional motion vector: - MV[0], - MV[1], - MV[0] + C0_BV[0], - MV[0] + C0_BV[1], - MV[1] + C1_BV[0], - MV[1] + C1_BV[1] In other words, MVP candidates with multiple motion vectors can be constructed using the motion vectors MV[0] and MV[1] of the target block, as well as MV[0] + C0_BV[0], MV[0] + C0_BV[1], MV[1] + C1_BV[0], and MV[1] + C1_BV[1] derived from the motion vectors of the target block and the bidirectional block vectors of the corresponding block. However, this is just one example that can be applied to one embodiment, and various combinations can be generated depending on which direction of the motion vector is taken from the motion vectors of the target block and which direction of the block vector is taken from the block vectors of the corresponding block. Since the block vectors are used for intra-frame image prediction, the reference image in the above example can be a reference image specified by the motion vectors of the target block.

[0396] When the corresponding block is in intra-frame mode (when the corresponding block is an intra-frame block), that is, when intra-frame mode is applied to the corresponding block, motion vectors can be obtained from the corresponding block specified by motion vectors in different directions, or a candidate with multiple motion vectors can be derived by targeting another candidate included in the MVP candidate list without using the target block. Alternatively, even when intra-frame mode is applied to the corresponding block, if the block vectors of the corresponding block are available, the block vectors of the corresponding block can be used to obtain additional motion information. That is, as described above, various additional motion vectors can be generated by combining motion vectors and block vectors.

[0397] Specifically, assuming the motion vector of the target block is MV[X] (X = 0..1), and the block vector when the corresponding block is in intra-frame mode is CX_BV[Y] (X=0..1, Y=0..1), then the additional motion vector can be calculated as MV[X] + CX_BV[Y]. Candidates with multiple motion information can be composed of the following motion vectors, or can be composed of a subset of the motion vectors listed below. When the corresponding block is in IBC mode and IBC mode supports unidirectional block vectors, a single CX_BV can be used to calculate the additional motion vector: - MV[0], - MV[1], - MV[0] + C0_BV[0], - MV[0] + C0_BV[1], - MV[1] + C1_BV[0], - MV[1] + C1_BV[1] This is one example that can be applied to an embodiment, and various combinations can be generated depending on which direction of the motion vector is taken from the motion vector of the target block and which direction of the block vector is taken from the block vector of the corresponding block. Since the block vectors are used for intra-frame image prediction, the reference image in the above example can be a reference image specified by the motion vector of the target block.

[0398] In the examples above, the availability of block vectors in intra-frame mode refers to the case where the block vector is stored for the corresponding block. For example, block vectors might be available in the following modes: - Intra-frame TMP: A pattern for generating predicted blocks by matching the template of the current block with the template of a reference block in the current image to find blocks with small errors.

[0399] - SGPM-IBC: SGPM is an intra-frame technique that performs different predictions for each region divided by partitioning. Intra-frame / intra-frame combined predictions are performed for each partitioned region, and SGPM-IBC can apply intra / IBC, IBC / intra, or IBC / IBC predictions.

[0400] The above-described modes are examples of intra-frame modes in which block vectors can be used. Besides the examples described above, block vectors can also be used and stored in intra-frame modes through combinations of intra-frame TMP with another intra-frame mode, combinations of IBC mode with another intra-frame mode, etc. In other words, any intra-frame mode that uses and stores block vectors, even if it is not one of the modes listed above, falls within the scope of the embodiments of this disclosure.

[0401] As described above, MVP candidates with multiple motion vectors can be generated by considering the target block and the motion vectors of the corresponding blocks specified by the motion vectors of the target block, but the specific method can be applied differently depending on the number of allowed additional motion vectors. As an example, when bidirectional prediction is applied to the target block CAND[i], there are motion vectors in either the L0 or L1 direction, and when bidirectional prediction is applied to the corresponding blocks in each direction, a total of 6 motion vectors may be available. Depending on the number of allowed additional motion vectors (the number of allowed additional reference blocks), a subset of these motion vectors can be selected to construct candidates with multiple motion vectors. More specifically, candidates with multiple motion information can be generated using the following method: - When a target candidate A exists and bidirectional prediction is applied to it, the motion vectors of the target candidate A in the L0 and L1 directions can be included first, and when the corresponding block specified by the motion vectors in the L0 and / or L1 directions is in inter-frame mode, intra-frame mode including block vectors, or IBC mode, the motion vectors of that block or the block vectors can be used to derive additional motion vectors.

[0402] - When the corresponding block specified by the motion vector in the LX (X = 0 or 1) direction is not in inter-frame mode, intra-frame mode including block vectors, or IBC mode, another corresponding block can be derived, and additional motion vectors can be generated by repeating the same process using the motion vector in the L(1-X) direction.

[0403] - Motion vectors in the LX (X = 0 or 1) direction can be searched first, and when the corresponding block specified by the motion vector in that direction is in inter-frame mode, intra-frame mode including block vectors, or IBC mode, the motion vector or block vector of that block can be used to derive additional motion vectors. When the corresponding block in the first searched direction allows additional motion vectors (such as when performing unidirectional prediction), it can be determined whether additional motion vectors for corresponding blocks in the remaining directions are allowed.

[0404] - When applying bidirectional prediction to target candidates, to determine the initial search direction, the distance between the current image and the reference images of the target candidate in the L0 and L1 directions can be considered. As an example, the direction with the shorter distance can be preferred. Figure 27 It shows the distance between the current image and the reference image. In other words, it shows the distance between the current image and the reference image through comparison. Figure 27 In the equations d0 and d1, the motion vectors with smaller values ​​can be selected preferentially.

[0405] Motion vectors in the L0 and L1 directions included in the target candidate can be included in the MVP candidate with multiple motion vectors, and as additional motion vectors, motion vectors in shorter distance directions can be considered, taking into account the distance between the current image and a reference image including additional reference blocks derived from the corresponding blocks in each of the L0 and L1 directions. That is, by comparing... Figure 27 In the D0 and D1 parameters, the motion vector with the smaller value can be selected first.

[0406] Motion vectors in the L0 and L1 directions included in the target candidate can be included in the MVP candidate with multiple motion vectors, and as additional motion vectors, motion vectors in directions with shorter distances can be preferentially selected, taking into account the distance between the reference image in each of the L0 and L1 directions and the reference image including additional reference blocks derived from the corresponding blocks in each of the L0 and L1 directions. That is, by comparing Figure 27 In the (D0-d0) and (D1-d1) values, the motion vectors in the direction with smaller values ​​can be selected preferentially.

[0407] - When the reference blocks in each of the L0 and L1 directions included in the target candidate are different from each other, such as an inter-frame block and an IBC block, or an intra-frame mode including block vectors, the motion vector of the inter-frame block can be preferentially selected. Alternatively, it can be defined in the order of motion vector of the inter-frame block, block vector of the IBC block, and block vector of the INTRA block.

[0408] - When the reference images in each of the L0 and L1 directions included in the target candidate have a list of reference images reordered based on a specific cost, the motion vector indicating the reference image at the top of the list, i.e., with a lower cost, can be preferentially selected.

[0409] - Motion information of the blocks specified by motion vectors in the L0 and L1 directions included in the target candidates can be obtained sequentially. As an example, when applying bidirectional prediction to all / part of the blocks specified by motion vectors in the L0 and L1 directions, one of the motion vectors can be selected and used as an additional reference vector.

[0410] - When applying bidirectional prediction to the corresponding block specified by the motion vector of the target candidate in the LX direction, the motion vector of the corresponding block in the LX direction can be used when calculating the additional motion vector.

[0411] - When calculating additional motion vectors, the motion vectors in the L0 direction included in the corresponding block can be used.

[0412] - Among the motion vectors included in the corresponding block, when calculating additional motion vectors, the motion vectors of a reference image that is closer to the current image can be used.

[0413] In the methods listed above, in addition to the motion information initially included, additional motion information can be used as additional motion vectors by utilizing the distance between each candidate reference image to be compared and the current image, such that the motion information of candidates with shorter distances to the current image is used as additional motion vectors. Alternatively, when applying a weighted sum between reference blocks specified by each motion vector, the motion vector indicating the reference block to which a larger weight is applied can be used as the additional motion vector. Alternatively, the encoding / decoding process can be simplified by including motion vectors in a fixed direction.

[0414] Simultaneously, the number of candidates with multiple motion vectors that can be included in the MVP candidate list can be predetermined. Alternatively, in order to consider candidates that can be included in the MVP list after candidates with multiple motion vectors, the number of MVP candidates including candidates with multiple motion vectors can be set to N-2. Here, 2 is an example, and it can also be changed and applied to N-1, N-3, etc., taking into account the positions of multiple reference blocks.

[0415] Similar to the example above, multiple candidates with multiple motion vectors can be generated and included in the MVP candidate list. As an example, the number of candidates with multiple motion vectors that can be included in the MVP candidate list can be predetermined. Alternatively, to consider candidates that can be included in the MVP candidate list after those with multiple motion vectors, multiple multiple motion vectors can be generated until the number of MVP candidates including those with multiple motion vectors is filled to N-2. Here, 2 is just an example; it can be changed and applied to N-1, N-3, etc., taking into account the positions of multiple reference blocks.

[0416] Example 5

[0417] This embodiment describes a method for deriving additional motion information based on the position of a predicted block specified by the motion vector of the target block CAND[i].

[0418] To compress the amount of motion information data at the image level, the motion vectors included in the current block are stored in units of a predetermined size (e.g., 16x16) after image-level encoding / decoding, instead of being stored in units of CU blocks. Figure 28 This is a diagram illustrating an example of the position of a reference block specified by the motion vector of the current block. Reference Figure 28When the motion vector of the current block indicates a specific location within the encoded / decoded reference image, examples are shown where the reference block of the current block is included in a single motion information storage unit (Case 1), and examples are shown where it is included in multiple storage units (Case 2). Figure 28 Case 1 illustrates the case where the motion vector of CAND[i] points to a single storage cell within the reference image that stores motion information. Case 2 illustrates the case where the motion vector points to a location within the reference image that includes multiple storage cells. Typically, the location specified by the motion vector is set to the (0,0) position of the block, and the motion vector of the block containing the center sample located at (width / 2, height / 2) can be used as the motion vector of the corresponding block. However, when pointing to a location with a different motion vector, as in Case 2, an inaccurate motion vector may be derived.

[0419] In embodiments of this disclosure, the motion vector stored at the position specified by the motion vector of CAND[i] can be used as an additional motion vector, or it can be used to calculate an additional motion vector. Specifically, the motion vector of a block that includes a specific position within the corresponding block can be used, and as an example, various modifications can be made, such as using the motion vector of a block that includes a sample at the center position (width / 2, height / 2). Furthermore, the motion vector of a block that includes a sample at a position shifted by a predefined offset can also be used as an additional motion vector.

[0420] Figure 29 This is a diagram illustrating an example of an offset that can be applied in one embodiment.

[0421] like Figure 29 As shown in (a), one of the four offset candidates based on the center position shown in Equation 7 can be applied to find the block that includes the corrected position.

[0422] [Formula 7]

[0423] (width / 2, height / 2)

[0424] (width / 2, height / 2) + (offset, 0)

[0425] (width / 2, height / 2) + (-offset, 0)

[0426] (width / 2, height / 2) + (0, offset)

[0427] (width / 2, height / 2) + (0, -offset)

[0428] The offset can have an integer value, and as an example, values ​​such as 1, 2, 4, 8, and 16 can be applied. Furthermore, the motion vector at the corrected position, as shown in Equation 7, can be represented as C_MV[X][0]..C_MV[X][K-1], etc. Here, X represents the predicted direction and can have a value of 0 or 1, and K can represent the number of available corrected position candidates.

[0429] When a target block CAND[i] exists, multiple motion vectors can be derived by applying various correction values. Motion vectors derived in this way can be used to construct multiple additional motion vectors as follows: - MV[0], - MV[1], - MV[0] + C_MV[0][0], - MV[0] + C_MV[0][1] Figure 30 This is a diagram illustrating an example of the process of correcting the position by applying an offset within the corresponding block and deriving the motion vector at each corrected position.

[0430] refer to Figure 30 The position can be corrected within the corresponding block specified by the motion vector MV[0] pointed to by the target block CAND[i] using the offset, and two motion vectors MV[0] + C_MV[0][0] and MV[0] + C_MV[0][1] can be derived at each corrected position.

[0431] Alternatively, when a target block CAND[i] exists, multiple candidates with additional motion vectors can be generated by applying various offsets. Table 10 below shows an example of generating two candidates with additional motion vectors. Referring to Table 10, two different MVP candidates can be generated using the motion vectors MV_i[0] and MV_i[1] of the target block CAND[i] and the motion vectors C_MV[0][0] and C_MV[0][1] of the corresponding block with applied offsets.

[0432] [Table 10]

[0433] Alternatively, when there is a target block CAND[i] and another target block CAND[j] (i != j), various MVP candidates can be generated by applying the motion vectors C_MV[0][0] and C_MV[0][1] of the corresponding blocks with different offsets. Table 11 below shows an example of generating two candidates with additional motion vectors using the motion vectors MV_i[X] and MV_j[X] of different target blocks. Here, X has a value of 0 or 1, meaning the x-axis or y-axis.

[0434] [Table 11]

[0435] In other words, multiple candidates can be generated not only using the motion vector MV[0] included in the target block CAND[i], but also by using offset correction positions within the corresponding block specified by the motion vector. Alternatively, multiple candidates can be generated by deriving from one or more predefined sample positions within the corresponding block. As an example, the motion vector of a block including samples of each vertex position, i.e. (0,0), (width-1,0), (0,height-1), (width-1,height-1), or a subset of these positions, can each be used as an additional motion vector, or used to compute an additional motion vector. Alternatively, modifications such as selecting the representative motion vector covered by the largest number of samples and using it as a candidate can also be made. Furthermore, it is apparent that when generating additional motion vectors, redundancy checks between candidates can be performed, enabling the consideration of various candidates. Thus far, for ease of description, the various embodiments have been described separately, but it is apparent that combinations of two or more embodiments are also feasible, and necessary modifications resulting from combinations of embodiments may also be included within the scope of this disclosure or its embodiments.

[0436] Figure 31 Examples of content streaming systems to which embodiments of the present disclosure can be applied are shown.

[0437] refer to Figure 31 The content streaming system using the embodiments of this disclosure may mainly include an encoding server, a streaming server, a web server, media storage, a user device, and a multimedia input device.

[0438] An encoding server generates a bitstream by compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, and then sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.

[0439] Bitstreams can be generated by applying the encoding method or bitstream generation method of the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.

[0440] A streaming server sends multimedia data to a user's device via a web server based on a user's request, and the web server acts as a medium to notify the user what services are available. When a user requests a service from the web server, the web server delivers it to the streaming server, and the streaming server sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which in this case controls the commands / responses between each device in the content streaming system.

[0441] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store bitstreams for specific time periods.

[0442] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs), digital TVs, desktop computers, digital signage, etc.).

[0443] In a content streaming system, each server can be operated as a distributed server, and in this case, data received from each server can be distributed and processed.

[0444] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined and implemented as a device, and the technical features of the device claims of this disclosure can be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this disclosure can be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this disclosure can be combined and implemented as a method.

[0445] Industrial applicability

[0446] The embodiments of this disclosure can be used to encode / decode images.

Claims

1. A decoding method, comprising: Obtain image information from the bitstream; Based on the acquired image information, the prediction mode applied to the current block is determined to be the prediction mode using the motion vector predictor (MVP); Construct an MVP candidate list that includes the MVP candidates of the current block; as well as Based on at least one MVP candidate from the MVP candidate list, generate a predicted block for the current block. The construction of the MVP candidate list includes including MVP candidates with regular motion vectors and additional motion vectors in the MVP candidate list.

2. The method according to claim 1, wherein, The acquired image information includes at least one of information related to whether the additional motion vectors are enabled or information related to the maximum number of the additional motion vectors.

3. The method according to claim 1, further comprising: Based on the cost of calculating the MVP candidates, the MVP candidate list is reordered. The reordering includes applying weights to neighboring samples of the reference block specified by the conventional motion vector and neighboring samples of the reference block specified by the additional motion vector to calculate the cost of MVP candidates with the conventional motion vector and the additional motion vector.

4. The method according to claim 1, wherein, The construction of the MVP candidate list includes: checking for redundancy between an MVP candidate with the regular motion vector and the additional motion vector and another MVP candidate in the MVP candidate list.

5. The method according to claim 1, wherein, The construction of the MVP candidate list includes: based on the prediction mode applied to the current block being AMVP (Advanced MVP) mode, including MVP candidates with the regular motion vector and the additional motion vector in one of the MVP candidate list in the L0 direction and the MVP candidate list in the L1 direction.

6. The method according to claim 1, wherein, The final MVP candidate among the MVP candidates is an MVP candidate with the regular motion vector and the additional motion vector, and the acquired image information includes motion vector difference information for at least one of the regular motion vector and the additional motion vector.

7. An encoding method, comprising: The prediction mode applied to the current block is determined to be the prediction mode using the motion vector predictor (MVP); Construct an MVP candidate list that includes the MVP candidates of the current block; Based on the final MVP candidates in the MVP candidate list, generate the prediction block for the current block; as well as The image information, including information about the prediction pattern, is encoded. The construction of the MVP candidate list includes including MVP candidates with regular motion vectors and additional motion vectors in the MVP candidate list.

8. The method according to claim 7, wherein, The encoded image information includes at least one of information related to whether the additional motion vectors are enabled or information related to the maximum number of the additional motion vectors.

9. The method according to claim 7, further comprising: The MVP candidate list is reordered based on the cost calculated for each MVP candidate. The reordering process includes applying weights to neighboring samples of the reference block specified by the conventional motion vector and neighboring samples of the reference block specified by the additional motion vector to calculate the cost of an MVP candidate with the conventional motion vector and the additional motion vector.

10. The method according to claim 7, wherein, The construction of the MVP candidate list includes: checking for redundancy between an MVP candidate with the regular motion vector and the additional motion vector and another MVP candidate in the MVP candidate list.

11. The method according to claim 7, wherein, The construction of the MVP candidate list includes: based on the prediction mode applied to the current block being AMVP (Advanced MVP) mode, including MVP candidates with the regular motion vector and the additional motion vector in one of the MVP candidate list in the L0 direction and the MVP candidate list in the L1 direction.

12. The method according to claim 7, wherein, The encoding of the image information includes: encoding motion vector difference information for at least one of the regular motion vector and the additional motion vector, based on the fact that the final MVP candidate is an MVP candidate with the regular motion vector and the additional motion vector.

13. A computer-readable storage medium storing a bitstream generated by an encoding method. in, The encoding method includes: The prediction mode applied to the current block is determined to be the prediction mode using the motion vector predictor (MVP); Construct an MVP candidate list that includes the MVP candidates of the current block; Based on the final MVP candidates in the MVP candidate list, generate the predicted block for the current block; and The image information, including information about the prediction pattern, is encoded. The construction of the MVP candidate list includes including MVP candidates with regular motion vectors and additional motion vectors in the MVP candidate list.

14. A method for transmitting data for an image, comprising: Obtain the bitstream of the image, wherein the bitstream is generated based on: determining the prediction mode applied to the current block as a prediction mode using a motion vector predictor (MVP); constructing an MVP candidate list including MVP candidates for the current block; generating a prediction block for the current block based on the final MVP candidates in the MVP candidate list; and encoding image information including information about the prediction mode; and Transmitting data including the bit stream, The construction of the MVP candidate list includes including MVP candidates with regular motion vectors and additional motion vectors.