Method and computer-readable storage medium

By constructing a motion vector predictor candidate list with multiple motion vectors and applying weight information, the method addresses inefficiencies in inter-frame prediction and oversmoothing, enhancing video encoding and decoding quality for high-resolution video.

WO2026089458A1PCT designated stage Publication Date: 2026-04-30LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-10-21
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing video compression technologies face challenges in effectively encoding and decoding high-resolution, high-quality video due to inefficiencies in inter-frame prediction and oversmoothing issues, particularly in handling multiple motion vectors and interpolation filters.

Method used

Incorporating a method to construct a motion vector predictor candidate list with multiple motion vectors and applying weight information, along with limiting interpolation filters to improve inter-frame prediction performance and reduce oversmoothing.

Benefits of technology

Enhances inter-frame prediction efficiency and compression performance by utilizing various motion vectors and signaling weight information, reducing oversmoothing and improving overall video encoding and decoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025016754_30042026_PF_FP_ABST
    Figure KR2025016754_30042026_PF_FP_ABST
Patent Text Reader

Abstract

A method according to an embodiment comprises the steps of: acquiring image information from a bitstream; configuring a motion information candidate list comprising motion information candidates for the current block, on the basis of the image information; and generating a prediction block for the current block, on the basis of at least one motion information candidate within the motion information candidate list, wherein the step of configuring the motion information candidate list comprises adding a motion information candidate having basic motion information and additional motion information to the motion information candidate list, and the image information comprises weight information applied to a weighted sum of a basic prediction block derived on the basis of the basic motion information and an additional prediction block derived on the basis of the additional motion information.
Need to check novelty before this filing date? Find Prior Art

Description

Method and computer-readable storage medium

[0001] The present disclosure relates to a method for encoding / decoding image information, a computer-readable storage medium for storing a bitstream, and a method for transmitting data.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields, and accordingly, high-efficiency video compression technologies are being discussed.

[0003] Various image compression technologies exist, such as inter-prediction technology that predicts pixel values ​​in the current picture from previous or subsequent pictures, intra-prediction technology that predicts pixel values ​​in the current picture using pixel information within the current picture, and entropy coding technology that assigns short codes to values ​​with high frequency and long codes to values ​​with low frequency; by utilizing these image compression technologies, image data can be effectively compressed for transmission or storage.

[0004] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.

[0005] The present disclosure provides a method for improving inter-frame prediction performance by utilizing various motion vectors, wherein, in the process of constructing a motion vector predictor candidate list in an inter-prediction mode, a candidate including multiple motion vectors is included in the list.

[0006] In addition, it provides a method for signaling weight information applied to multiple motion vectors and a method for limiting the use of interpolation filters to reduce oversmoothing caused by weighted sums between prediction blocks.

[0007] A method according to one embodiment comprises: a step of acquiring image information from a bitstream; a step of forming a motion information candidate list including motion information candidates for a current block based on the image information; and a step of generating a prediction block for the current block based on at least one motion information candidate in the motion information candidate list; wherein the step of forming the motion information candidate list includes adding a motion information candidate having basic motion information and additional motion information to the motion information candidate list, and the image information includes weight information applied to a weighted sum of a basic prediction block derived based on the basic motion information and an additional prediction block derived based on the additional motion information.

[0008] A method according to one embodiment comprises: a step of configuring a motion information candidate list including motion information candidates for a current block; a step of generating a prediction block for the current block based on at least one motion information candidate in the motion information candidate list; and a step of encoding image information for the current block; wherein the step of configuring the motion information candidate list includes adding a motion information candidate having basic motion information and additional motion information to the motion information candidate list, and the image information includes weight information applied to a weighted sum of a basic prediction block derived based on the basic motion information and an additional prediction block derived based on the additional motion information.

[0009] A computer-readable storage medium storing a bitstream generated by an encoding method according to one embodiment, wherein the encoding method comprises: a step of configuring an MVP candidate list including MVP candidates for a current block; a step of generating a prediction block for the current block based on a final MVP candidate in the MVP candidate list; and a step of encoding image information for the current block; wherein the step of configuring the MVP candidate list includes adding an MVP candidate having basic motion information and additional motion information to the MVP candidate list, and the image information includes weight information applied to a weighted sum of a basic prediction block derived based on the basic motion information and an additional prediction block derived based on the additional motion information.

[0010] A method for transmitting data for an image according to one embodiment comprises: acquiring a bitstream for the image, wherein the bitstream is generated based on the steps of: configuring a motion information candidate list including motion information candidates for a current block; generating a prediction block for the current block based on at least one motion information candidate in the motion information candidate list; and encoding image information for the current block; and transmitting data including the bitstream. The step of configuring the motion information candidate list includes adding a motion information candidate having basic motion information and additional motion information to the motion information candidate list, and the image information includes weight information applied to a weighted sum of a basic prediction block derived based on the basic motion information and an additional prediction block derived based on the additional motion information.

[0011] According to the present disclosure, in the process of constructing a list of motion vector predictor candidates in an inter-prediction mode, the inter-frame prediction performance can be improved by including a candidate containing multiple motion vectors in the list and utilizing various motion vectors.

[0012] In addition, compression performance can be improved through a method of signaling weight information applied to multiple motion vectors and a method of limiting the use of interpolation filters to reduce over-smoothing caused by weighted sums between prediction blocks.

[0013] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0014] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0015] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.

[0016] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.

[0017] FIG. 4 shows an example of a video / image decoding method to which an embodiment of the present disclosure can be applied.

[0018] FIG. 5 shows an example of a video / image encoding method to which an embodiment of the present disclosure can be applied.

[0019] Figure 6 is a diagram showing an example of a search area used in intra-template matching.

[0020] FIGS. 7 and FIGS. 8 illustrate examples of inter-prediction-based video / image encoding methods to which embodiments of the present disclosure may be applied.

[0021] FIGS. 9 and FIGS. 10 illustrate examples of inter-prediction-based video / image decoding methods to which embodiments of the present disclosure may be applied.

[0022] FIG. 11 illustrates an exemplary inter-prediction procedure to which an embodiment of the present disclosure may be applied.

[0023] FIG. 12 is a diagram showing examples of blocks used to construct a merge candidate list.

[0024] Figure 13 is a diagram showing four movements that can be expressed in an affine motion model.

[0025] Figure 14 is a diagram showing an example of a control point motion vector used in affine motion prediction.

[0026] Figure 15 is an example showing the inheritance of control point motion vectors.

[0027] Figure 16 is a drawing showing an example of a surrounding block for the current block.

[0028] Figure 17 illustrates the process of SbTMVP.

[0029] Figure 18 illustrates GPM segments grouped at the same angle.

[0030] Figure 19 illustrates the left and upper neighbor blocks used in CIIP weight derivation.

[0031] Figure 20 exemplarily shows available IPM candidates for GPM including inter and intra predictions.

[0032] Figure 21 is a diagram illustrating a method for deriving motion vectors from template matching.

[0033] FIG. 22 is a diagram showing a case where multiple reference blocks are used in one embodiment.

[0034] FIGS. 23 and 24 are flowcharts illustrating an example of a method for including motion information candidates containing multiple motion information in a motion information candidate list, in a decoding method and an encoding method according to one embodiment.

[0035] FIG. 25 is a diagram showing an example of a method for configuring an MVP candidate list in one embodiment.

[0036] FIG. 26 is a diagram illustrating an example of a method for generating a candidate having multiple motion vectors according to one embodiment.

[0037] FIG. 27 is a diagram illustrating an example of a signaling / parsing method when a multi-reference block mode is included as one of the general merge modes in a method according to one embodiment.

[0038] FIGS. 28 and 29 are drawings illustrating examples of a signaling / parsing method of weight information applied to a multiple reference block in a method according to one embodiment.

[0039] FIGS. 30 to 32 are flowcharts illustrating examples of a method using an MVP index and a weight index in a decoding method or encoding method according to one embodiment.

[0040] FIG. 33 is a diagram illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.

[0041] The present disclosure is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals have been used for similar components in the description of each drawing.

[0042] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0043] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0044] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “having” are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0045] The present disclosure relates to video / video coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the VVC (versatile video coding) standard. Additionally, the methods / embodiments disclosed herein may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268).

[0046] This specification presents various embodiments regarding video / image coding, and unless otherwise noted, said embodiments may be performed in combination with one another.

[0047] In this specification, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image of a specific time period, and "slice" or "tile" is a unit that constitutes a part of a picture in coding. A slice or tile may contain one or more coding tree units (CTUs). A picture may consist of one or more slices or tiles. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of ​​CTUs having a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of ​​CTUs having a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged continuously according to the CTU raster scan, whereas tiles within a picture may be arranged continuously according to the tile's raster scan. A single slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively contained in a single NAL unit. Meanwhile, a single picture may be divided into two or more subpictures. A subpicture may be a rectangular area of ​​one or more slices within a picture.

[0048] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample generally represents a pixel or its value, and it may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component.

[0049] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0050] In this specification, “A or B” may mean “only A,” “only B,” or “both A and B.” Alternatively, in this specification, “A or B” may be interpreted as “A and / or B.” For example, in this specification, “A, B or C” may mean “only A,” “only B,” “only C,” or “any combination of A, B and C.”

[0051] As used herein, a slash ( / ) or a comma may mean “and / or.” For example, “A / B” may mean “A and / or B.” Accordingly, “A / B” may mean “only A,” “only B,” or “both A and B.” For example, “A, B, C” may mean “A, B or C.”

[0052] In this specification, “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” Additionally, in this specification, the expressions “at least one of A or B” or “at least one of A and / or B” may be interpreted as synonymous with “at least one of A and B.”

[0053] Additionally, in this specification, “at least one of A, B and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” may mean “at least one of A, B and C.”

[0054] Additionally, parentheses used in this specification may mean “for example.” Specifically, where indicated as “prediction (intra-prediction),” “intra-prediction” may be proposed as an example of “prediction.” In other words, “prediction” in this specification is not limited to “intra-prediction,” and “intra-prediction” may be proposed as an example of “prediction.” Furthermore, even when indicated as “prediction (i.e., intra-prediction),” “intra-prediction” may be proposed as an example of “prediction.”

[0055] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.

[0056] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0057] Referring to FIG. 1, the video / image coding system may include a first device (source device) and a second device (receiving device).

[0058] A source device can transmit encoded video / image information or data in the form of a file or streaming to a receiving device via a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.

[0059] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include a computer, a tablet, a smartphone, etc., and may generate video / images (electronically). For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.

[0060] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0061] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The storage medium may be a computer-readable storage medium and may store data non-transitory. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0062] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.

[0063] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0064] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.

[0065] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0066] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units (PU). For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure.

[0067] For example, a single coding unit may be divided into multiple coding units with a deeper depth based on a quad tree structure, a binary tree structure, and / or a terrestrial structure. In this case, for example, the quad tree structure may be applied first and the binary tree structure and / or terrestrial structure may be applied later. Alternatively, the binary tree structure may be applied before the quad tree structure. A coding procedure according to the present specification may be performed based on a final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of a lower depth so that a coding unit of the optimal size may be used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below.

[0068] As another example, the processing unit may further include a Prediction Unit (PU) or a Transform Unit (TU). In this case, the Prediction Unit and the Transform Unit may each be divided or partitioned from the aforementioned final coding unit. The Prediction Unit may be a unit for sample prediction, and the Transform Unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.

[0069] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used as a term corresponding to a pixel or pel of a picture (or image).

[0070] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).

[0071] The prediction unit (220) performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit (220) can generate various information regarding the prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding the prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0072] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples, i.e., the reference samples, may be located near the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and a plurality of directional modes. The non-directional mode may include at least one DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional modes may be used. The intra prediction unit (222) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0073] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. Temporal surrounding blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and a reference picture containing temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0074] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, such as SCC (screen content coding). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within the picture can be signaled based on information regarding the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal.

[0075] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously restored pixels. Additionally, the transformation process may be applied to a pixel block of the same size in a square, or to a block of variable size that is not square.

[0076] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.

[0077] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information required for video / image restoration (e.g., values ​​of syntax elements, etc.) together or separately, in addition to the quantized transform coefficients.

[0078] Encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream at the level of a Network Abstraction Layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).

[0079] Quantized transformation coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). An adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (221) or the intra-prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstruction unit or a reconstruction block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0080] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0081] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (200) and the decoding device, and can also improve encoding efficiency.

[0082] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information within the current picture was derived (or encoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (270) can store restoration samples of the blocks restored within the current picture and transmit them to the intra-prediction unit (222).

[0083] Video information output in the form of a bitstream from the encoding device (200) can be transmitted to the decoding device (300).

[0084] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.

[0085] Video information transmitted in the form of a bitstream from the encoding device (200) can be received by the decoding device (300).

[0086] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).

[0087] The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoding device chipset or processor) according to an embodiment. Additionally, the memory (360) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0088] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.

[0089] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based on information regarding the parameter sets and / or the general constraint information. The signaling / receiving information and / or syntax elements described below in this specification may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and the residual value for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310).

[0090] Meanwhile, the decoding device according to the present specification may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoding device (video / image / picture information decoding device) and a sample decoding device (video / image / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).

[0091] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.

[0092] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0093] The prediction unit (320) can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. The prediction unit (320) can determine whether an intra prediction or an inter prediction is applied to the current block based on information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter prediction mode.

[0094] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.

[0095] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block according to the prediction mode, or may be located at a certain distance from the current block. In intra prediction, the prediction modes may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0096] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the inter-prediction mode for the current block.

[0097] The adder (340) can generate a restoration signal (restoration picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the restoration block.

[0098] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0099] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (360), specifically to the DPB of memory (360). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0100] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter-prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (332) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (331).

[0101] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.

[0102] FIG. 4 shows an example of a video / image decoding method to which an embodiment of the present disclosure can be applied.

[0103] In video coding, the pictures constituting the video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, and based on this, not only forward prediction but also reverse prediction can be performed during inter-prediction.

[0104] In FIG. 4, S400 may be performed in the entropy decoding unit (310) of the aforementioned decoding device (300), S410 may be performed in the prediction unit (330), S420 may be performed in the residual processing unit (320), S430 may be performed in the addition unit (340), and S440 may be performed in the filtering unit (350). S400 may include a decoding procedure according to the present disclosure, S410 may include an inter / intra prediction procedure according to the present disclosure, S420 may include a residual processing procedure according to the present disclosure, S430 may include a block / picture restoration procedure according to the present disclosure, and S440 may include an in-loop filtering procedure according to the present disclosure.

[0105] Referring to FIG. 4, the decoding device acquires image / video information from a bitstream (S400), performs a prediction based on the acquired image / video information (S410), and can restore a picture through residual processing (S420, inverse quantization and inverse transformation of the quantized transformation coefficients) (S430).

[0106] A modified restored picture can be generated by applying an in-loop filtering procedure (S440) to the restored picture generated through the above restoration procedure, and the modified restored picture can be output as a decoded picture and also stored in the buffer or memory of the decoding device to be used as a reference picture in the inter-prediction procedure when decoding the next picture. In some cases, the above in-loop filtering procedure may be omitted, in which case the restored picture can be output as a decoded picture and also stored in the buffer or memory of the decoding device to be used as a reference picture in the inter-prediction procedure when decoding a subsequent picture.

[0107] The in-loop filtering procedure (S440) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, and some or all of these may be omitted. Additionally, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture. Alternatively, for example, the ALF procedure may be performed after the deblocking filtering procedure is applied to the restored picture. This may be performed in the same manner in the encoding device.

[0108] FIG. 5 shows an example of a video / image encoding method to which an embodiment of the present disclosure can be applied.

[0109] In FIG. 5, the prediction step (S500) may be performed in the prediction unit (220) of the aforementioned encoding device (200), residual processing (S510) based on the prediction result may be performed in the residual processing unit (230), and the step (S520) of encoding image information including prediction information and residual information may be performed in the entropy encoding unit (240). S500 may include an inter / intra prediction procedure according to the present disclosure, S510 may include a residual processing procedure according to the present disclosure, and S520 may include an encoding procedure according to the present disclosure.

[0110] The encoding procedure may optionally include not only a procedure for encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, but also a procedure for generating a restored picture for the current picture and a procedure for applying in-loop filtering to the restored picture.

[0111] The encoding device (200) can derive (modified) residual samples from quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), and can generate a restored picture based on the (modified) residual samples and the predicted samples which are the outputs of S500. The restored picture thus generated may be identical to the restored picture generated by the decoding device (300) described above. A modified restored picture may be generated through an in-loop filtering procedure on the restored picture, which may be stored in a buffer or memory, and, as in the case of the decoding device, may be used as a reference picture in the inter-prediction procedure during the subsequent encoding of the picture.

[0112] As described above, depending on the case, part or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) may be encoded in the entropy encoding unit (240) and output in the form of a bitstream, and the decoding device (300) may perform the in-loop filtering procedure in the same way as the encoding device based on the filtering-related information.

[0113] Through this in-loop filtering procedure, noise generated during video / image coding, such as blocking artifacts and ringing artifacts, can be reduced, and subjective / objective image quality can be improved. In addition, by performing the in-loop filtering procedure in both the encoding device (200) and the decoding device (300), the same prediction results can be derived in both the encoding device (200) and the decoding device (300), the reliability of picture coding can be increased, and the amount of data that must be transmitted for picture coding can be reduced.

[0114] As described above, the picture restoration procedure can be performed in the encoding device (200) as well as the decoding device (300). Restoration blocks can be generated based on intra prediction / inter prediction for each block unit, and a restored picture containing the restoration blocks can be generated. If the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based solely on intra prediction. Meanwhile, if the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra prediction or inter prediction. In this case, inter prediction may be applied to some blocks within the current picture / slice / tile group, and intra prediction may be applied to the remaining blocks.

[0115] The color component of the picture may include a luminance component and a chroma component, and unless explicitly limited in the present disclosure, embodiments according to the present disclosure may be applied to the luminance component and the chroma component.

[0116] Meanwhile, when intra prediction is performed, the prediction unit (220, 330) of the encoding device (200) / decoding device (300) can derive a reference sample according to the intra prediction mode of the current block among the surrounding samples of the current block, and can generate a prediction sample of the current block based on the reference sample.

[0117] For example, (i) a prediction sample can be derived based on the average or interpolation of neighboring reference samples of the current block, and (ii) the prediction sample can be derived based on reference samples existing in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. Case (i) can be called a non-directional mode or non-angular mode, and case (ii) can be called a directional mode or angular mode.

[0118] In addition, Linear interpolation intra prediction (LIP) may be applied to perform intra prediction on the current block by linearly interpolating prediction sample values ​​generated based on the intra prediction mode of the current block.

[0119] In addition, a provisional prediction sample of the current block may be derived based on filtered surrounding reference samples, and a prediction sample of the current block may be derived by performing a weighted sum of the provisional prediction sample and at least one reference sample derived according to the intra prediction mode among the existing surrounding reference samples, that is, unfiltered surrounding reference samples. This prediction may be referred to as PDPC (Position Dependent intra Prediction Combination).

[0120] In addition, intra-prediction coding can be performed by selecting the reference sample line with the highest prediction accuracy among the surrounding multiple reference sample lines of the current block, deriving a prediction sample using a reference sample located in the prediction direction from that line, and signaling the used reference sample line to the decoding device. This case can be referred to as multi-reference line intra prediction (MRL) or MRL-based intra prediction.

[0121] In addition, the current block can be divided into vertical or horizontal subpartitions to perform intra prediction based on the same intra prediction mode, while utilizing surrounding reference samples derived at the subpartition level. That is, in this case, the intra prediction mode for the current block is applied equally to the subpartitions, but intra prediction performance can be improved depending on the circumstances by deriving and utilizing surrounding reference samples at the subpartition level. This prediction method may be referred to as intra sub-partitions (ISP) or ISP-based intra prediction.

[0122] In addition, when the prediction direction based on the prediction sample points among surrounding reference samples, that is, when the prediction direction points to a fractional sample location, the value of the prediction sample can also be derived through the interpolation of multiple reference samples located around the prediction direction (around the fractional sample location).

[0123] The MPM list for deriving the aforementioned intra prediction mode may be configured differently depending on the intra prediction type. Alternatively, the MPM list may be configured commonly regardless of the intra prediction type.

[0124] SGPM (Spatial Geometric Partitioning Mode) is an intra mode similar to the intercoding tool of GPM, where two prediction parts are generated in the intra prediction process. In this mode, a candidate list containing one partition split and two intra prediction modes is created for each entry, and the partition mode and three intra prediction modes are used to form combinations. The length of the candidate list can be set to 16, and selected candidate indices can be signaled.

[0125] The candidate list is reordered using a template, where the SAD between the template's prediction and restoration is used for sorting. The template size can be fixed at 1.

[0126] For each partition mode, an IPM list for each part is derived using the same intra-inter GPM list derivation. The size of the IPM list can be set to 3. In the list, the TIMD derivation mode can be replaced with two derivation modes in the horizontal and vertical directions.

[0127] SGPM mode can be applied with limited block sizes as follows:

[0128] 4<=width<=64, 4<=height<=64, width <height*8, height<width*8, width*height> =32

[0129] Adaptive blending is also used in spatial GPM, and the blending depth τ can be derived as follows:

[0130] - If min(width, height)==4, 1 / 2 τ is selected

[0131] - else if min(width, height)==8, τ is selected

[0132] - else if min(width, height)==16, 2 τ is selected

[0133] - else if min(width, height)==32, 4 τ is selected

[0134] - else, 8 τ is selected

[0135] Intra-block copy (IBC) is a method that can significantly improve the coding efficiency of screen content materials. Since the IBC mode is implemented as a block-level coding mode, block matching (BM) can be performed in the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector can be used to represent the displacement from the current block to a reference block already reconstructed within the current block. The luminance block vector of an IBC-coded CU can be of integer resolution (or precision). The chroma block vector can be rounded to integer resolution. When combined with Adaptive Motion Vector Resolution (AMVR), the IBC mode can switch between 1-Pel (pixel) and 4-Pel (pixel) motion vector resolutions. An IBC-coded CU can be treated as a third prediction mode rather than intra or inter prediction modes. The IBC mode can be applied to CUs with both width and height of 64 luminance samples or less.

[0136] Hash-based motion estimation can be performed on the encoder side for IBC. The encoder can perform BD checks on blocks whose width and height are not greater than 16 luma samples. For non-merge mode, block vector search can be performed using hash-based search first. If hash search does not return a valid candidate, block matching-based local search can be performed.

[0137] In hash-based search, hash key matching (32-bit CRC) between the current block and the reference block can be extended to any allowed block size. The hash key calculation for all locations in the current picture is based on 4x4 sub-blocks. For a larger current block, if all hash keys of the 4x4 sub-blocks match the hash keys of the corresponding reference locations, the hash key can be determined to match that of the reference block.

[0138] If the hash keys of multiple predicted blocks are found to match those of the current block, the block vector cost of each matched reference is calculated, and the minimum cost can be selected.

[0139] In block matching search, the search range can be set to cover both previous and current CTUs. At the CU level, the IBC mode is signaled by a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as follows:

[0140] - IBC Skip / Merge Mode: The merge candidate index can be used to indicate which block vectors are used to predict the current block from a list of neighboring candidate IBC-coded blocks. The merge list can include space, HMVP, and pairwise candidates.

[0141] - IBC AMVP Mode: Block vector differencing can be coded in the same way as motion vector differencing. Block vector prediction methods can use two candidates as predictors (if IBC coded): one from the left neighbor and one from the upper neighbor. If neither neighbor is available, a default block can be used as the predictor. A flag can be signaled to indicate the block vector predictor index.

[0142] Intra Template Matching Prediction (IntraTMP) is a special intra prediction mode that copies the optimal prediction block from the reconstructed portion of the current frame where an L-shaped template matches the current template. Within a predefined search range, the encoder searches the reconstructed region of the current frame for the template most similar to the current template and uses that block as the prediction block. Subsequently, the encoder signals the use of this mode, and the same prediction operation is performed on the decoder side.

[0143] Figure 6 is a diagram showing an example of a search area used in intra-template matching.

[0144] The prediction signal is generated by matching other blocks within the predefined search areas of FIG. 6 with causal neighbors using the L-shape, top-only, or left-only of the current block. As illustrated in FIG. 6, there may be a total of six predefined search areas (i.e., R1–R6), which include not only some of the reconstructed samples within the current CTU located at the top, left, bottom-left, and top-right of the current block, but also samples reconstructed from the top CTU and left CTU:

[0145] The Sum of Absolute Differences (SAD) is used as a cost function.

[0146] The given search order of the six search regions is utilized (i.e., R4, R5, R6, R1, R2, R3). Within each region, the decoder generates a list of up to 19 template-matching block vector candidates sorted in ascending order according to the Template Cost (SAD). The supported modes are as follows:

[0147] 1. Single Predictor: A single predictor is selected from the list of candidates.

[0148] 2. Fusion of Multiple Predictors: Multiple predictors are fused to derive the final prediction block. Fusion weights can be calculated based on the template matching cost of each predictor, or a weight derivation method based on a Wiener filter can be used.

[0149] 3. Sub-pixel Precision: When using a single predictor, it supports 1 / 2 pixel, 1 / 4 pixel, and 3 / 4 pixel precision, providing 8 directions each.

[0150] 4. Linear Filter Model: Applies a linear filter learned between the reference template and the current template to the reference block. This mode can be applied to a single predictor that does not use subpixel precision.

[0151] To ensure a fixed number of SAD comparisons per pixel, the size of all search ranges (SearchRange_w, SearchRange_h) is set proportionally to the block size (BlkW, BlkH). That is:

[0152] SearchRange_w = min(64,a * BlkW)

[0153] SearchRange_h = min(64,a * BlkH)

[0154] Here, 'a' is a constant that controls the trade-off between gain and complexity, and can be set to a = 5.

[0155] To accelerate the template matching process, the search range of all search areas can be subsampled by a factor of three. After finding the optimal match, a refinement process is performed. This refinement is carried out through a second template matching search around the optimal match for the reduced range.

[0156] Intra-template matching can be enabled in CUs with a width and height of 64 or less. The maximum CU size for intra-template matching can be set.

[0157] The intra-template matching prediction mode can be signaled at the CU level via a dedicated flag when DIMD is not used for the current CU.

[0158] Meanwhile, when inter-prediction is applied, the prediction unit of the encoding device / decoding device can perform inter-prediction on a block-by-block basis to derive prediction samples. Inter-prediction may represent a prediction derived in a manner dependent on data elements (e.g., sample values, or motion information) of picture(s) other than the current picture. When inter-prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on the reference picture pointed to by the reference picture index.

[0159] At this time, to reduce the amount of motion information transmitted in inter-prediction mode, motion information of the current block can be predicted in block, sub-block, or sample units based on the correlation of motion information between surrounding blocks and the current block. The motion information may include motion vectors and / or reference picture indices. The motion information may further include inter-prediction type information (L0 prediction, L1 prediction, Bi prediction, etc.). When inter-prediction is applied, surrounding blocks may include spatial surrounding blocks existing within the current picture and temporal surrounding blocks existing in the reference picture.

[0160] The reference picture containing the above reference block and the reference picture containing the above temporal surrounding block may be the same or different. The above temporal surrounding block may be called by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal surrounding block may be called collocated picture (colPic). For example, a list of motion information candidates may be constructed based on the surrounding blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.

[0161] Inter-prediction can be performed based on various prediction modes; for example, in the case of skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected surrounding block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected surrounding block is used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.

[0162] The above motion information may include L0 motion information and / or L1 motion information depending on the inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a pair (Bi) prediction. Here, the L0 motion vector may represent a motion vector associated with reference picture list L0 (L0), and the L1 motion vector may represent a motion vector associated with reference picture list L1 (L1). Reference picture list L0 may include pictures that are prior to the current picture in output order as reference pictures, and reference picture list L1 may include pictures that are subsequent to the current picture in output order. The aforementioned previous pictures may be called forward (reference) pictures, and the aforementioned subsequent pictures may be called reverse (reference) pictures.

[0163] The above reference picture list L0 may include additional pictures as reference pictures that are output later than the current picture. In this case, within the reference picture list L0, the previous pictures may be indexed first and the subsequent pictures may be indexed next. The above reference picture list L1 may include additional pictures as reference pictures that are output earlier than the current picture. In this case, within the reference picture list L1, the subsequent pictures may be indexed first and the previous pictures may be indexed next. Here, the output order may correspond to the picture order count (POC) order.

[0164] FIGS. 7 and FIGS. 8 illustrate examples of inter-prediction-based video / image encoding methods to which embodiments of the present disclosure may be applied.

[0165] Referring to FIG. 7, the encoding device (200) can perform inter prediction for the current block (S600). The encoding device can derive the inter prediction mode and motion information of the current block and generate prediction samples of the current block. Here, the procedures for determining the inter prediction mode, deriving motion information, and generating prediction samples may be performed simultaneously, or one procedure may be performed before the other. For example, as shown in FIG. 8, the inter prediction unit (221) of the encoding device (200) may include a prediction mode determination unit (221a), a motion information derivation unit (221b), and a prediction sample derivation unit (221c), and the prediction mode determination unit (221a) may determine the prediction mode for the current block, the motion information derivation unit (221b) may derive motion information of the current block, and the prediction sample derivation unit (221c) may derive prediction samples of the current block.

[0166] For example, the inter-prediction unit of the encoding device can search for a block similar to the current block within a certain area (search area) of reference pictures through motion estimation, and derive a reference block whose difference from the current block is minimal or below a certain standard. Based on this, it can derive a reference picture index pointing to the reference picture where the reference block is located, and derive a motion vector based on the positional difference between the reference block and the current block. The encoding device can determine a mode applied to the current block among various prediction modes. The encoding device can compare the RD costs for the various prediction modes and determine the optimal prediction mode for the current block.

[0167] For example, when a skip mode or merge mode is applied to the current block, the encoding device may construct a merge candidate list described below and derive a reference block among the reference blocks pointed to by the merge candidates included in the merge candidate list, wherein the difference between the sample of the current block, i.e., the difference in SAD or SATD, is minimum or below a certain standard. In this case, a merge candidate associated with the derived reference block is selected, and merge index information pointing to the selected merge candidate is generated and can be signaled to the decoding device. Movement information of the current block can be derived using the movement information of the selected merge candidate.

[0168] As another example, when the (A)MVP mode is applied to the current block, the encoding device may construct the (A)MVP candidate list described below and use the motion vector of a selected motion vector predictor candidate among the motion vector predictor (mvp) candidates included in the (A)MVP candidate list as the motion vector predictor for the current block. In this case, for example, a motion vector pointing to a reference block derived by the motion estimation described above may be used as the motion vector for the current block, and among the motion vector predictor candidates, the motion vector predictor candidate having the smallest difference from the motion vector of the current block may become the selected motion vector predictor candidate. A Motion Vector Difference (MVD), which is the difference obtained by subtracting the motion vector predictor from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the decoding device. Additionally, when the (A)MVP mode is applied, the value of the reference picture index may be composed of reference picture index information and separately signaled to the decoding device.

[0169] The encoding device can derive residual samples based on predicted samples (S610). The encoding device can derive residual samples by comparing the original samples of the current block with the predicted samples.

[0170] The encoding device can encode image information including prediction information and residual information (S620). The encoding device can output the encoded image information in the form of a bitstream. The prediction information is information regarding prediction and may include prediction mode information (e.g., skip flag, merge flag, or merge index, etc.) and / or motion information. The motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. Additionally, the motion information may include information regarding the aforementioned MVD and / or reference picture index information. Additionally, the motion information may include information indicating whether L0 prediction, L1 prediction, or paired (bi) prediction is applied. Residual information is information regarding residual samples. The residual information may include information regarding quantized transform coefficients for residual samples.

[0171] The output bitstream can be stored in a (digital) storage medium and transmitted to a decoding device, or it can be transmitted to a decoding device via a network.

[0172] Meanwhile, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is intended to induce the same prediction results in the encoding device as those performed in the decoding device, thereby increasing coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and utilize it as a reference picture for inter-prediction. As previously described, in-loop filtering procedures, etc., may be further applied to the reconstructed picture.

[0173] FIGS. 9 and FIGS. 10 illustrate examples of inter-prediction-based video / image decoding methods to which embodiments of the present disclosure may be applied.

[0174] A video / image decoding procedure based on inter-prediction may include, for example, the following in general.

[0175] Referring to FIGS. 9 and 10, the decoding device (300) can perform an operation corresponding to the operation performed by the encoding device (200). The decoding device can perform a prediction for the current block based on the received prediction information and derive prediction samples.

[0176] Specifically, the decoding device can determine a prediction mode for the current block based on the received prediction information (S700). The prediction mode determination unit (332a) of the decoding device (300) can determine which inter-prediction mode is applied to the current block based on the prediction mode information within the prediction information.

[0177] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block or whether (A)MVP mode is determined. Alternatively, one of various inter-prediction mode candidates can be selected based on the mode index. The inter-prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter-prediction modes described below.

[0178] The decoding device can derive movement information of the current block based on the determined inter-prediction mode (S710). For example, the movement information derivation unit (332b) of the decoding device (300) can construct a merge candidate list described below when a skip mode or merge mode is applied to the current block, and select one merge candidate among the merge candidates included in the merge candidate list. This selection can be performed based on the selection information (merge index) described above. The movement information of the current block can be derived using the movement information of the selected merge candidate. The movement information of the selected merge candidate can be used as the movement information of the current block.

[0179] As another example, when an (A)MVP mode is applied to the current block, the decoding device may construct the (A)MVP candidate list described below and use the motion vector of a selected MVP candidate among the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. This selection may be performed based on the selection information (MVP flag or MVP index) described above. In this case, the MVD of the current block can be derived based on information regarding the MVD, and the motion vector of the current block can be derived based on the MVP and MVD of the current block. Additionally, the reference picture index of the current block can be derived based on reference picture index information. Within the reference picture list regarding the current block, the picture pointed to by the reference picture index can be derived as the reference picture referenced for inter-prediction of the current block.

[0180] Meanwhile, movement information of the current block may be induced without constructing a candidate list, in which case the movement information of the current block may be induced according to the procedure initiated in the prediction mode. In this case, the construction of the candidate list as described above may be omitted.

[0181] The decoding device can generate prediction samples for the current block based on the movement information of the current block (S720). In this case, the prediction sample derivation unit (332c) of the decoding device (300) can derive a reference picture based on the reference picture index of the current block and derive prediction samples for the current block using samples of the reference block that the movement vector of the current block points to on the reference picture. In this case, as described below, a prediction sample filtering procedure may be further performed on all or some of the prediction samples of the current block depending on the case.

[0182] In other words, the inter-prediction unit (332) of the decoding device (300) may include a prediction mode determination unit (332a), a motion information induction unit (332b), and a prediction sample induction unit (332c). The prediction mode determination unit (332a) may determine a prediction mode for the current block based on the prediction mode information received, the motion information induction unit (332b) may induce motion information (motion vector and / or reference picture index, etc.) for the current block based on the information regarding motion information received, and the prediction sample induction unit (332c) may induce or generate prediction samples for the current block.

[0183] The decoding device generates residual samples for the current block based on the received residual information (S730). The decoding device (300) generates restoration samples for the current block based on the prediction samples and residual samples, and can generate a restoration picture based thereon (S740). Subsequently, an in-loop filtering procedure, etc., may be further applied to the restoration picture as described above.

[0184] FIG. 11 illustrates an exemplary inter-prediction procedure to which an embodiment of the present disclosure may be applied.

[0185] Referring to FIG. 11, as described above, the inter prediction procedure (S600) may include a step of determining an inter prediction mode, a step of deriving motion information according to the determined prediction mode, and a step of performing a prediction based on the deriving motion information (generating a prediction sample). As described above, the inter prediction procedure may be performed in an encoding device and a decoding device. In this document, the term "coding device" may include an encoding device and / or a decoding device.

[0186] Referring to FIG. 11, the coding device determines an inter-prediction mode for the current block (S800). Various inter-prediction modes may be used for the prediction of the current block within the picture. For example, various modes such as merge mode, skip mode, Motion Vector Prediction (MVP) mode, affine mode, sub-block merge mode, and MMVD (Merge with MVD) mode may be used. Decoder-side Motion Vector Refinement (DMVR) mode, Adaptive Motion Vector Resolution (AMVR) mode, Bi-prediction with CU-level Weight (BCW), and Bi-Directional Optical Flow (BDOF) may be used as auxiliary modes. Additionally, according to one embodiment, the above-described inter-prediction mode may include a Multi-Hypethesis Prediction (MHP) mode. The Multi-Hypethesis Prediction mode represents a method of performing prediction by weighting additional prediction blocks generated based on additional motion information to the inter-prediction block. The Multi-Hypethesis Prediction mode will be described in detail later.

[0187] In the present disclosure, an affine mode may be referred to as an affine motion prediction mode. Additionally, an MVP mode may be referred to as an AMVP (Advanced Motion Vector Prediction) mode. In the present disclosure, some modes and / or motion information candidates derived by some modes may be included as one of the motion information related candidates of other modes. For example, an HMVP candidate may be added as a merge candidate for a merge / skip mode, or may be added as a motion vector predictor candidate for an AMVP mode. When an HMVP candidate is used as a motion information candidate for a merge mode or a skip mode, the HMVP candidate may be referred to as an HMVP merge candidate.

[0188] Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be received by the decoding device by being included in a bitstream. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter-prediction mode may be indicated through hierarchical signaling of flag information.

[0189] In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether skip mode is applied, and if skip mode is not applied, a merge flag may be signaled to indicate whether merge mode is applied, and if merge mode is not applied, MVP mode may be applied, or additional flags for further distinction may be signaled. The affine mode may be signaled as an independent mode, or as a mode dependent on merge mode or MVP mode, etc. For example, the affine mode may include affine merge mode and affine MVP mode.

[0190] The coding device can derive motion information for the current block (S810). The motion information can be derived based on the inter prediction mode determined in the aforementioned step. The coding device can perform inter prediction using the motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure.

[0191] For example, an encoding device can use the original block within the original picture for the current block to search for highly correlated similar reference blocks in fractional pixel units within a defined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference between phase-based sample values. For example, block similarity can be calculated based on the SAD between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information can be signaled to a decoding device according to various methods based on inter-prediction modes.

[0192] The coding device can generate prediction samples by performing inter-prediction based on movement information for the current block (S820). The current block containing the prediction samples may be called a prediction block.

[0193] Meanwhile, information indicating whether the aforementioned List0 (L0) prediction, list1 (L1) prediction, or bi-prediction is being used in the current block (current coding unit) may be signaled. This information may be referred to as motion prediction direction information, inter prediction direction information, or inter prediction indication information, and may be configured / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element may indicate whether the aforementioned List0 (L0) prediction, List1 (L1) prediction, or bi-prediction is being used in the current block (current coding unit). For convenience of explanation in this document, the inter prediction type (L0 prediction, L1 prediction, or BI prediction) pointed to by the inter_pred_idc syntax element may be indicated as the motion prediction direction. L0 prediction may be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, depending on the value of the inter_pred_idc syntax element, prediction types such as those shown in Table 1 below can be represented.

[0194] [Table 1]

[0195]

[0196] As described above, a picture may contain one or more slices. A slice may have one of the slice types, including intra (I) slices, predictive (P) slices, and bi-predictive (B) slices. These slice types may be indicated based on slice type information. For blocks within an I slice, only intra prediction may be used for prediction, and no inter prediction may be used. Of course, even in this case, the original sample values ​​may be coded and signaled without prediction. For blocks within a P slice, either intra prediction or inter prediction may be used; if inter prediction is used, only uni prediction may be used. Meanwhile, for blocks within a B slice, either intra prediction or inter prediction may be used; if inter prediction is used, up to bi-prediction may be used.

[0197] L0 and L1 may contain reference pictures that were encoded / decoded prior to the current picture. For example, L0 may contain reference pictures that are prior to and / or later than the current picture in the POC order, and L1 may contain reference pictures that are later and / or earlier than the current picture in the POC order. In this case, L0 may be assigned a reference picture index relatively lower than the reference pictures prior to the current picture in the POC order, and L1 may be assigned a reference picture index relatively lower than the reference pictures later than the current picture in the POC order. For B slices, paired prediction may be applied, and in this case, either unidirectional paired prediction or bidirectional paired prediction may be applied. Bidirectional paired prediction may be referred to as true paired prediction.

[0198] Inter prediction can be performed using motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block within the original picture for the current block to search for similar reference blocks with high correlation in fractional pixel units within a defined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference between phase-based sample values. For example, block similarity can be calculated based on the SAD between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information can be signaled to the decoding device according to various methods based on the inter prediction mode.

[0199] When merge mode is applied, the movement information of the current prediction block is not transmitted directly; instead, the movement information of surrounding prediction blocks is used to induce the movement information of the current prediction block. Therefore, the movement information of the current prediction block can be indicated by transmitting flag information indicating that merge mode has been used and a merge index indicating which surrounding prediction blocks were utilized. The above merge mode may also be referred to as regular merge mode.

[0200] To perform a merge mode, the encoder may search for merge candidate blocks to derive movement information of the current prediction block. For example, up to five merge candidate blocks may be used, but the present invention is not limited thereto. Also, the maximum number of merge candidate blocks may be transmitted in the slice header or tile group header, but the present invention is not limited thereto. After finding the merge candidate blocks, the encoder may generate a merge candidate list and select the merge candidate block with the smallest cost among them as the final merge candidate block.

[0201] FIG. 12 is a diagram showing examples of blocks used to construct a merge candidate list.

[0202] The present invention provides various embodiments of a merge candidate block that constitutes a merge candidate list.

[0203] For example, the merge candidate list may use five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate may be used. As a specific example, for the spatial merge candidate, the blocks shown in FIG. 12 may be used as spatial merge candidates. Hereinafter, the spatial merge candidate or the spatial MVP candidate described below may be referred to as SMVP, and the temporal merge candidate or the temporal MVP candidate described below may be referred to as TMVP.

[0204] The merge candidate list for the current block can be constructed, for example, based on the following procedure:

[0205] - Insert spatial merge candidates derived by exploring spatial neighboring blocks into the merge candidate list

[0206] - Insert temporal merge candidates derived by searching temporal neighboring blocks into the merge candidate list

[0207] - Compare the current number of merge candidates with the maximum number of merge candidates

[0208] - If the current number of merge candidates is less than the maximum number of merge candidates, insert additional merge candidates into the merge candidate list.

[0209] To explain the aforementioned procedure in detail, the coding device (encoding device / decoding device) searches for spatial surrounding blocks of the current block and inserts the derived spatial merge candidates into a merge candidate list. For example, the spatial surrounding blocks may include the blocks surrounding the lower-left corner of the current block, the left surrounding block, the upper-right corner surrounding block, the upper surrounding block, and the upper-left corner surrounding block. However, as this is an example, additional surrounding blocks such as the right surrounding block, the lower surrounding block, and the lower-right surrounding block may be used as the spatial surrounding blocks in addition to the spatial surrounding blocks described above. The coding device can search for available blocks based on priority to detect available blocks and derive movement information of the detected blocks as spatial merge candidates. For example, the encoder and decoder can search the five blocks shown in FIG. 12 in the order A1, B1, B0, A0, and B2, and sequentially index the available candidates to form a merge candidate list.

[0210] The coding device searches for temporal neighbor blocks of the current block and inserts the derived temporal merge candidates into the merge candidate list. Temporal neighbor blocks may be located on a reference picture that is different from the current picture where the current block is located. The reference picture where the temporal neighbor blocks are located may be called a collocated picture or a col picture. Temporal neighbor blocks may be searched in the order of the lower-right corner neighbor blocks and the lower-right center blocks of the co-located blocks for the current block on the col picture.

[0211] Meanwhile, when motion data compression is applied, specific motion information can be stored as representative motion information for each fixed storage unit in the col picture. In this case, there is no need to store motion information for all blocks within the fixed storage unit, thereby achieving the effect of motion data compression. In this case, the fixed storage unit may be predetermined, for example, as a 16x16 sample unit or an 8x8 sample unit, or size information regarding the fixed storage unit may be signaled from the encoder to the decoder. When motion data compression is applied, the motion information of temporal neighbor blocks can be replaced with the representative motion information of the fixed storage unit where the temporal neighbor blocks are located. That is, from an implementation perspective, in this case, rather than the prediction block located at the coordinates of the temporal neighbor block, a temporal merge candidate can be derived based on the motion information of a prediction block covering a position that is arithmetically right-shifted by a fixed value based on the coordinates of the temporal neighbor block (top-left sample position) and then arithmetically left-shifted. For example, if the above fixed storage unit is 2 n x2 nIn the case of sample units, if the coordinates of the temporal neighbor block are (xTnb, yTnb), then the modified position is ((xTnb>>n)<<n), (yTnb> >n)< <n))에 위치하는 예측 블록의 움직임 정보가 시간적 머지 후보를 위하여 사용될 수 있다. 구체적으로 예를 들어, 일정 저장 단위가 16x16 샘플 단위인 경우, 시간적 주변 블록의 좌표가 (xTnb, yTnb)라 하면, 수정된 위치인 ((xTnb> Movement information of prediction blocks located at (yTnb>>4)<<4)) can be used for temporal merge candidates. Or, for example, if a certain storage unit is an 8x8 sample unit, and the coordinates of temporal neighbor blocks are (xTnb, yTnb), movement information of prediction blocks located at modified positions ((xTnb>>3)<<3)) and (yTnb>>3)<<3)) can be used for temporal merge candidates.

[0212] The coding device can compare the current number of merge candidates with the maximum number of merge candidates. The maximum number of merge candidates can be predefined or signaled from the encoder to the decoder. For example, the encoder can generate information regarding the maximum number of merge candidates, encode it, and transmit it to the decoder as a bitstream. Once the maximum number of merge candidates is reached, the process of adding subsequent candidates may not proceed.

[0213] If, as a result of comparison, the number of current merge candidates is less than the maximum number of merge candidates, the coding device inserts additional merge candidates into the merge candidate list. The additional merge candidates may include at least one of, for example, the history-based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate (when the slice / tile group type of the current slice / tile group is of type B), and / or zero-vector merge candidate described below.

[0214] If, as a result of comparison, the number of current merge candidates is not less than the maximum number of merge candidates, the coding device may terminate the construction of the merge candidate list. In this case, the encoder may select the optimal merge candidate among the merge candidates that make up the merge candidate list based on the rate-distortion (RD) cost, and may signal selection information (e.g., merge index) pointing to the selected merge candidate to the decoder. The decoder may select the optimal merge candidate based on the merge candidate list and the selection information.

[0215] As previously described, the motion information of the selected merge candidate can be used as the motion information of the current block, and prediction samples of the current block can be derived based on the motion information of the current block. The encoder can derive residual samples of the current block based on the prediction samples and can encode residual information regarding the residual samples and transmit it to the decoder. As previously described, the decoder can generate reconstructed samples based on the residual samples derived from the transmitted residual information and the prediction samples, and generate a reconstructed picture based thereon.

[0216] When skip mode is applied, movement information of the current block can be derived in the same way as when merge mode is applied. However, when skip mode is applied, the residual signal for the corresponding block is omitted, and therefore, the predicted samples can be directly used as reconstruction samples.

[0217] History-based merge candidate derivation (HMVP) merge candidates can be added to the merge list following spatial merge candidates and temporal merge candidates. In this method, movement information of previously encoded blocks is stored in a table and used as the MVP of the current coding unit (CU). A table containing multiple HMVP candidates is maintained during the encoding and decoding process. The table is initialized (emptied) when a new coding tree unit (CTU) row begins. Whenever there is a non-subblock encoded CU, the associated movement information is added to the last entry of the table as a new HMVP candidate.

[0218] In VVC, the size S of the HMVP table is set to 5, indicating that a maximum of 5 history-based MVP (HMVP) candidates can be added to the table. When inserting a new move candidate into the table, a restricted First-In, First-Out (FIFO) rule is applied, and a duplicate check is performed first to verify whether the same HMVP exists in the table. If a duplicate HMVP is found, that HMVP is removed from the table, and all subsequent HMVP candidates are moved forward.

[0219] HMVP candidates can be used in the process of constructing the merge candidate list. Multiple recent HMVP candidates within the table are examined sequentially and inserted into the merge candidate list after TMVP candidates. Spatial or temporal duplicate checks with merge candidates are also performed for HMVP candidates.

[0220] To reduce the number of redundant check operations, the following simplification is introduced:

[0221] The number of HMVP candidates used to generate the merge list is set to M if (N ≤ 4) and (8 - N) otherwise, where N represents the number of existing candidates in the merge list and M represents the number of available HMVP candidates in the table.

[0222] When the total number of possible merge candidates reaches the value of the maximum allowed merge candidate number minus 1, the process of constructing the merge candidate list using HMVP ends.

[0223] In this specification, a pair-wise average merge candidate may be referred to as a pair-wise average candidate or a pair-wise candidate. A pair-wise average candidate is generated by averaging predefined pair-wise candidates within an existing merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}. Here, the numbers represent merge indices within the merge candidate list.

[0224] The average motion vector is calculated individually for each reference list. If both motion vectors are available in a single reference list, the average is calculated even if the two motion vectors point to different reference pictures. If only one motion vector is available, that vector is used directly. If no motion vectors are available, the corresponding list remains invalid.

[0225] If the merge list is not filled even after the pairwise average merge candidates are added, zero MVPs are inserted at the end of the list until the maximum number of merge candidates is reached.

[0226] The previously mentioned MMVD mode (Merge mode with MVD) is explained. In addition to the merge mode, which directly uses implicitly derived motion information in the prediction samples of the current block, a merge mode utilizing the difference in motion vectors (MMVD) can be used. Since similar motion information derivation methods are used in both skip mode and merge mode, MMVD can also be applied to skip mode. After signaling the skip flag and merge flag, the MMVD flag information (e.g., mmvd_flag) can be signaled to indicate whether to use MMVD mode for the current block.

[0227] In MMVD mode, after a merge candidate is selected, the candidate may be further refined based on the signaled MVD information. When MMVD is applied to the current block (i.e., when mmvd_flag is 1), additional information regarding the MMVD may be signaled. This additional information may include a merge candidate flag (e.g., mmvd_merge_flag) indicating whether the first or second candidate in the merge candidate list is used with the motion vector difference, a distance index (e.g., mmvd_distance_idx) indicating the motion magnitude, and a direction index (e.g., mmvd_direction_idx) indicating the motion direction. In MMVD mode, one of the first two candidates in the merge list may be selected to be used as the MV standard. The merge candidate flag is signaled to indicate which candidate is being used.

[0228] The distance index specifies movement magnitude information and represents a predefined offset from the starting point. The offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset can be represented as shown in Table 2 below.

[0229] [Table 2]

[0230]

[0231] Here, if slice_fpel_mmvd_enabled_flag is 1, it indicates that the merge mode with motion vector difference uses integer sample precision in the current slice. If slice_fpel_mmvd_enabled_flag is 0, it indicates that the merge mode with motion vector difference can use fractional sample precision in the current slice. The slice_fpel_mmvd_enabled_flag syntax element can be signaled through the slice header or included in the slice header.

[0232] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent one of four directions, as shown in Table 3 below. The meaning of the MVD symbol may vary depending on the information of the starting MV. If the starting MV is a non-predictive MV or bidirectional MVs where both lists point to the same side of the current picture (i.e., the POCs of both references are either greater or smaller than the POC of the current picture), the sign in Table 3 below may represent the sign of the MV offset added to the starting MV. If the starting MVs are bidirectional predictive MVs and the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is smaller than the POC of the current picture), the sign in Table 3 below represents the sign of the MV offset added to the List 0 MV element of the starting MV, and the sign of the List 1 MV has the opposite value.

[0233] [Table 3]

[0234]

[0235] The two elements of the merge plus MVD offset MmvdOffset[x0][y0] can be derived as follows:

[0236] - MmvdOffset[ x0 ][ y0 ]

[0000] = ( MmvdDistance[ x0 ][ y0 ] << 2 ) * MmvdSign[ x0 ][ y0 ][0]

[0237] - MmvdOffset[ x0 ][ y0 ]

[0001] = ( MmvdDistance[ x0 ][ y0 ] << 2 ) * MmvdSign[ x0 ][ y0 ][1]

[0238] Meanwhile, the Motion Vector Prediction (MVP) mode may also be referred to as the Advanced Motion Vector Prediction (AMVP) mode. When the MVP mode is applied, a list of motion vector predictor (mvp) candidates can be generated using the motion vectors of the reconstructed spatial neighbor blocks and / or the motion vectors corresponding to the temporal neighbor blocks (or Col blocks). That is, the motion vectors of the reconstructed spatial neighbor blocks and / or the motion vectors corresponding to the temporal neighbor blocks can be used as motion vector predictor candidates. When paired prediction is applied, a list of mvp candidates for deriving L0 motion information and a list of mvp candidates for deriving L1 motion information can be generated and used separately.

[0239] The aforementioned prediction information (or information regarding prediction) may include selection information (e.g., MVP flag or MVP index) indicating the optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the list. In this case, the prediction unit of the decoding device may use the selection information to select the motion vector predictor of the current block from among the motion vector predictor candidates included in the motion vector candidate list.

[0240] The prediction unit of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can encode this to output it in the form of a bitstream. That is, the MVD can be obtained as the value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit of the decoding device can obtain the motion vector difference included in the information regarding the prediction, and derive the motion vector of the current block through the addition of the motion vector difference and the motion vector predictor. The prediction unit of the decoding device can obtain or derive a reference picture index indicating a reference picture, etc., from the information regarding the prediction. For example, the motion vector predictor candidate list can be configured as follows:

[0241] - Search for spatial candidate blocks for motion vector prediction and insert them into the prediction candidate list

[0242] - Check if the number of spatial candidate blocks is less than 2

[0243] - If the number of spatial candidate blocks is less than 2, search for temporal candidate blocks and add them to the prediction candidate list.

[0244] - If temporal candidate blocks are unavailable, use zero motion vectors.

[0245] - If the number of spatial candidate blocks is not less than 2, terminate the construction of the motion vector predictor candidate list.

[0246] Meanwhile, when MVP mode is applied, the reference picture index can be explicitly signaled. In this case, the reference picture index for L0 prediction (refidxL0) and the reference picture index for L1 prediction (refidxL1) can be signaled separately. For example, when MVP mode is applied and paired prediction (BI prediction) is applied, both information regarding refidxL0 and information regarding refidxL1 can be signaled.

[0247] When the MVP mode is applied, information regarding the Motion Vector Difference (MVD) derived from the encoding device as described above may be signaled or encoded and transmitted to the decoding device. The information regarding the MVD may include, for example, information indicating the x and y components for the absolute value and sign of the MVD. In this case, information indicating whether the absolute value of the MVD is greater than 0, whether it is greater than 1, and the remainder of the MVD may be signaled in stages. For example, information indicating whether the absolute value of the MVD is greater than 1 may be signaled only when the value of the flag information indicating whether the absolute value of the MVD is greater than 0 is 1.

[0248] For example, information regarding the MVD can be composed of the following syntax, encoded in an encoding device, and signaled to a decoding device.

[0249] [Table 4]

[0250]

[0251] For example, MVD[compIdx] can be derived based on abs_mvd_greater0_flag[compIdx] * ( abs_mvd_minus2[compIdx] + 2 ) * ( 1 2 * mvd_sign_flag[compIdx]). Here, compIdx (or cpIdx) represents the index of each component and can have a value of 0 or 1. A compIdx value of 0 can represent the x component, and a compIdx value of 1 can represent the y component. However, this is merely an example, and values ​​for each component can be represented using a coordinate system other than the x and y coordinate systems.

[0252] Meanwhile, an MVD (MVDL0) for L0 prediction and an MVD (MVDL1) for L1 prediction may be signaled separately, and information regarding the MVD may include information regarding MVDL0 and / or information regarding MVDL1. For example, if MVP mode is applied to the current block and BI prediction is applied, both information regarding the MVDLO and information regarding MVDL1 may be signaled.

[0253] Meanwhile, when BI prediction is applied, symmetric MVD may be used in consideration of coding efficiency. In this case, some of the signaling of motion information may be omitted. For example, when symmetric MVD is applied to the current block, information regarding refidxL0, information regarding refidxL1, and information regarding MVDL1 may not be signaled from the encoding device to the decoding device, but may be derived internally. For example, when MVP mode and BI prediction are applied to the current block, flag information indicating whether symmetric MVD is applied (e.g., symmetric MVD flag information or sym_mvd_flag syntax element) may be signaled, and when the value of the flag information is 1, the decoding device may determine that symmetric MVD is applied to the current block.

[0254] When the symmetric MVD mode is applied (i.e., when the value of the symmetric MVD flag information is 1), information regarding mvp_l0_flag, mvp_l1_flag, and MVDL0 may be explicitly signaled, and as described above, the signaling of information regarding refidxL0, information regarding refidxL1, and information regarding MVDL1 may be omitted and derived internally. For example, refidxL0 may be derived as an index pointing to the previous reference picture closest to the current picture in POC order within referene picture list 0 (which may be called list 0 or L0). refidxL1 may be derived as an index pointing to the subsequent reference picture closest to the current picture in POC order within reference picture list 1 (which may be called list 1 or L1). Alternatively, for example, refidxL0 and refidxL1 may both be derived as 0. Alternatively, for example, the above refidxL0 and refidxL1 can each be derived as minimum indices having the same POC difference in relation to the current picture. As a specific example, [POC of the current picture] - [POC of the first reference picture indicated by refidxL0] is called the first POC difference, and [POC of the second reference picture indicated by refidxL1] is called the second POC difference. Only when the first POC difference and the second POC difference are identical, the value of refidxL0 pointing to the first reference picture is derived as the value of refidxL0 of the current block, and the value of refidxL1 pointing to the second reference picture is derived as the value of refidxL1 of the current block.In addition, for example, if there are multiple sets where the first POC difference and the second POC difference are identical, the refidxL0 and refidxL1 of the set with the minimum difference among them can be derived as the refidxL0 and refidxL1 of the current block.

[0255] MVDL1 can be derived as -MVDL0. For example, the final MV for the current block can be derived as shown in Equation 1 below.

[0256] [Equation 1]

[0257]

[0258] Conventional video coding systems use only a single motion vector to represent the movement of an encoding block (using a translation motion model). While this method can represent optimal movement at the block level, it does not represent the actual optimal movement of each pixel; therefore, encoding efficiency can be improved if the optimal motion vector can be determined at the pixel level. To this end, we describe the affine motion prediction method, which encodes using an affine motion model.

[0259] Figure 13 is a diagram showing four movements that can be expressed in an affine motion model.

[0260] The affine motion prediction method can represent motion vectors at the pixel level of a block using two, three, or four motion vectors.

[0261] An affine motion model can represent four types of motion, as illustrated in FIG. 13. An affine motion model that represents three types of motion (translation, scale, and rotate) among the motions that an affine motion model can represent is called a similarity (or simplified) affine motion model, and the proposed methods are described below based on the similarity affine motion model. However, the disclosed embodiments are not limited to said motion model.

[0262] Figure 14 is a diagram showing an example of a control point motion vector used in affine motion prediction.

[0263] As illustrated in FIG. 14, affine motion prediction can determine the motion vector of a pixel location contained in a block using two or more control point motion vectors (CPMV). At this time, the set of motion vectors is called the affine motion vector field (MVF) and can be determined by the following equations.

[0264] For a 4-parameter affine motion model, the motion vector at the sample position (x,y) of the block can be derived by Equation 2 below.

[0265] [Equation 2]

[0266]

[0267] For a 6-parameter affine motion model, the motion vector at the sample position (x,y) of the block can be derived by Equation 3 below.

[0268] [Equation 3]

[0269]

[0270] Here is the CPMV of the CP at the top-left corner position of the encoding block, and is the CPMV of the CP at the top-right corner position, and is the CPMV of the CP at the bottom-left corner position. And W corresponds to the width of the current block, and H corresponds to the height of the current block, and is the motion vector at position {x, y}.

[0271] During the encoding / decoding process, the affine MVF can be determined at the pixel level or at the level of a predefined subblock. When determined at the pixel level, a motion vector is obtained based on each pixel value, and when determined at the subblock level, the motion vector of the corresponding block is obtained based on the pixel value of the center of the subblock (the center bottom-right, i.e., the bottom-right sample among the four central samples). For example, as shown in FIG. 14, it is possible for the affine MVF to be determined at the level of a 4x4 subblock. However, this is merely an example, and the size of the subblock applied to affine prediction can be varied in many ways.

[0272] If affine prediction is available, motion models applicable to the current block may include three models: a Translational motion model, a 4-parameter affine motion model, and a 6-parameter affine motion model. Here, the Translational motion model may represent a model in which existing block-unit motion vectors are used, the 4-parameter affine motion model may represent a model in which two CPMVs are used, and the 6-parameter affine motion model may represent a model in which three CPMVs are used.

[0273] Affine motion prediction may include affine MVP (or affine inter) mode and affine merge. In affine motion prediction, motion vectors of the current block can be derived on a sample-by-sample or sub-block basis.

[0274] In Affine Merge mode, CPMV can be determined based on the affine motion model of neighboring blocks coded according to affine motion prediction. Neighboring blocks affine-coded in the search order can be used in Affine Merge mode. If one or more neighboring blocks are coded by affine motion prediction, the current block can be coded by Affine Merge.

[0275] That is, when the affine merge mode is applied, the CPMVs of the current block can be derived using the CPMVs of neighboring blocks. In this case, the CPMVs of neighboring blocks may be used as they are for the current block, or they may be modified based on the size of the neighboring blocks and the size of the current block to be used as the CPMVs for the current block.

[0276] Meanwhile, in the case of an affine merge where an MV is derived at the subblock level, it may be called a subblock merge mode, which can be indicated based on the merge subblock flag (merge_subblock_flag (value 1)). In this case, the affine merging candidate list described below may also be called a subblock merging candidate list. In this case, the subblock merging candidate list may further include candidates derived by SbTMVP. In this case, the candidate derived by sbTMVP can be used as a candidate for index 0 of the subblock merging candidate list. In other words, the candidate derived by sbTMVP may be positioned ahead of the inherited affine candidates and constructed affine candidates described below within the subblock merging candidate list.

[0277] When Affine Merge mode is applied, an Affine Merge candidate list may be constructed to derive CPMVs for the current block. The Affine Merge candidate list may include, for example, at least one of the following candidates:

[0278] 1) Inherited affine candidates

[0279] 2) Constructed affine candidates

[0280] 3) Zero MVs Candidates

[0281] Here, inherited affine candidates are candidates derived based on the CPMVs of neighboring blocks when neighboring blocks are coded in affine mode, constructed affine candidates are candidates derived by constructing CPMVs based on the MVs of neighboring CP blocks for each CPMV unit, and zero MVs candidates may represent candidates composed of CPMVs whose value is 0. For example, zero MVs candidates may be optionally inserted into the candidate list when the current number of candidates is less than the maximum number of candidates.

[0282] Figure 15 is an example showing the inheritance of control point motion vectors.

[0283] Two inherited affine candidates can be derived from the affine motion model of the surrounding blocks. One of these can be derived from the left surrounding CUs, and the other can be derived from the upper CUs.

[0284] Referring again to Figure 12 described earlier, for the left predictor, the scan order is A0 -> A1, and for the upper predictor, the scan order is B0 -> B1 -> B2. Only the first inherited candidate on each side can be selected, and no pruning check is performed between the two inherited candidates.

[0285] When a surrounding affine CU is identified, the control point motion vectors are used to derive CPMVP candidates from the affine merge list of the current CU. Referring to Fig. 15, when the surrounding lower-left block A is coded in affine mode, motion vectors v2, v3, and v4 of the upper-left corner, upper-right corner, and lower-left corner of the CU containing block A are obtained. When block A is coded in a 4-parameter affine model, two CPMVs of the current CU are calculated based on v2 and v3. When block A is coded in a 6-parameter affine model, three CPMVs of the current CU are calculated based on v2, v3, and v4.

[0286] Figure 16 is a drawing showing an example of a surrounding block for the current block.

[0287] The constructed affine candidate refers to a candidate constructed by combining neighbor translational motion information of each control point. Motion information for the control point can be derived from designated spatial and temporal neighbors as shown in FIG. 16.

[0288] CPMV k (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, blocks are checked in the order B2 -> B3 -> A2, and the MV of the first available block is used. For CPMV2, blocks are checked in the order B1 -> B0, and for CPMV3, blocks are checked in the order A1 -> A0. TMVP can be used as CPMV4 if possible.

[0289] After the MVs of the four control points are acquired, affine merge candidates can be constructed based on motion information. The following combinations of the control point MVs can be used in order for construction:

[0290] {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4},

[0291] {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}

[0292] A combination of 3 CPMVs forms a 6-parameter affine merge candidate, and a combination of 2 CPMVs forms a 4-parameter affine merge candidate. To skip the motion scaling procedure, if the reference indices of the control points are different, the related combination of control point MVs is not used.

[0293] The subblock-based temporal motion vector prediction (SbTMVP) method can be used. Similar to Temporal Motion Vector Prediction (TMVP), SbTMVP uses the motion field of the collocated picture to improve motion vector prediction and merge modes for the CUs of the current picture.

[0294] The same collocated picture used in TMVP is also used in SbTMVP.

[0295] SbTMVP differs from TMVP in the following two main aspects:

[0296] 1. TMVP predicts movement at the CU level, but SbTMVP predicts movement at the sub-CU level.

[0297] 2. TMVP obtains the motion vector of the collocated block from the collocated picture (the collocated block is the block located at the bottom-right or center (bottom-right center) position relative to the current CU), whereas SbTMVP obtains temporal motion information from the collocated picture after applying a motion shift. Here, the motion shift is obtained from the motion vector of one of the spatially adjacent blocks of the current CU.

[0298] Figure 17 illustrates the process of SbTMVP.

[0299] SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, the spatially adjacent block A1 shown in Fig. 17(a) is examined. If A1 has a motion vector using a collocated picture as a reference picture, this motion vector (which may be referred to as the temporal motion vector (tempVM)) is selected as the motion shift to be applied. If such a motion is not identified, the motion shift is set to (0,0).

[0300] In the second step, the motion shift identified in the first step is applied (i.e., added to the coordinates of the current block) to obtain motion information (motion vectors and reference indices) at the sub-CU level from the collocated picture as illustrated in FIG. 17(b). In the example of FIG. 17(b), it is assumed that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of the corresponding block (the minimum motion grid covering the central sample) in the collocated picture is used to derive the motion information of the sub-CU. The central sample (bottom-right central sample) may correspond to the bottom-right sample among the four central samples within the sub-CU if the horizontal and vertical lengths of the sub-block are even. After the motion information of the collocated sub-CU is identified, it is converted into the motion vectors and reference indices of the current sub-CU in a manner similar to the TMVP process of HEVC.

[0301] In the second step, the motion shift identified in the first step is applied (i.e., added to the coordinates of the current block) to obtain motion information (motion vectors and reference indices) at the sub-CU level from the collocated picture as illustrated in FIG. 17 (b). In the example of FIG. 17 (b), it is assumed that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of the corresponding block (the minimum motion grid covering the central sample) in the collocated picture is used to derive the motion information of the sub-CU. The central sample (bottom-right central sample) may correspond to the bottom-right sample among the four central samples within the sub-CU when the horizontal and vertical lengths of the sub-block are even.

[0302] After the motion information of the collocated sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC. At this time, temporal motion scaling may be applied, which serves to align the reference picture of the temporal motion vector with the reference picture of the current CU.

[0303] Bilinear interpolation and sample padding are explained. The resolution of the motion vector (MV) can be 1 / 16 of a luminance sample unit. Samples corresponding to fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, search points with integer sample offsets are set around the initial fractional pixel MV, so samples corresponding to those fractional positions must be interpolated during the DMVR search process.

[0304] To reduce computational complexity, a bi-linear interpolation filter is used in the DMVR search process to generate fractional samples. Another important benefit is that when using a bi-linear filter, even when applying a 2-sample search range, it is not necessary to access more reference samples compared to a standard motion compensation process.

[0305] After the refined MV is obtained through the DMVR search process, a standard 8-tap interpolation filter is applied to generate the final prediction value. To avoid accessing more reference samples than in the standard motion compensation process, samples that are not required in the original MV-based interpolation process but are required in the refined MV-based interpolation process are provided by padding from available samples.

[0306] Previously, when use_integer_mv_flag was set to 0 in the slice header, the Motion Vector Difference (MVD) between the motion vector of the CU and the predicted motion vector was signaled in quarter-luma-sample units. In the disclosed embodiment, an Adaptive Motion Vector Resolution (AMVR) scheme is introduced at the CU level. AMVR enables the MVD of the CU to be encoded at different precisions or resolutions. Depending on the current mode of the CU (general AMVP mode or affine AMVP mode), the MVD of the CU can be adaptively selected from the following precisions:

[0307] - Standard AMVP Mode: Quarter-luma-sample, integer-luma-sample, or four-luma-sample

[0308] - Affine AMVP Mode: Quarter Luma Sample, Integer Luma Sample, or 1 / 16 Luma Sample

[0309] The CU-level MVD resolution indication is conditionally signaled only if the current CU contains at least one non-zero MVD component. If all MVD components (i.e., horizontal and vertical MVDs for reference lists L0 and L1) are zero, the quarter-luminance sample MVD resolution is automatically inferred. If at least one non-zero MVD component exists in the CU, the first flag is signaled to indicate whether a quarter-luminance sample precision MVD is used for that CU. If the first flag is zero, no further signaling is required, and a quarter-luminance sample precision MVD is used for the current CU. Otherwise, the second flag is signaled to indicate whether integer luminance sample or 4-luminance sample MVD precision is used for a standard AMVP CU. The same second flag is also used to indicate integer luminance sample or 1 / 16-luminance sample MVD precision for an affine AMVP CU.

[0310] To ensure that the reconstructed motion vector has the intended precision (quarter-luminance, integer-luminance, or 4-luminance), the motion vector predictor (MVP) of the corresponding CU is rounded to the same precision as the MVD before being added to the MVD. At this time, the motion vector predictor is rounded toward 0, which means that negative MVPs are rounded toward positive infinity and positive MVPs are rounded toward negative infinity.

[0311] The encoding device (200) determines the optimal motion vector precision for the current CU through an RD (rate-distortion) cost evaluation. To avoid performing three CU-level RD checks for every MVD precision, in one example, the RD check for MVD precision excluding quarter-luminance samples may be performed conditionally.

[0312] In the general AMVP mode, the RD cost of the quarter-luminous sample and integer-luminous sample precision is calculated first. Then, by comparing whether the RD cost of the integer-luminous sample precision is smaller than that of the quarter-luminous sample, it is determined whether it is necessary to further evaluate the RD cost of the 4-luminous sample MVD precision. If the RD cost of the quarter-luminous sample is significantly smaller than that of the integer-luminous sample, the RD evaluation for the 4-luminous sample MVD precision is omitted.

[0313] In the case of Affine AMVP mode, after evaluating the RD cost for Affine Merge / Skip mode, Merge / Skip mode, General AMVP mode with quarter luma sample precision, and Affine AMVP mode with quarter luma sample precision, if Affine Inter mode is not selected, Affine Inter mode with 1 / 16 luma sample or integer sample precision is not evaluated.

[0314] In addition, affine parameters obtained in affine inter mode with quarter-luminous sample precision are used as starting points for search in affine inter mode with 1 / 16-luminous sample and integer-luminous sample precision.

[0315] Geometric partitioning mode (GPM) may be supported for inter-prediction. GPM is a type of merge mode that can be signaled using CU-level flags. Other merge modes may include Regular Merge mode, MMVD mode, CIIP mode, and Sub-block Merge mode. Each possible CU size wxh = 2 m x 2 n A total of 64 partitions can be supported for . Here, m,n ∈ {36}, and 8x64 and 64x8 are excluded.

[0316] When this mode is used, the CU can be divided into two parts by a geometrically positioned straight line. The position of the dividing line can be mathematically derived from the angle and offset parameters of a specific partition. Each part of the geometric partition of the CU can be inter-predicted using its own motion. Only a single prediction is allowed for each partition, that is, each part has one motion vector and one reference index. The uni-prediction motion constraint can be applied to ensure that only two motion-compensated predictions are required for each CU, just like the existing bi-prediction.

[0317] Figure 18 illustrates GPM segments grouped at the same angle.

[0318] A single predicted movement of each partition can be induced using a process as illustrated in Fig. 18.

[0319] When a GPM is used for the current block, a geometric partitioning index representing the partition mode of the geometric partitioning (angle and offset) and two merge indices (one per partition) may be additionally signaled. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, and the syntax binarization of the GPM merge indices can be specified. After predicting each part of the geometric partitioning, the sample values ​​can be adjusted according to the geometric partitioning edge using adaptive weighted blending processing, such as blending along the geometric partitioning edge described below. This is a prediction signal for the entire CU, and the transformation and quantization processes can be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the GPM can be stored in motion field storage for geometric partitioning mode.

[0320] Combined inter and intra prediction (CIIP) may be applied to the current block. An additional flag (e.g., ciip_flag) may be signaled to indicate whether CIIP mode is applied to the current CU. For example, when the CU is coded in merge mode, if the CU contains at least 64 luminance samples (i.e., the product of the CU width and CU height is greater than or equal to 64) and both the CU width and CU height are less than 128 luminance samples, an additional flag may be signaled to indicate whether CIIP mode is applied to the current CU.

[0321] CIIP prediction combines the inter prediction signal and the intra prediction signal. The inter prediction signal P_inter in CIIP mode can be derived using the same inter prediction process applied in regular merge mode, and the intra prediction signal P_intra can be derived according to the regular intra prediction process in planner mode. Then, the intra and inter prediction signals can be combined using a weighted average. Figure 19 exemplarily illustrates the neighbor blocks used in CIIP weight derivation. Here, the weight values ​​can be calculated as follows, depending on the coding modes of the neighbor blocks to the left and above:

[0322] - isIntraTop is set to 1 if the upper neighbor is available and intra-coded, and isIntraTop is set to 0 otherwise

[0323] - isIntraLeft is set to 1 if the left neighbor is available and intra-coded, and isIntraLeft is set to 0 otherwise

[0324] - If (isIntraTop + isIntraLeft) is 2, wt is set to 3

[0325] - Otherwise, if (isIntraTop + isIntraLeft) is 1, wt is set to 2.

[0326] - Otherwise, wt is set to 1

[0327] CIIP prediction can be constructed as shown in Equation 4 below.

[0328] [Equation 4]

[0329]

[0330] In a GPM including inter and intra predictions, the final predicted samples can be generated by weighting the inter-predicted samples and the intra-predicted samples for each GPM-separated region. Inter-predicted samples are derived by the inter GPM, whereas intra-predicted samples can be derived by an index signaled from an intra prediction mode (IPM) candidate list and an encoding device. For example, the size of the IPM candidate list can be predefined as 3.

[0331] Figure 20 exemplarily shows available IPM candidates for GPM including inter and intra predictions.

[0332] Available IPM candidates may be parallel angle mode, perpendicular angle mode, and planar mode with respect to the GPM block boundaries, as illustrated in FIG. 20 (a), (b), and (c). Furthermore, as illustrated in FIG. 20 (d), GPMs including inter and intra predictions may be limited to reduce signaling overhead for IPMs and prevent an increase in the size of the intra prediction circuit of the hardware decoder. Additionally, direct motion vectors and IPM storage in the GPM blending region may be introduced to further improve coding performance.

[0333] In DIMD and adjacent mode-based IPM derivation, parallel modes can be registered first. Therefore, if there are no identical IPM candidates in the list, up to two IPM candidates derived from the decoder-side intra-mode derivation (DIMD) method and / or neighbor blocks can be registered.

[0334] In deriving neighbor modes, there are up to 5 locations for available neighbor blocks, but they may be limited by the GPM block boundary angles already in use in the GPM (GPM-TM) using template matching.

[0335] Table 5 shows the locations of available neighboring blocks for deriving IPM candidates according to the GPM block boundary angle. A and L may represent the upper and left sides of the predicted block.

[0336] [Table 5]

[0337]

[0338] GPM-Intra can be combined with GPM-MMVD (GPM with merge with motion vector difference). To further improve coding performance, TIMD can be used for IPM candidates of GPM-Intra. Parallel mode can be registered first, and then TIMD, DIMD, and IPM candidates of neighboring blocks can be registered.

[0339] Template Matching (TM) is a method for deriving decoder-side motion vectors (MV) and is a method for finding the optimal match between the template of the current CU (i.e., the top and / or left adjacent block of the current CU) within the current picture and the block within the reference picture (i.e., a block of the same size as the template) in order to refine the motion information of the current coding unit (CU).

[0340] FIG. 21 illustrates a method for deriving motion vectors in template matching. For example, as shown in FIG. 21, a more appropriate motion vector is searched around the initial motion of the current CU within the [-8, +8] pel search range. Additionally, the search step size may be determined according to the AMVR mode, and the TM may be applied continuously with the bidirectional matching process in merge mode.

[0341] In AMVP mode, MVP candidates are determined based on template matching errors to select the candidate with the minimum difference between the template of the current block and the template of the reference block, and then TM is performed only on the MVP candidate for MV refinement. TM refines the MVP candidate and is performed using an iterative diamond search within a [-8, +8] pel search range starting from full-pel MVD precision (or 4-pel precision for 4-pel AMVR mode). The AMVP candidate can be further refined through a cross search using full-pel MVD precision (or 4-pel precision for 4-pel AMVR mode) and then sequentially refined to half-pel and quarter-pel precision according to Table 6 specified in AMVR mode.

[0342] [Table 6]

[0343]

[0344] This search process ensures that AMVP candidates maintain the same MV precision specified by AMVR mode even after the TM process. During the search process, if the difference between the previous minimum cost and the current minimum cost in the iteration process is smaller than a threshold equal to the area of ​​the block, the search process is terminated.

[0345] In merge mode, a similar search method is applied to merge candidates designated as merge indices. As shown in Table 6 above, TM may be performed up to 1 / 8 Pell MVD precision, or steps exceeding Banpel MVD precision may be skipped. This is determined by whether the alternative interpolation filter used when AMVR is in Banpel mode is used according to the merged motion information.

[0346] In addition, when TM mode is enabled, template matching may operate as an independent process or as an additional MV refinement process between block-based and sub-block-based bidirectional matching (BM) methods. This is determined by whether the BM can be applied by satisfying the activation conditions.

[0347] In Multi-Hypothesis Prediction mode, one or more additional motion compensation prediction signals are signaled in addition to the existing bidirectional prediction signal, and the resulting total prediction signal can be obtained by sample-wise weighted superposition. Bidirectional prediction signal p bi and using the first additional inter prediction signal / assumption h3, the result prediction signal p3 can be obtained as follows.

[0348] [Equation 5]

[0349]

[0350] The weight factor α can be specified by the new syntax element add_hyp_weight_idx according to the mapping as shown in Table 7 below.

[0351] [Table 7]

[0352]

[0353] Similarly to the above, one or more additional prediction signals may be used. The resulting total prediction signal can be repeatedly accumulated with each additional signal as shown in Equation 6 below.

[0354] [Equation 6]

[0355]

[0356] The resulting overall prediction signal is the final p n (i.e., p with the largest index n) nIt can be obtained as ). For example, up to two additional prediction signals may be used (i.e., n is limited to 2).

[0357] The motion parameters of each additional prediction hypothesis can be explicitly signaled by specifying the reference index, motion vector predictor index, and motion vector difference, or implicitly signaled by specifying the merge index. A separate multi-hypothesis merge flag can distinguish between these two signaling modes.

[0358] The bilateral matching AMVP-merge mode is described below. A bidirectional predictor consists of an AMVP predictor in one direction and a merge predictor in the other direction. This mode can be activated for a coding block if the selected merge predictor and AMVP predictor satisfy the DMVR condition. Here, the DMVR condition includes that for the current picture, there exists at least one reference picture from the past and at least one reference picture from the future, and the distance from the two reference pictures to the current picture is the same. In this case, bilateral matching MV refinement is applied starting from the merge MV candidate and the AMVP MVP. Otherwise, if the template matching function is enabled, template matching MV refinement is applied to the merge predictor or AMVP predictor with the higher template matching cost.

[0359] The AMVP portion of this mode is signaled as a standard unidirectional AMVP, that is, the reference index and MVD are signaled. And when template matching is used, it has the derived MVP index, and when template matching is disabled, the MVP index is signaled.

[0360] For the AMVP direction LX (where X can be 0 or 1), the merge portion of the other direction (1-LX) is implicitly derived by minimizing the two-sided matching cost between the AMVP predictor and the merge predictor (i.e., a pair of AMVP and merge motion vectors). For all merge candidates in the list of merge candidates having motion vectors in the other direction (1-LX), the two-sided matching cost is calculated using the MV of the corresponding merge candidate and the AMVP MV. The merge candidate with the smallest cost is selected. Starting from the MV of the selected merge candidate and the AMVP MV, two-sided matching refinement is applied to the coding block.

[0361] The third pass of the multi-pass DMVR (8x8 sub-PU BDOF refinement) is enabled for blocks coded in AMVP-merge mode.

[0362] This mode is indicated by a flag, and when the mode is activated, the AMVP direction LX is indicated by an additional flag.

[0363] If the pairwise matching (BM) AMVP-merge mode is used for the current block and template matching is enabled, the MVD is not signaled. An additional AMVP-merge MVP pair is introduced. The merge candidate list is sorted in ascending order based on the BM cost. An index (0 or 1) is signaled to indicate which merge candidate from the sorted merge candidate list to use. If only one candidate exists in the merge candidate list, the pair of AMVP MVP and merge MVP is padded without pairwise matching MV refinement.

[0364] Local Illumination Compensation (LIC) is an inter-prediction technique for modeling the local light variation between the current block and its predicted block as a function of the light variation between the current block template and the reference block template. The parameters of this function can be denoted by scale α and offset β, which form a linear equation for compensating for light variation, namely α*p[x]+β, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. When wrap-around motion compensation is enabled, the MV must be clipped considering the wrap-around offset.

[0365] Since α and β can be derived based on the current block template and the reference block template, no encoding overhead is required for them, and only the LIC flag is signaled to indicate the use of LIC in AMVP mode.

[0366] Local light compensation is used in Inter-CU in the following aspects:

[0367] - Intra-neighbor samples can be used to derive LIC parameters.

[0368] - LIC is disabled for blocks with fewer than 32 luma samples.

[0369] - For both non-subblock mode and affine mode, the derivation of LIC parameters is performed based on template block samples corresponding to the current CU, and is not performed based on some template block samples corresponding to the top-left 16X16 unit.

[0370] - Samples of the reference block template are generated by performing motion compensation (MC) using block MV without rounding MV to integer pixel units.

[0371] In the case of paired prediction inter-CU, two sets of LIC parameters are derived individually for L0 and L1 prediction samples, respectively. An iterative approach is applied to derive the L0 and L1 LIC parameters. Specifically, the L0 LIC parameters are first derived by minimizing the difference between the L0 template prediction T0 and the template T. Then, the samples of the template T are updated by subtracting the corresponding samples of T0. Next, the L1 parameters are calculated by minimizing the difference between the L1 template prediction T1 and the updated template. Finally, the L0 parameters are calculated once again in the same manner.

[0372] Meanwhile, when paired prediction is applied to the current block, prediction samples can be derived based on a weighted average. Applying a weighted average to paired prediction can be called BCW (Bi-prediction with CU-level weight). When paired prediction is applied, paired prediction signals (paired prediction samples) can be derived through the weighted average of the L0 prediction signal and the L1 prediction signal as shown in Equation 7 below.

[0373] [Equation 7]

[0374]

[0375] For example, when applying a weighted average to paired predictions, 5 weights may be allowed, and the available weights may be set to w ∈ {-2, 3, 4, 5, 10}. For each paired prediction CU, the weight w can be determined according to one of the following two methods:

[0376] - In the case of a non-merge CU, the weight index is signaled after the motion vector difference.

[0377] - In the case of a merge CU, a weighted index is inferred based on surrounding blocks and a selected merge candidate index.

[0378] Weighted average pair prediction is applied only to CUs with a luma sample count of 256 or more (i.e., when the width Х height of the CU is 256 or more). For low-delay pictures, all 5 weights are used, and for non-low-delay pictures, only 3 weights (w ∈ {3, 4, 5}) are used.

[0379] In the encoding device (200), a fast search algorithm for quickly searching for weight indices without increasing complexity may be applied. The relevant algorithm can be summarized as follows:

[0380] - When combined with AMVR (Adaptive Motion Vector Resolution), unequal weights for 1-pel and 4-pel motion vector precision are conditionally checked only when the current picture is low-delay.

[0381] - When combined with an affine mode, Affine ME (motion estimation) is performed on asymmetric weights only if the affine mode is selected as the current optimal mode.

[0382] - If the two reference pictures are identical, asymmetric weights are checked only conditionally.

[0383] In addition, asymmetric weights are not explored if specific conditions are met based on the difference in POC between the current picture and the reference picture, coding QP, and temporal level.

[0384] BCW weight indices are coded using a single context-coded bin followed by bypass-coded bins. The first context-coded bin indicates whether equal weights are used, and if asymmetric weights are used, additional bins can be signaled via bypass to indicate that asymmetric weights are being used.

[0385] Weighted Prediction (WP) is a coding tool for efficiently encoding video content containing fading. WP can signal weights and offsets for each list (L0, L1) for each reference picture, and the weights and offsets of the corresponding reference picture are applied during the motion compensation process. WP and BCW are techniques designed to be suitable for different types of video content. To reduce decoder design complexity, when the CU uses WP, the BCW weight index is not signaled, and in this case, the weight w is considered to be 4 (equal weight).

[0386] In the case of a merged CU, the weighted index is inferred based on neighboring blocks and selected merge candidate indices; this method can be applied not only to general merge mode but also to inherited affine merge mode. In the configured affine merge mode, affine motion information is constructed based on motion information obtained from up to three blocks. In this case, the procedure for deriving the BCW index of the CU is as follows:

[0387] 1) Divide the range of the BCW index {0,1,2,3,4} into three groups {0}, {1,2,3}, and {4}. If the BCW index of all control points belongs to the same group, proceed to step 2. Otherwise, the BCW index of the currently configured candidate is set to 2.

[0388] 2) If two or more control points have the same BCW index, the corresponding index value is assigned to the candidate. Otherwise, the BCW index of the currently configured candidate is set to 2.

[0389] The disclosed embodiment provides a method to improve compression performance by utilizing various motion vectors when applied to an inter-frame prediction process. To this end, not only unidirectional prediction or paired prediction but also more than that may be allowed. In one embodiment, various motion vectors are allowed by permitting a motion vector predictor (MVP) that includes multiple motion vectors, and the accuracy of the prediction blocks can be improved through a weighted sum between the prediction blocks indicated by each motion vector.

[0390] According to one embodiment, in the process of configuring MVP candidates, a candidate having multiple motion vectors may be included in the candidate list. The candidate having multiple motion vectors may represent a candidate that includes an additional motion vector in addition to a unidirectional or bidirectional motion vector (basic motion vector). In the disclosed embodiment, "bidirectional" refers to a bi-direction and may represent the L0 direction and the L1 direction. Therefore, it does not necessarily mean only different directions relative to the current picture, but also includes cases where they are the same direction. That is, bidirectional prediction or bidirectional motion vector may refer to paired prediction or paired prediction motion vector.

[0391] The number of the above additional motion vectors may be predefined or signaled.

[0392] The above multiple motion vectors can be divided into regular motion vectors and additional motion vectors, and each motion vector may be included as part of the information within the regular motion information and additional motion information. Since each motion information may include BCW index, LIC flag, and interpolation filter information, etc., during the process of generating the final prediction block, the information may be modified, stored, and utilized according to the characteristics of the block.

[0393] In the process of generating the final prediction block, an additional weight index may be signaled for a weighted sum between the basic prediction block and the additional prediction block. The weight index may be modified, such as by changing and applying the signaled weight value, taking into account the size and shape of the block, whether bidirectional prediction is performed, the BCW index, and the LIC flag.

[0394] The above mode having multiple motion vectors can operate as a prediction mode and as a part of the general merge mode tool.

[0395] Alternatively, the mode having the multiple motion vectors may be signaled as a single prediction mode, as a mode distinct from the merge mode and the inter mode.

[0396] In one embodiment, a method for including a candidate having multiple motion vectors is described in the process of configuring an MVP candidate in an inter-frame prediction mode. In an inter-frame prediction mode, motion vector predictors are used in various modes, such as AMVP mode and merge mode. Since higher accuracy of the predictor contributes to improved compression performance, it is effective to configure various candidates. As described above, in this embodiment, the motion vector predictor (MVP) is used to derive the motion vector of the current block and can be used in various modes, such as merge mode as well as AMVP mode. Accordingly, the MVP may be included in the merge candidate list or in the MVP candidate list configured in AMVP mode. For example, if an MVP included in the merge candidate list is selected by an index, the current block can use the selected MVP as the motion vector, and if an MVP included in the MVP candidate list is selected by an index, the current block can use the selected MVP and the motion vector derived based on the MVD. In the embodiments described below, for convenience of explanation, any candidate list containing an MVP may be referred to as the MVP candidate list.

[0397] In addition, an MVP candidate having multiple motion vectors according to one embodiment may be applied not only to AMVP mode or merge mode, but also to modes such as affine mode, AMVP-merge mode, SMVD (Symmetric MVD), sbTMVP, MMVD mode, GPM mode, GPM-Inter / Intra mode, CIIP mode, TM mode, etc. However, the modes / tools listed above are merely examples to which an MVP candidate having multiple motion vectors may be applied according to the disclosed embodiment, and among the modes or tools not listed above, if there is a mode or tool that uses an MVP candidate, an MVP candidate having multiple motion vectors according to the disclosed embodiment may be applied.

[0398] As previously mentioned, the number of multiple motion vectors included in the MVP candidate can be determined by the number of predefined additional motion vectors. For example, referring to the syntax in Table 8 below, information indicating whether multiple motion vectors are supported or allowed in SPS (e.g., sps_multiple_predictor_enabled_flag) and information indicating the maximum number of additional motion vectors that can be added when multiple motion vectors are allowed, i.e., when the flag is '1' (e.g., sps_max_num_additional_predictor_minus1) may be signaled. This means that 'sps_max_num_additional_predictor_minus1 + 1' additional motion information can be included in the unidirectional or bidirectional motion information that an existing block may have.

[0399] [Table 8]

[0400]

[0401] It is obvious that the location of information indicating whether multiple motion vectors are allowed and information indicating the maximum number of motion vectors that can be added may be included in a higher parameter set other than SPS, such as VPS, PPS, APS, PH, SH, etc. For example, if a flag indicating whether multiple motion vectors are allowed is located in PPS, the allowance of multiple motion vectors can be determined on a picture-by-picture basis rather than by sequence.

[0402] In addition, information indicating whether the multiple motion vectors are allowed and information indicating the maximum number of motion vectors that can be added may be located at different levels, so that the maximum number of multiple motion vectors can be specified differently depending on the application unit. For example, as shown in the example of Table 8 above, sps_multiple_predictor_enabled_flag indicating whether multiple motion vectors are allowed exists in SPS, and sps_max_num_additional_predictor indicating the maximum number of multiple motion vectors is signaled at different levels such as picture, slice, CTU, CU, etc., so that the maximum number of multiple motion vectors can be variably specified differently for each unit where sps_max_num_additional_predictor is located.

[0403] In addition, the maximum number of multiple motion vectors can be set according to a predefined value without being separately signaled. For example, as illustrated in the syntax of Table 9 below, the maximum number of motion vectors that can be added when the value of the information indicating whether multiple motion vectors are allowed (e.g., sps_multiple_predictor_enabled_flag) is 1 can be predefined as a specific integer value without separate signaling.

[0404] [Table 9]

[0405]

[0406] For example, a specific value of 1 or 2 can be used as the maximum number of motion vectors to be added. This means that one or two additional motion information can be included in the unidirectional or bidirectional motion information that the existing block can have.

[0407] FIG. 22 is a diagram showing a case where multiple reference blocks are used in one embodiment.

[0408] In one embodiment, a unidirectional or bidirectional motion vector may be defined as a regular motion vector, and a reference block obtained using the regular motion vector may be defined as a regular reference block. Additionally, a motion vector added thereto may be defined as an additional motion vector, and a reference block obtained using the motion vector may be defined as an additional reference block. A multiple motion vector may be defined to include a regular motion vector and an additional motion vector, and a multiple reference block may be defined to include a regular reference block and an additional reference block.

[0409] In addition, depending on the case, multiple reference blocks or multiple prediction blocks may mean additional reference blocks or additional prediction blocks, and multiple motion vectors may mean additional motion vectors.

[0410] Additionally, a prediction block generated based on a basic reference block is referred to as the basic prediction block, a prediction block generated based on an additional reference block is referred to as the additional prediction block, and the basic prediction block and the additional prediction block may be collectively referred to as a multiple prediction block. Furthermore, a prediction block generated based on the basic prediction block and the additional prediction block (e.g., a weighted sum or a weighted average) may be referred to as the final prediction block.

[0411] Alternatively, it is also possible to refer to all blocks as primary reference blocks, additional reference blocks, or multiple reference blocks until the final predicted block for the current block is generated.

[0412] Alternatively, a combination of basic reference blocks, multiple reference blocks, or additional reference blocks and basic prediction blocks, multiple prediction blocks, or additional prediction blocks may be used.

[0413] Referring to FIG. 22, if multiple motion vectors are allowed for the current block C, bidirectional basic reference blocks P0 and P1 are obtained, and additional reference blocks P2 and P3 can be obtained based on additional motion vectors. For example, a final prediction block can be generated by applying a weighted sum to the basic reference blocks and additional reference blocks.

[0414] For the sake of convenience of explanation, it was assumed that two reference blocks are obtained for each prediction direction, but it goes without saying that the number of reference blocks obtained per direction may change.

[0415] FIGS. 23 and FIGS. 24 are flowcharts illustrating an example of a method for including motion information candidates containing multiple motion information in a motion information candidate list, in a decoding method and an encoding method according to one embodiment. The decoding method of FIG. 23 and the encoding method of FIG. 24 can be performed by the aforementioned decoding device (300) and encoding device (200). The decoding device (300) and the encoding device (200) may each include at least one memory and a processor electrically connected to the at least one memory, and the methods described below may be performed by the processor.

[0416] In the embodiments described below, the focus is on the content not previously described in order to avoid redundant explanations, and the descriptions below alone do not support the embodiments of the decoding method and the encoding method. The descriptions regarding the operation of the aforementioned decoding device (300) and encoding device (200), the descriptions regarding the decoding method and the encoding method (e.g., FIGS. 1 to 10), and the descriptions regarding various prediction modes, prediction types, and prediction tools ( FIGS. 11 to 22) may be applied in the embodiments described below in the same way as long as they do not conflict with one another.

[0417] Referring to FIG. 23, a decoding method according to one embodiment may include the step of obtaining image information from a bitstream (S900), the step of configuring a motion information candidate list including motion information candidates for a current block based on the image information (S910), and the step of generating a prediction block based on at least one motion information candidate (S920).

[0418] Referring to FIG. 24, an encoding method according to one embodiment includes the steps of: configuring a motion information candidate list including motion information candidates for a current block (S1000); generating a prediction block for a current block based on at least one motion information candidate in the motion information candidate list (S1010); and encoding image information for a current block (S1020).

[0419] Here, motion information candidates may include information obtained from previously restored surrounding blocks of the current block. For example, motion information candidates may include motion vectors of previously restored surrounding blocks. If merge mode is applied to the current block, motion information candidates may be included in merge candidates, and if AMVP mode is applied to the current block, motion information candidates may be included in MVP candidates. Alternatively, it is also possible to refer to motion information candidates as MVP candidates for merge mode.

[0420] Additionally, the motion information candidate may include multiple motion candidates having basic motion information and additional motion information. That is, the step of configuring the motion information candidate list (S910, S1000) may include adding a motion information candidate having basic motion information and additional motion information to the motion information candidate list.

[0421] Additionally, the image information may include weight information applied to the weighted sum of a basic prediction block derived based on basic motion information and an additional prediction block derived based on additional motion information. The weight information may include a weight index pointing to one of a plurality of weight candidates.

[0422] The image information obtained from the bitstream, or the image information encoded by the encoding device (200), may include information regarding prediction. Based on the prediction mode indicated by the information regarding prediction being a prediction mode using MVP, for example, the aforementioned AMVP mode or merge mode, the prediction mode applied to the current block may be determined as a prediction mode using motion information candidate or MVP candidate.

[0423] The image information obtained from the bitstream may include information indicating whether multiple motion vectors are supported or allowed (e.g., sps_multiple_predictor_enabled_flag). Based on the information indicating that multiple motion vectors are allowed (e.g., the value of sps_multiple_predictor_enabled_flag is 1), an MVP candidate list including MVP candidates with multiple motion vectors can be constructed.

[0424] A prediction block is generated based on the final MVP candidate within the MVP candidate list. The final MVP candidate refers to the MVP candidate used for the prediction of the current block among the MVP candidates within the MVP candidate list. The encoding device (200) can encode video information containing information representing the final MVP candidate (e.g., index information) and transmit it to the decoding device (300) in the form of a bitstream. The decoding device (300) can obtain the corresponding information from the transmitted bitstream and derive the final MVP candidate. If the final MVP candidate has multiple motion vectors, a prediction block can be generated using an additional reference block along with a basic reference block. For example, a prediction block can be generated by applying weights to the basic reference block and the additional reference block. Alternatively, a final prediction block can be generated by applying weights to the basic prediction block and the additional prediction block.

[0425] Meanwhile, it goes without saying that the description of the decoding method and encoding method described above may also be applied to an embodiment. For example, a decoding method according to an embodiment may generate a restoration block after generating a prediction block, based on the generated prediction block and residual information included in the image information (which may be omitted depending on the prediction mode). An encoding method according to an embodiment may generate residual information (which may be omitted depending on the prediction mode) based on the generated prediction block, and encode image information including information regarding the prediction and residual information, and transmit it in the form of a bitstream.

[0426] FIG. 25 is a diagram showing an example of a method for configuring an MVP candidate list in the disclosed embodiment.

[0427] The example in FIG. 25 illustrates the process of constructing an MVP candidate list in basic merge mode. Referring to FIG. 25, spatial neighboring candidates (S911), temporal candidates (S912), non-adjacent spatial neighboring candidates (S913), HMVP candidates (S914), candidates with multiple motion vectors (S915), and pairwise candidates (S916) may be included in the MVP candidate list.

[0428] However, the types, order, and number of MVP candidates shown in FIG. 25 are merely examples applicable to one embodiment, and it is possible to change the types, order, or number differently, or to replace existing MVP candidates with MVP candidates having multiple motion vectors. For example, an MVP candidate including multiple motion vectors may be included in the MVP candidate list by replacing a pairwise average candidate.

[0429] Meanwhile, after the MVP candidate list is configured, a template matching cost-based reordering can be performed on each candidate within the MVP candidate list. In the process of calculating the template matching cost for the reordering of candidates, the cost for unidirectional reference blocks can be calculated based on the difference between the adjacent sample of the current block and the adjacent sample of the reference block, and the cost for bidirectional reference blocks can be calculated based on the difference between the adjacent sample of the current block and the adjacent sample of each reference block.

[0430] For candidates containing multiple motion vectors, the cost can be calculated using adjacent samples of the current block and adjacent samples of the reference blocks represented by each motion vector, or the computational complexity can be reduced by using only a subset of reference blocks. In this case, if the final prediction block is determined by a weighted sum for each reference block, the cost can be calculated by applying the weight for each reference block to its adjacent samples; if only a subset of reference blocks is used, it is possible to calculate the cost by applying a modified form of weight. For example, if only a subset of reference blocks is used for cost calculation, a method can be applied in which the weight for the relevant reference blocks is increased and the weight for the reference blocks not used in the calculation is decreased. However, this is merely one example, and it goes without saying that various other methods of applying modified weights can also be applied.

[0431] In addition, in the process of constructing the MVP candidate list, multiple candidates containing multiple motion vectors may be constructed, and a method of reordering based on template matching costs and selecting only some candidates may be applied to these candidates. The same cost calculation method as in the example above may be applied to the process of calculating the template matching cost. Furthermore, since a candidate containing multiple motion vectors implies containing at least two reference blocks, reordering and selecting only some candidates can be achieved by calculating the bilateral matching cost using each prediction block. The cost can be calculated based on the difference value between each reference block, and it is possible to calculate it for multiple reference blocks or to select some reference blocks and calculate it only for the selected reference blocks.

[0432] Meanwhile, in the process of constructing the MVP candidate list, a redundancy check between MVP candidates with multiple motion vectors and other candidates can be performed as follows:

[0433] Duplication can be determined based on the number of reference blocks for each candidate.

[0434] - If each candidate to be compared has the same number of reference blocks, duplicates can be determined using the reference index and motion vector representing the motion information of each candidate.

[0435] - If there are multiple candidates with multiple motion vectors within the MVP candidate list, and a weighted sum between reference blocks obtained from each motion vector can be applied, and the weights between the reference blocks of each candidate are different, they can be considered as non-duplicate candidates.

[0436] When constructing multiple candidates with multiple motion vectors, the motion vectors included in each candidate can be restricted to differ at least once from the motion vectors included in the preceding candidate with multiple motion vectors within the list. This makes it possible to determine duplication based solely on motion information between candidates, without the need for redundancy checks that consider weights.

[0437] When applying the aforementioned method to merge mode and AMVP mode, signaling can be performed as follows.

[0438] The MVP candidate having the above multiple motion vectors can be applied to existing merge modes and each tool that can be combined with merge modes (Sub-block based MERGE mode, CIIP mode, GPM mode, etc.), and it is possible to include the candidate having multiple motion vectors in the process of configuring the list of MVP candidates for each mode.

[0439] Additionally, it is possible to implement it as another merge mode within the general merge mode. For example, as a merge mode with multiple motion vectors, a flag indicating the mode (e.g., multi_pred_merge_flag) can be signaled. When the value of the flag is '1', the MVP candidate list may consist of candidates with multiple motion vectors, and a merge index indicating a specific candidate within the MVP candidate list may be signaled. When the process of reordering using the aforementioned template-based cost or bidirectional-based cost and selecting only some candidates is performed, the merge index may be omitted or indicate a candidate within the list containing only some candidates.

[0440] Candidates with multiple motion vectors can also be applied in AMVP mode. In AMVP mode, since MVP candidate lists can be configured for each direction, it is possible to include candidates with multiple motion vectors within each MVP candidate list for the L0 and L1 directions.

[0441] Figure 26 is a diagram showing an example of an MVP candidate list configured for each direction in AMVP mode.

[0442] When additional motion vectors are allowed in AMVP mode, the number of motion vectors that can be had for each direction is limited, so that 2 to 4 motion vectors can be included in the end. In the example of FIG. 26, Case 1 represents a case where a total of 4 MVPs are obtained through a combination of candidate CAND[2] with multiple motion vectors in the L0 direction MVP candidate list and candidate CAND[2] with multiple motion vectors in the L1 direction MVP candidate list. Case 2 represents a case where a total of 3 MVPs are obtained through a combination of candidate CAND[0] with a single motion vector in the L0 direction MVP candidate list and candidate CAND[2] with multiple motion vectors in the L1 direction MVP candidate list.

[0443] The application method in the above AMVP mode may be modified as follows, taking into account encoding / decoding efficiency. Specifically, additional motion vectors for the AMVP mode may be applied only to blocks where bidirectional prediction is applied. Additionally, limited multiple motion vectors may be allowed by constructing candidates with multiple motion vectors, limited to the MVP candidate construction process for a specific direction (L0 or L1). That is, by allowing n (n is a natural number greater than or equal to 1) additional motion vectors only in a specific prediction direction, the total number of additional motion vectors may be limited to n.

[0444] Meanwhile, it is possible to construct a candidate with multiple motion vectors without signaling whether multiple motion vectors are allowed, or to signal a flag indicating the mode to allow multiple motion vectors. In the latter case, when the mode flag is signaled and the flag has a value of '1', an MVP candidate with multiple motion vectors can be included in the MVP candidate list. The flag may indicate that a candidate with multiple motion vectors can be constructed consisting of the number of allowed motion vectors among the MVP candidates, or that an MVP candidate list consisting only of candidates with multiple motion vectors can be constructed. Alternatively, it may indicate the number of multiple motion vectors that can be included within the MVP candidate list. Additionally, an mvp_index or mvp_flag pointing to a specific candidate within the MVP candidate list may be signaled; if a process of reordering using the aforementioned template-based cost or bilateral-based cost and selecting only some candidates is performed, the mvp_index or mvp_flag may be omitted or point to a candidate within the list containing only some candidates.

[0445] Meanwhile, in AMVP mode, if the final MVP candidate—that is, the MVP candidate used for predicting the current block—is a candidate with multiple motion vectors, MVD information corresponding to each motion vector can be signaled. Then, similar to the existing AMVP mode, the final MV can be calculated using the derived MVP information and the signaled MVD information. At this time, the signaling method for MVD information can be applied in various ways, such as indicating the magnitude of the vector or indexing the MVD similarly to MMVD to derive it from information based on distance and direction. Additionally, for an MVP with multiple motion vectors, the amount of signaling information can be reduced by limiting the MVD information to the basic (regular) motion vector rather than the additional motion vectors and signaling only a portion of the information. That is, the additional motion vector predictors included in the MVP candidate can be utilized as motion vectors directly without MVD. Alternatively, when the final MVP candidate has multiple motion vectors, variations are also possible, such as transmitting MVD information for some motion vectors while not transmitting MVD information for others, or not transmitting MVD information at all.

[0446] The method described in the disclosed embodiment may be similarly applied to AMVP-Merge. Additional motion vectors may be included in each direction of the AMVP-Merge mode, and it is possible to allow additional motion vectors by a modified method, such as applying them exclusively to the AMVP mode or merge mode directions. When additional motion vectors are included in the AMVP mode direction, the candidate pointed to by mvp_index or mvp_flag may include multiple motion vectors. In this case, modifications such as additionally signaling MVD information or not signaling the additional motion vectors are possible, and when additional motion vectors are included in the merge mode direction, the candidate pointed to by the merge index may include multiple motion vectors.

[0447] As described above, when multiple reference blocks exist, each reference block is divided into a basic reference block and an additional reference block, and the average sum between each reference block can be applied to generate the final prediction block. As the variety of images increases, various types of reference blocks may be required, and weighted sums can serve the role of generating these various types of reference blocks. Below, a method for signaling / parsing the index of weight information required for the weighted sum of multiple reference blocks is described. In the disclosed embodiment, signaling of information may include encoding of said information, and parsing of information may include obtaining said information from a bitstream. Furthermore, even without separate description, a description of signaling certain information may include parsing said information, and a description of parsing certain information may include signaling said information.

[0448] A mode including multiple reference blocks exists as a prediction mode, and when the prediction mode is used, weight index information can be signaled / parsed.

[0449] FIG. 27 is a diagram illustrating an example of a signaling / parsing method when a multi-reference block mode is included as one of the general merge modes in a method according to one embodiment.

[0450] Referring to FIG. 27, for example, whether a regular merge mode is applied can be defined by a merge flag (regular_merge_flag), and whether a multi-reference block mode is applied can be defined by a multi-mvp flag (multi_mvp_flag). When the flag is TRUE, an additional merge index (additional_merge_idx) to indicate an additional reference block is signaled / parsed, and a weight index (weight_idx) for a weighted sum between the basic reference block and the additional reference block is signaled / parsed. However, this is just one example, and the signaling / parsing order or position of the multi-mvp flag may be changed.

[0451] Additionally, it is possible to omit the signaling / parsing of the additional merge index. When the signaling / parsing of the additional merge index is omitted, it can be derived as a modified form of the merge index (merge_idx). For example, it can be calculated as additional_merge_idx = merge_idx + offset. In this case, offset can be specified as a constant such as 1 or 2. As another example, the additional merge index (additional_merge_idx) may not be used, and the merge index (merge_idx) may be used as a candidate index that includes both the primary reference block and the additional reference block.

[0452] The weight index (weight_idx) signaled / parsed when the multi_mvp_flag is TRUE may be in the range of 0 to N-1, where N is a positive integer. For example, N can be 4. The weight represented by each index can be used as a weight for a weighted sum between the base reference block and the additional reference block. That is, it can be used for a weighted sum between the base reference block (the result block to which the weighted sum or mean sum is applied in the case of bidirectional prediction) and the additional reference block (the result block to which the weighted sum or mean sum is applied in the case of bidirectional prediction).

[0453] The weights can be set as follows depending on the characteristics of the current block:

[0454] - Weight information may be assigned based on the size and shape of the current block. For example, if the size of the block (e.g., the product of the width and height of the block) is greater than a predefined threshold, the weight of the base reference block may be set higher. In this case, the threshold may have a value of 128 or 256. However, the above threshold is merely an example and may have other thresholds.

[0455] Weight information can be derived based on whether unidirectional or bidirectional prediction is applied to the primary reference block and the additional reference block. For example, if unidirectional and bidirectional prediction blocks are mixed, the weight of the bidirectional prediction block can be set higher.

[0456] Weight information can be derived based on the BCW index, LIC flag, and parameters included in the basic motion information and additional motion information. For example, the weight of a prediction block containing information that includes a BCW index other than the default value can be set higher. Similarly, the weight of a prediction block containing motion information where the LIC flag is TRUE or the LIC parameter is not the default value can be set higher.

[0457] Weight information can be derived based on the distance (POC Difference) between the reference picture and the current picture included in the basic motion information and additional motion information. For example, the weight of a prediction block containing motion information with a small POC difference can be set higher. In this case, if bidirectional prediction is applied to the basic prediction block or the additional prediction block, the average value of the POC difference between the reference picture and the current picture in each direction can be used. Alternatively, the POC difference with the reference picture located closer between the two reference pictures may be used.

[0458] - Weights may be derived to different values ​​by considering the number of taps in the interpolation filter or the type of filter included in the basic motion information and additional motion information. For example, the weight of a prediction block containing information with a longer number of interpolation filter taps can be set higher. As another example, when the interpolation filters applied to the basic prediction block and the additional prediction block differ in characteristics—such as a smoothing filter and a sharpening filter—the weight of the prediction block containing smoothing filter information can be set higher.

[0459] Below, information regarding the weights that can be applied to the weighted sum between each reference block when multiple reference blocks exist is described.

[0460] For example, when a merge index or MVP index exists for a multi-reference block, the index may point to a combination of two candidates, and each candidate may contain unidirectional or bidirectional movement information. When bidirectional prediction is applied to each candidate, a prediction block may be generated by weighted summing the prediction blocks of each direction, and the final prediction block may be generated by weighted summing the basic prediction block and the additional prediction block.

[0461] FIGS. 28 and 29 illustrate an example of a signaling / parsing method for weight information applied to a multiple reference block in a method according to one embodiment. For example, the weight information may include a weight index pointing to one of a plurality of weight candidates. The method of FIGS. 28 and 29 may be applied to a decoding method and an encoding method according to one embodiment and may be performed by a decoding device (300) and an encoding device (200).

[0462] As described above, whether the multi-reference block mode is applied can be defined by the multi-mvp flag (multi_mvp_flag). Referring to FIG. 28, when the flag is TRUE, the MVP index (e.g., merge_idx) and weight index (e.g., weight_idx) can be signaled or parsed. When the multi-mvp flag is FALSE, a prediction operation according to the prediction mode applied to the current block can be performed.

[0463] Referring to FIG. 29, when multiple mvp flags are TRUE, the MVP index (e.g., merge_idx) can be signaled or parsed, and when a defined condition is satisfied, the weight index can be encoded or obtained. In the embodiments described below, the encoding or obtaining of the weight index is described as signaling or parsing the weight index.

[0464] However, FIGS. 28 and 29 are examples applicable to one embodiment, and the signaling / parsing order or position of multiple MVP flags or MVP indices may be changed.

[0465] The signaling / parsing conditions of the weight index may include conditions related to the block size (e.g., width, height, area, aspect ratio, etc.), picture size, quantization parameter, temporal ID, etc. For example, if the current block size is smaller than a threshold, the weight index may not be signaled or parsed, and a predefined weight may be assigned to perform a weighted sum of the base prediction block and the additional prediction block. As another example, if the quantization parameter is larger than a threshold, the weight index may not be signaled or parsed, and a predefined weight may be assigned to perform a weighted sum of the base prediction block and the additional prediction block. In this case, 37 may be used as the threshold for the quantization parameter. However, this is merely an example, and it goes without saying that other thresholds may also be used.

[0466] Candidates for weight indices can be defined as N (N > 1). The weights pointed to by indices 0 to N-1 can each have different values, for example, N can be 2. The weight applied to the basic prediction block can be called the first weight, and the weight applied to the additional prediction block can be called the second weight. The aforementioned multiple weight candidates may include the first weight or the second weight, and the sum of the first weight and the second weight may have a fixed value. Additionally, the first weight may have a larger value than the second weight.

[0467] For example, the first weight and the second weight can be expressed as (MW) and W, respectively. M is an integer greater than W, and for example, it can be 8. That is, in the example described below, the case where the sum of the weights applied to the basic prediction block and the weights applied to the additional prediction block is 8 is used as an example. However, the size of the sum of the weights can be changed to 16, 32, etc. When the sum is 8, W can have a value from 1 to 7, and when the sum is 16, W can have a value from 1 to 15.

[0468] For example, when the sum of the weights is 8, the value W[2] = {4, 3} can be assigned. Using the assigned weights, the final prediction block can be generated as follows.

[0469] [Equation 8]

[0470] Final Prediction Block = ((8-W[i]) * Regular Prediction Block + W[i] * Additional Prediction Block + 4) >> 3, i=0..1

[0471] The above weight candidate {4, 3} is an example, and the weight applied to the basic prediction block can be set higher by setting the weight candidate to {3, 2}. Alternatively, it is also possible to set the weight applied to the basic prediction block and the additional prediction block to be the same. In addition, the number of weight indices can be changed.

[0472] In addition, it is possible to have a single weight without signaling of the index, in which case a larger weight can be applied to the basic prediction block. For example, a weight of 3 can be applied for the additional prediction block and a weight of 5 for the basic prediction block.

[0473] FIGS. 30 to 32 are flowcharts illustrating examples of a method using an MVP index and a weight index in a decoding method or encoding method according to one embodiment. The decoding method or encoding method of FIGS. 30 to 32 may be performed by the aforementioned decoding device (300) or encoding device (200). Furthermore, it goes without saying that the description of the decoding method or encoding method illustrated in FIGS. 23 and 24 may also be applied to the example.

[0474] Referring to FIG. 30, a multiple MVP candidate list is constructed. The multiple MVP candidate list may include MVP candidates that include multiple motion vectors.

[0475] If reordering candidates using template cost is possible, MVP candidates can be reordered in ascending order of template cost. In this case, the template cost can be calculated based on the result of a weighted sum of the template areas of the basic prediction block and the additional prediction block, and weight information can be used when performing the weighted sum of each template area. Additionally, the candidate reordering process may be omitted.

[0476] Generate a primary prediction block and an additional prediction block using the candidate pointed to by the MVP index within the reordered MVP candidate list.

[0477] The final prediction block is generated by applying the weights indicated by the weight index to the basic prediction block and the additional prediction block and performing a weighted sum.

[0478] In the case where the aforementioned method is a decoding method, the MVP index and weight index may be included in the image information obtained from the bitstream.

[0479] In the case where the above-described method is an encoding method, the method may further include encoding image information including an MVP index and a weight index.

[0480] The above-described method can be simplified as shown in FIG. 31.

[0481] Referring to Fig. 31, a list of multiple MVP candidates is constructed, and if reordering of candidates using template cost is possible, the MVP candidates are reordered in ascending order of template cost. At this time, the template cost can be calculated based on the weighted sum result of the template area of ​​the basic prediction block and the template area of ​​the additional prediction block, and when weighting each template area, predefined weight information rather than signaled weight information can be used. In addition, the reordering process of the candidates may be omitted.

[0482] Generate a primary prediction block and an additional prediction block using the candidate pointed to by the MVP index within the reordered MVP candidate list.

[0483] The final prediction block is generated by applying the weights pointed to by the weight index to the basic prediction block and the additional prediction block.

[0484] In the candidate reordering process using the above template cost, the predefined weight information may be equal-weights for each prediction block. Alternatively, the predefined weight information may be the value indicated by weight index 0. This simplified decoding process can contribute to reducing the complexity of the encoding process. Specifically, when there are N weight candidates, the reordering result of the MVP candidate may differ depending on each weight; therefore, the encoding process must perform MVP candidate reordering for N weight candidates, which can increase encoding time. As mentioned above, using fixed weights requires performing MVP candidate reordering only once, making it efficient in terms of the performance-complexity tradeoff.

[0485] The above-described method can be further simplified as shown in FIG. 32.

[0486] Referring to FIG. 32, in the step of weighted summing the basic prediction block and the additional prediction block, a separate weight index may not be used, and the same weight may be applied to each prediction block, or a predefined weight may be applied. In addition, regarding the weights applied during the candidate reordering process using the template cost and / or the weights applied to the weighted summing of the basic prediction block and the additional prediction block, it is also possible to use predefined weights without a weight index only when certain conditions are satisfied, as explained earlier with reference to FIG. 28.

[0487] Meanwhile, the image information encoded in the encoding method and obtained from the bitstream in the decoding method may include a motion information index pointing to at least one motion information candidate within a motion information candidate list. For example, the motion information index may include an MVP index or a merge index, and the MVP index and the merge index may be used in combination.

[0488] When an MVP index (e.g., merge_idx) for a multi-reference block exists, the index may point to a combination of two candidates, each candidate may contain unidirectional or bidirectional motion information. When the motion information of each candidate, i.e., the motion vector, points to a fractional-pel location, a prediction block can be generated using a predefined interpolation filter. When using the interpolation filter, predefined filter coefficients such as 8 taps or 12 taps can be used, and when the AMVR resolution is 1 / 2 samples and the location actually pointed to by the motion vector is a half-pixel (1 / 2-pel) location, a specific interpolation filter, i.e., a Gaussian filter, in which all filter coefficients are greater than or equal to 0 can be applied.

[0489] Table 10 below shows examples of interpolation filter coefficients and Gaussian filter coefficients (when hpelIfIdx == 1) based on an 8-tab DCTIF (a standard interpolation filter).

[0490] [Table 10]

[0491]

[0492] General interpolation filters contribute to generating prediction blocks similar to the current block through more accurate pixel correction at fractional pixel positions, and Gaussian filters help generate various prediction blocks by creating smoothed prediction blocks with minimal changes in values ​​between pixels within the block. However, since multiple reference blocks are weighted sums using up to four reference blocks, image quality degradation may occur due to oversmoothing between pixels within the prediction blocks.

[0493] Accordingly, in one embodiment, when generating a prediction block of a multi-reference block, it may be configured to apply a general interpolation filter rather than a specific interpolation filter. Specifically, a specific interpolation filter may not be applied based on the fact that the motion information candidate pointed to by the motion information index includes basic motion information and additional motion information. For example, when the multi-reference block mode is applied, information indicating whether a specific interpolation filter is applied (e.g., hpelIfIdx) may be induced to 0 or signaled / parsed.

[0494] In the case of merge mode, during the process of deriving each MVP candidate of a multi-reference block, even if an adjacent block uses a specific interpolation filter, the current block to which the multi-reference block mode is applied can be configured to apply a general interpolation filter instead of the specific interpolation filter. For example, when the multi-reference block mode is applied, even if information indicating whether a specific interpolation filter is applied to an adjacent block (e.g., hpelIfIdx) is set to 1, it can be changed to 0 and applied.

[0495] Furthermore, the aforementioned example is not limited to multi-reference blocks but can also be applied to MHPs or CIIPs containing multiple reference blocks. That is, when MHP mode or CIIP mode is applied, the use of a specific interpolation filter may be restricted. In this case, the use of the specific interpolation filter may be restricted in the same way even if it is not limited to being applied at a half-pixel location.

[0496] In addition, when multiple reference blocks exist, the precision of motion information may be applied differently to the basic reference block and the additional reference block. Here, the precision of motion information may be referred to as AMVR resolution, MVD resolution, MV resolution, or MVD precision.

[0497] As described above, since the multi-reference block is weighted summed using up to four reference blocks, image quality degradation may occur due to oversmoothing of pixels within the prediction block. In this case, to prevent image quality degradation caused by oversmoothing, in addition to limiting the use of a specific interpolation filter, motion precision for the multi-reference block can be applied differently to the primary reference block and the additional reference block. For example, the motion precision can be maintained as is for the primary reference block, while the motion precision for the additional reference block can be changed to an integer pixel. For example, the first motion precision pointed to by the AMVR index signaled or parsed for the current block can be applied as is to the primary reference block, and the second motion precision of an integer pixel can be applied to the additional reference block. The motion precision of an integer pixel can have a predefined value.

[0498] The examples or embodiments described so far may be combined with one another, and any modifications required by the combination of embodiments may also be included within the scope of the disclosed invention or the disclosed embodiments.

[0499] FIG. 33 is a diagram illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.

[0500] Referring to FIG. 33, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0501] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.

[0502] The bitstream may be generated by an encoding method or a bitstream generation method to which the embodiment(s) of this specification are applied, stored in a computer-readable non-transient storage medium, or transmitted by a transmission method or a transmission device. The transmission device may include at least one processor for generating a bitstream and a transmitter for transmitting the generated bitstream. The streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0503] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.

[0504] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.

[0505] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.

[0506] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.

[0507] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be implemented as a device, and the technical features of the device claims in this specification may be combined to be implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a device, and the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a method.

[0508] An embodiment according to the present disclosure can be used to encode / decode images.

Claims

1. A step of acquiring image information from a bitstream; Based on the above image information, a step of constructing a motion information candidate list including motion information candidates for the current block; and The method includes the step of generating a prediction block for the current block based on at least one motion information candidate within the motion information candidate list. The step of constructing the above list of motion information candidates is, It includes adding motion information candidates having basic motion information and additional motion information to the said motion information candidate list, and The above video information is, A method comprising weight information applied to a weighted sum of a basic prediction block derived based on the basic motion information and an additional prediction block derived based on the additional motion information.

2. In Paragraph 1, The above weighting information is, A method comprising a weight index pointing to one of a plurality of weight candidates.

3. In Paragraph 1, The above weighting information is, A method obtained based on satisfying a set condition.

4. In Paragraph 3, The above-determined conditions are, A method comprising at least one of a condition regarding the size of the current block or a condition regarding a quantization parameter for the current block.

5. In Paragraph 2, A method in which the first weight applied to the above basic prediction block has a value greater than the second weight applied to the above additional prediction block.

6. In Paragraph 5, The above plurality of weight candidates include the first weight or the second weight, and A method in which the sum of the first weight and the second weight has a predetermined value.

7. In Paragraph 1, The above video information is, A method comprising a motion information index pointing to at least one motion information candidate within the above motion information candidate list.

8. In Paragraph 7, The step of generating a prediction block for the above current block is, Based on the motion information indicating fractional pixel positions, it includes applying an interpolation filter, A method of not applying a specific interpolation filter among the interpolation filters based on the fact that the motion information candidate pointed to by the motion information index includes the basic motion information and additional motion information.

9. In Paragraph 1, A first motion precision is applied to the above basic motion information, and a second motion precision is applied to the above additional motion information, and The above second motion precision is a method having the resolution of integer samples.

10. A step of constructing a motion information candidate list including motion information candidates for the current block; A step of generating a prediction block for the current block based on at least one motion information candidate in the motion information candidate list; and The method includes the step of encoding image information for the current block; and The step of constructing the above list of motion information candidates is, It includes adding motion information candidates having basic motion information and additional motion information to the said motion information candidate list, and The above video information is, A method comprising weight information applied to a weighted sum of a basic prediction block derived based on the basic motion information and an additional prediction block derived based on the additional motion information.

11. In Paragraph 10, The above weighting information is, A method comprising a weight index pointing to one of a plurality of weight candidates.

12. In Paragraph 10, The above weighting information is, A method that is encoded based on satisfying a defined condition.

13. In Paragraph 12, The above-determined conditions are, A method comprising at least one of a condition regarding the size of the current block or a condition regarding a quantization parameter for the current block.

14. In Paragraph 11, A method in which the first weight applied to the above basic prediction block has a value greater than the second weight applied to the above additional prediction block.

15. In Paragraph 14, The above plurality of weight candidates include the first weight or the second weight, and A method in which the sum of the first weight and the second weight has a predetermined value.

16. In Paragraph 10, The above video information is, A method comprising a motion information index pointing to at least one motion information candidate within the above motion information candidate list.

17. In Paragraph 16, The step of generating a prediction block for the above current block is, Based on the motion information indicating fractional pixel positions, it includes applying an interpolation filter, A method of not applying a specific interpolation filter among the interpolation filters based on the fact that the motion information candidate pointed to by the motion information index includes the basic motion information and additional motion information.

18. In Paragraph 10, A first motion precision is applied to the above basic motion information, and a second motion precision is applied to the above additional motion information, and The above second motion precision is a method having the resolution of integer samples.

19. A computer-readable storage medium for storing a bitstream generated by an encoding method, The above encoding method is, A step of constructing an MVP candidate list including MVP candidates for the current block; A step of generating a prediction block for the current block based on the final MVP candidate within the above MVP candidate list; and The method includes the step of encoding image information for the current block; and The step of constructing the above MVP candidate list is, Includes adding an MVP candidate with basic movement information and additional movement information to the above MVP candidate list, The above video information is, A storage medium comprising weight information applied to the weighted sum of a basic prediction block derived based on the above basic motion information and an additional prediction block derived based on the above additional motion information.

20. In a method for transmitting data regarding an image, Acquiring a bitstream for the above image, wherein the bitstream is generated based on the steps of: configuring a motion information candidate list including motion information candidates for a current block; generating a prediction block for the current block based on at least one motion information candidate in the motion information candidate list; and encoding image information for the current block; and The method includes the step of transmitting data including the bitstream above; and The step of constructing the above list of motion information candidates is, It includes adding motion information candidates having basic motion information and additional motion information to the said motion information candidate list, and The above video information is, A method comprising weight information applied to a weighted sum of a basic prediction block derived based on the basic motion information and an additional prediction block derived based on the additional motion information.

Citation Information

Patent Citations

  • Video coding method, apparatus, and non-transitory computer readable medium

    JP2024137953A

  • Battery pack, method for manufacturing battery pack, and vehicle including battery pack

    KR1020230170599A

  • Spool, tightening device with spool, and coupling method of spool and lace

    KR1020240093304A

  • Lighting apparatus and lamp of vehicle having the same

    KR1020240159867A

  • Operating method for electronic apparatus for providing service and electronic apparatus supporting thereof

    KR1020250033123A