Image coding method and apparatus using motion vectors

By signaling motion vector differences and utilizing short-term reference pictures for inter prediction, the method improves image/video coding efficiency, addressing the need for efficient compression of high-resolution and high-quality content, including VR and AR.

JP2026086606APending Publication Date: 2026-05-26LG ELECTRONICS INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2026-02-06
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, including VR and AR content, has led to a need for highly efficient image/video compression technologies to reduce transmission and storage costs while effectively handling diverse image characteristics.

Method used

A method and apparatus for improving image/video coding efficiency by signaling information regarding motion vector differences, particularly L0 and L1 motion vector differences, and performing inter prediction based on reference picture marking, including short-term reference pictures.

Benefits of technology

This approach enhances overall image/video compression efficiency by efficiently deriving and signaling motion vector differences, reducing coding complexity, and enabling efficient inter prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086606000001_ABST
    Figure 2026086606000001_ABST
Patent Text Reader

Abstract

Provides an image decoding method. [Solution] The method includes the steps of: receiving image information from a bitstream, including residual information and prediction-related information including information about MVD; deriving an interprediction mode for the current block in the current picture based on the prediction-related information; deriving motion information for the current block based on the interprediction mode; generating a prediction sample based on the motion information; generating a residual sample based on the residual information; and generating a reconstructed sample of the current picture based on the prediction sample and the residual sample. Based on the POC difference between each of the short-term reference pictures and the current picture, a symmetric motion vector difference reference index is derived; based on the information about MVD, MVDL0 for L0 prediction is derived; and based on MVDL0, MVDL1 for L1 prediction is derived.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to an image coding method and apparatus using motion vectors.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing (conventional) image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] As a result, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.

[0005] Also, inter prediction in image / video coding (encoding) can include procedures for SMVD (Symmetric Motion Vector Difference) reference indexes and / or procedures for MMVD (Merge Motion Vector Difference), and there is a discussion on a technique for performing the above procedures in consideration of reference picture marking (e.g., short-term or long-term reference).

Summary of the Invention

[0006] According to one embodiment of this document, a method and apparatus for improving image / video coding efficiency are provided.

[0007] According to one embodiment of this document, a method and apparatus for performing efficient interpretation in an image / video coding system are provided.

[0008] According to one embodiment of this paper, a method and apparatus for signaling information regarding motion vector differences in interpretation are provided.

[0009] According to one embodiment of this document, a method and apparatus are provided for signaling information regarding the L0 motion vector difference and the L1 motion vector difference when dual prediction is applied to the current block.

[0010] According to embodiments of this document, a method and apparatus for signaling the SMVD flag are provided.

[0011] According to one embodiment of this document, the prediction procedure may be performed based on the type of reference picture for biprediction.

[0012] According to one embodiment of this document, procedures relating to an SMVD reference index may be performed based on reference picture marking.

[0013] According to one embodiment of this document, procedures relating to an SMVD reference index may be performed using a short-term reference picture (a picture marked as being used for short-term referencing).

[0014] According to one embodiment of this document, a video / image decoding method is provided that is performed by a decoding device.

[0015] According to one embodiment of this document, a decoding device for performing video / image decoding is provided.

[0016] According to one embodiment of this document, a video / image encoding method is provided that is performed by an encoding device.

[0017] According to one embodiment of this document, an encoding device for performing video / image encoding is provided.

[0018] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded video / image information generated by a video / image encoding method disclosed in at least one of the embodiments of this document.

[0019] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded information or encoded video / image information, causing a decoding device to perform a video / image decoding method disclosed in at least one of the embodiments of this document. [Effects of the Invention]

[0020] According to this document, it is possible to improve the overall image / video compression efficiency.

[0021] According to this document, information regarding motion vector differences can be efficiently signaled.

[0022] According to this document, when biprediction is applied to the current block, the L1 motion vector difference can be efficiently derived.

[0023] According to this document, the information used to derive the L1 motion vector difference is signaled based on the type of reference picture, and therefore the complexity of the coding can be reduced.

[0024] According to an embodiment of this document, by using short-term reference pictures for deriving a reference picture index for SMVD, efficient inter prediction can be performed.

[0025] The effects that can be obtained through a specific example of this document are not limited to the effects listed above. For example, there can be various technical effects that a person having ordinary skill in the related art can understand or derive from this document. Accordingly, the specific effects of this document can include various effects that can be understood or derived from the technical features of this document, without being limited to those explicitly described in this document.

Brief Description of Drawings

[0026] [Figure 1] It is a diagram schematically showing an example of a video / image coding system applicable to an embodiment of this document. [Figure 2] It is a diagram schematically explaining the configuration of a video / image encoding device applicable to an embodiment of this document. [Figure 3] It is a diagram schematically explaining the configuration of a video / image decoding device applicable to an embodiment of this document. [Figure 4] It is a diagram showing an example of an inter prediction-based video / image encoding method. [Figure 5] It is a diagram showing an example of an inter prediction-based video / image decoding method. [Figure 6] It is a diagram illustratively showing an inter prediction procedure. [Figure 7] It is a diagram schematically showing a method for constructing a merge candidate list according to this document. [Figure 8] It is a diagram schematically showing a method for constructing an MVP candidate list according to this document. [Figure 9] It is a diagram explaining SMVD. [Figure 10] This diagram illustrates how to derive motion vectors in interpretation prediction. [Figure 11] This figure shows the MVD derivation (induction) process (process) of MMVD according to one embodiment of this document. [Figure 12] This figure shows the MVD derivation process for MMVD according to another embodiment of this document. [Figure 13] This figure shows the MVD derivation process for MMVD according to another embodiment of this document. [Figure 14] This figure shows the MVD derivation process of MMVD according to one embodiment of this document. [Figure 15] This figure shows the MVD derivation process of MMVD according to one embodiment of this document. [Figure 16] This figure illustrates SMVD according to one embodiment of this document. [Figure 17] This flowchart shows a method for deriving MMVD according to one embodiment of this document. [Figure 18] This flowchart shows a method for deriving MMVD according to one embodiment of this document. [Figure 19] This flowchart shows a method for deriving MMVD according to one embodiment of this document. [Figure 20] This figure schematically illustrates an example of a video / image encoding method and related components according to the embodiments of this document. [Figure 21] This figure schematically illustrates an example of a video / image encoding method and related components according to the embodiments of this document. [Figure 22] This figure schematically illustrates an example of an image / video decoding method and related components according to the embodiment of this document. [Figure 23] This figure schematically illustrates an example of an image / video decoding method and related components according to the embodiment of this document. [Figure 24] This figure shows an example of a content streaming system to which the embodiments disclosed in this document may be applied. [Modes for carrying out the invention]

[0027] The disclosures in this document can be modified in various ways and may have various embodiments, but specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit this disclosure to any particular embodiment. The terms used in this document are used solely to describe specific embodiments and are not intended to limit the technical ideas of the embodiments described herein. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this document, terms such as “includes” or “has” are intended to indicate the existence of features, figures, stages, operations, components, parts, or combinations thereof described in the document, and should be understood not to preemptively exclude the possibility of the existence or addition of one or more different features, figures, stages, operations, components, parts, or combinations thereof.

[0028] On the other hand, each configuration shown in the drawings described in this document is shown independently for the convenience of describing its distinct characteristic functions, and does not mean that each configuration is embodied in separate hardware or separate software. For example, two or more configurations may be combined to form a single configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of disclosure of this document.

[0029] The embodiments described in this document will be explained below with reference to the attached drawings. Hereafter, the same reference numerals may be used for the same components in the drawings, and redundant descriptions of the same components may be omitted.

[0030] Figure 1 schematically shows an example of a video / image coding system to which the embodiments described in this document can be applied.

[0031] As shown in Figure 1, a video / image coding system may comprise a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.

[0032] The source device may comprise a video source, an encoding device, and a transmitter. The receiving device may comprise a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be provided in the encoding device. The receiver may be provided in the decoding device. The renderer may comprise a display unit, which may consist of a separate device or external component.

[0033] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.

[0034] An encoding device can encode input video / images. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.

[0035] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include various types of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit may include elements for generating media files via a predetermined file format and elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0036] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of the encoding device.

[0037] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0038] This document relates to video / image coding. For example, the methods / embodiments disclosed herein can be applied to methods disclosed in the VVC (Versatile Video Coding) standard. Furthermore, the methods / embodiments disclosed herein can be applied to methods disclosed in the EVC (Essential Video Coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of Audio Video coding Standard), or next-generation video / image coding standards (e.g., H.267 or H.268).

[0039] This document presents various embodiments of video / image coding, and unless otherwise noted, these embodiments may be combined with each other.

[0040] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a unit representing a single image at a specific time point in time, and "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more Coding Tree Units (CTUs). A single picture can consist of one or more slices or tiles. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture, and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.

[0041] On the other hand, a single picture can be divided into two or more subpictures. A subpicture can be a rectangular region of one or more slices within a picture.

[0042] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, and can represent only the luma component pixel / pixel value, or only the chroma component pixel / pixel value.

[0043] A unit can represent a basic unit of image processing. A unit can contain at least one of a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block can contain a sample (or sample array) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0044] In this document, "A or B" may mean "just A," "just B," or "both A and B." In other words, in this document, "A or B" may be interpreted as "A and / or B." For example, in this document, "A, B or C" may mean "just A," "just B," "just C," or "any combination of A, B and C."

[0045] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "just A", "just B", or "both A and B". For example, "A, B, C" can mean "A, B or C".

[0046] In this document, "at least one of A and B" may mean "just A," "just B," or "both A and B." Furthermore, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."

[0047] Furthermore, in this document, "at least one of A, B and C" may mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0048] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."

[0049] Technical features described individually within a single drawing in this document may be embodied individually or simultaneously.

[0050] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments described in this document can be applied. Hereinafter, the term "encoding device" may include an image encoding device and / or a video encoding device.

[0051] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a reconstructor or a reconstructed block generator. The image segmentation unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 described above can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (Decoded Picture Buffer) and may be configured by a digital storage medium. The above hardware components may also include the memory 270 as an internal / external component.

[0052] The image splitting unit 210 can split the input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, coding units can be recursively split from a coding tree unit (CTU) or a largeest coding unit (LCU) using a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary-tree structure. Alternatively, the binary-tree structure may be applied first. The coding procedure according to this disclosure may be performed based on a final coding unit that cannot be further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency based on image characteristics, or, if necessary, the coding unit may be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which are described later. As another example, the processing unit may further comprise a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above.The above prediction unit may be a unit of sample prediction, and the above conversion unit may be a unit for deriving conversion coefficients and / or a unit for deriving residual signals from conversion coefficients.

[0053] The term "unit" can sometimes be confused with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.

[0054] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the predicted signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes the predicted samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or CU unit. As will be described later in the explanation of each prediction mode, the prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy coding unit 240. The prediction-related information can be encoded by the entropy coding unit 240 and output in bitstream format.

[0055] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample can be located adjacent to the current block or at a distance, depending on the prediction mode. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 can also determine the prediction mode to apply to the current block using the prediction modes applied to adjacent blocks.

[0056] The interprediction unit 221 can derive predicted blocks relative to the current block based on reference blocks (reference sample arrays) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, BI prediction, etc.). In the case of interprediction, adjacent blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the above reference blocks and the reference picture containing the above temporal neighboring blocks may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs, etc., and the reference picture containing the above temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter-prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-prediction can be performed based on various prediction modes; for example, in skip mode and merge mode, the inter-prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In Motion Vector Prediction (MVP) mode, the motion vector of the adjacent block is used as the motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0057] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called Combined Inter and Intra Prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The above IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (Screen Content Coding). IBC basically performs predictions within the current picture, but can be done similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can use at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values ​​within the picture can be signaled based on information about the palette table and palette index.

[0058] The prediction signal generated via the above prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can apply a transformation technique to the residual signal to generate transformation coefficients. For example, the transformation technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), and CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when relational information between pixels is represented in a graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and based on that. The transformation process may also be applied to pixel blocks of the same size that are square, or to non-square blocks of variable size.

[0059] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of Network Abstraction Layer (NAL) units. The video / image information may further include information about various parameter sets, such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). The video / image information may also further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded via the encoding procedure described above and included in the bitstream.The bitstream described above can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which can be configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.

[0060] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there are no residuals for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.

[0061] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied during the picture encoding and / or restoration process.

[0062] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream form.

[0063] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.

[0064] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in already restored pictures. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.

[0065] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments described in this document can be applied. Hereinafter, the term "decoding device" may include an image decoding device and / or a video decoding device.

[0066] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-predictor 331 and an inter-predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering device 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may also include a Decoded Picture Buffer (DPB) and can be configured by a digital storage medium. The above hardware components may also be further equipped with Memory 360 as an internal / external component.

[0067] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image corresponding to the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transformation units can be derived from the coding unit. The reconstructed image signal decoded and output via the decoding device 300 can then be reproduced via a playback device.

[0068] The decoding device 300 can receive the signal output from the encoding device shown in Figure 2 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image reconstruction and the quantized values ​​of transformation coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoded information of adjacent and decoded blocks or symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, performs arithmetic decoding of the bins, and generates symbols corresponding to the values ​​of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values ​​after entropy decoding by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transform unit 322, adder 340, filtering unit 350, memory 360, inter-prediction unit 332, and intra-prediction unit 331.

[0069] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0070] The inverse transform unit 322 performs an inverse transform on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0071] The prediction unit can make predictions for the current block and generate a predicted block containing the predicted samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.

[0072] The prediction unit 330 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called Combined Inter and Intra Prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The above IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (Screen Content Coding). IBC basically performs predictions within the current picture, but can be done similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the above video / image information and signaled.

[0073] The intra-prediction unit 331 can predict the current block by referring to a sample within the current picture. The referenced sample can be located adjacent to or far from the current block, depending on the prediction mode. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to adjacent blocks.

[0074] The interprediction unit 332 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, BI prediction, etc.). In the case of interprediction, adjacent blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the information regarding the prediction may include information indicating the mode of interprediction for the current block.

[0075] The summing unit 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (which comprises an inter-prediction unit 332 and / or an intra-prediction unit 331). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block.

[0076] The summing unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.

[0077] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied during the picture decoding process.

[0078] The filtering unit 350 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0079] The restored picture stored (modified) in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in already restored pictures. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.

[0080] Embodiments described herein for the filtering unit 260, inter-prediction unit 221, and intra-prediction unit 222 of the encoding device 200 can be applied identically or in a corresponding manner to the filtering unit 350, inter-prediction unit 332, and intra-prediction unit 331 of the decoding device 300, respectively.

[0081] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived in both the encoding and decoding devices, and the encoding device can improve image coding efficiency by signaling the decoding device information about the residuals between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can generate a restored block containing restored samples by combining the residual block and the predicted block, thereby generating a restored picture containing the restored block.

[0082] The residual information described above can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information to a decoding device (via a bitstream) by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and then performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive residual samples (or residual blocks) by performing an inverse quantization / inverse transformation procedure based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.

[0083] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If the above quantization / inverse quantization is omitted, the above quantized transformation coefficients may be called transformation coefficients. If the above transformation / inverse transformation is omitted, the above transformation coefficients may also be called coefficients or residual coefficients, or for consistency of expression, they may still be called transformation coefficients.

[0084] In this document, quantized transformation coefficients and transformation coefficients can be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information can include information about the transformation coefficient(s), and such information about the transformation coefficient(s) can be signaled via residual coding syntax. Based on the residual information (or information about the transformation coefficient(s)), the transformation coefficient(s) can be derived, and the scaled transformation coefficient(s) can be derived via an inverse transformation (scaling) of the transformation coefficient(s). Based on an inverse transformation (transformation) of the scaled transformation coefficient(s), residual samples can be derived. This can be applied / expressed similarly in other parts of this document.

[0085] Intra prediction can represent a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived. The adjacent reference samples of the current block described above may include a total of 2 × nH samples adjacent to the left boundary and bottom-left of the current block of size nW × nH, a total of 2 × nW samples adjacent to the top boundary and top-right of the current block, and one sample adjacent to the top-left of the current block. Alternatively, the adjacent reference samples of the current block described above may also include multiple columns of upper adjacent samples and multiple rows of left adjacent samples. Furthermore, the adjacent reference samples of the current block may also include a total of nH samples adjacent to the right boundary of the current block, which has a size of nW × nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right side of the current block.

[0086] However, some of the adjacent reference samples in the current block may not yet be decoded or available. In this case, the decoder can construct adjacent reference samples to be used for prediction by substituting the unavailable samples as available samples. Alternatively, it can construct adjacent reference samples to be used for prediction through interpolation of available samples.

[0087] If neighboring reference samples are derived, (i) predicted samples can be derived based on the average or interpolation of neighboring reference samples in the current block, or (ii) predicted samples can be derived based on neighboring reference samples in the current block that are located in a specific (predicted) direction relative to the predicted sample. Case (i) is called the non-directional mode or non-angular mode, and case (ii) is called the directional mode or angular mode.

[0088] Furthermore, among the adjacent reference samples mentioned above, the predicted sample can also be generated by interpolation between a first adjacent sample located in the prediction direction of the intra-prediction mode of the current block and a second adjacent sample located in the opposite direction of the prediction direction, based on the predicted sample of the current block. In the case described above, this can be called Linear Interpolation Intra Prediction (LIP). Alternatively, chroma prediction samples can be generated based on chroma samples using a linear model. In this case, this can be called LM mode.

[0089] Alternatively, a temporary predicted sample for the current block can be derived based on the filtered adjacent reference samples, and the predicted sample for the current block can be derived by performing a weighted sum of the temporary predicted sample and at least one reference sample derived by the intra-prediction mode from the existing adjacent reference samples, i.e., the unfiltered adjacent reference samples. In the case described above, this can be called PDPC (Position Dependent Intra Prediction).

[0090] Alternatively, intra-predictive coding can be performed by selecting the reference sample line with the highest prediction accuracy (accuracy) from among the adjacent multi-reference sample lines in the current block, deriving a predicted sample using the reference sample located in the prediction direction on that line, and then instructing (signaling) the decoding device with the reference sample line used at that time. In the case described above, this can be called multi-reference line intra-prediction or MRL-based intra-prediction.

[0091] Furthermore, the current block can be divided into vertical or horizontal subpartitions, and intra-prediction can be performed based on the same intra-prediction mode. Adjacent reference samples can then be derived and used on a subpartition-by-subpartition basis. In other words, in this case, the intra-prediction mode for the current block is also applied to the subpartition, and by deriving and using adjacent reference samples on a subpartition-by-subpartition basis, intra-prediction performance can be improved in some cases. Such a prediction method can be called ISP (Intra Sub-Partitions) based intra-prediction.

[0092] The intra-prediction methods described above can be distinguished from intra-prediction modes and called intra-prediction types. These intra-prediction types can be referred to by a variety of terms, such as intra-prediction techniques or additional intra-prediction modes. For example, the above intra-prediction types (or additional intra-prediction modes, etc.) can include at least one of the above-mentioned LIP, PDPC, MRL, and ISP. General intra-prediction methods that exclude specific intra-prediction types such as LIP, PDPC, MRL, and ISP can be called normal intra-prediction types. Normal intra-prediction types can be generally applied when the above-mentioned specific intra-prediction types are not applicable, and predictions can be performed based on the aforementioned intra-prediction modes. On the other hand, post-processing filtering can also be performed on the derived prediction samples as needed.

[0093] Specifically, the intra-prediction procedure may include an intra-prediction mode / type determination step, an adjacent reference sample derivation step, and an intra-prediction mode / type-based predictive sample derivation step. Additionally, a post-filtering step may be performed on the derived predictive samples as needed.

[0094] When intraprediction is applied, the intraprediction mode applied to the current block can be determined by utilizing the intraprediction modes of adjacent blocks. For example, the decoder can select one of the MPM candidates in the Most Probable Mode (MPM) list derived based on the intraprediction modes of the current block's adjacent blocks (e.g., left and / or upper adjacent blocks) and additional candidate modes, based on the received MPM index, or it can select one of the remaining intraprediction modes not included in the above MPM candidates (and planar modes), based on the remaining (remaining) intraprediction mode information. The above MPM list may or may not include planar modes as candidates. For example, if the above MPM list includes planar modes as candidates, the above MPM list may have 6 candidates, and if the above MPM list does not include planar modes as candidates, the above MPM list may have 5 candidates. If the above MPM list does not include planar mode as a candidate, a not-planar flag (e.g., intra_luma_not_planar_flag) may be signaled to indicate that the intra-prediction mode of the current block is not planar mode. For example, the MPM flag may be signaled first, and the MPM index and not-planar flag may be signaled if the value of the MPM flag is 1. Also, the above MPM index may be signaled if the value of the not-planar flag is 1. Here, the reason why the above MPM list is configured not to include planar mode as a candidate is not because the above planar mode is not an MPM, but because planar mode is always considered as an MPM, so the flag (not-planar flag) is signaled first to check whether it is planar mode or not.

[0095] For example, whether the intra-prediction mode applied to the current block is within the MPM candidate (and planar mode) or within the remaining modes can be indicated based on the MPM flag (e.g., intra_luma_mpm_flag). A value of 1 for the MPM flag indicates that the intra-prediction mode for the current block is within the MPM candidate (and planar mode), and a value of 0 for the MPM flag indicates that the intra-prediction mode for the current block is not within the MPM candidate (and planar mode). A value of 0 for the not-planar flag (e.g., intra_luma_not_planar_flag) indicates that the intra-prediction mode for the current block is planar mode, and a value of 1 for the not-planar flag indicates that the intra-prediction mode for the current block is not planar mode. The above MPM index can be signaled in the form of mpm_idx or intra_luma_mpm_idx syntax elements, and the above remaining intra-prediction mode information can be signaled in the form of rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax elements. For example, the above remaining intra-prediction mode information can refer to one of the remaining intra-prediction modes that are not included in the above MPM candidate (and planar mode) from all intra-prediction modes, indexed in order of prediction mode number. The above intra-prediction mode is the intra-prediction mode for the luma component (sample). The intra-prediction mode information below may include at least one of the following: the MPM flag (e.g., intra_luma_mpm_flag), the not-planar flag (e.g., intra_luma_not_planar_flag), the MPM index (e.g., mpm_idx or intra_luma_mpm_idx), or the remaining intra-prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list may be referred to by various terms such as the MPM candidate list or candModeList.When MIP is applied to the current block, a separate mpm flag for MIP (e.g., intra_mip_mpm_flag), an mpm index (e.g., intra_mip_mpm_idx), and the remaining intra predictive mode information (e.g., intra_mip_mpm_remainder) may be signaled, while the above not planar flag is not signaled.

[0096] In other words, when an image is generally divided into blocks, the current block to be coded and its neighboring blocks will have similar image characteristics. Therefore, there is a high probability that the current block and its neighboring blocks will have the same or similar intra-prediction modes. Consequently, the encoder can utilize the intra-prediction mode of the neighboring block to encode the intra-prediction mode of the current block.

[0097] For example, an encoder / decoder can configure an MPM (Most Probable Modes) list for the current block. This MPM list can also be referred to as an MPM candidate list. Here, MPM can mean a mode used to improve coding efficiency during intra predictive mode coding by considering the similarity between the current block and adjacent blocks. As mentioned above, the MPM list can be configured to include planar modes, or it can be configured to exclude planar modes. For example, if the MPM list includes planar modes, the number of candidates in the MPM list is 6. If the MPM list does not include planar modes, the number of candidates in the MPM list is 5.

[0098] The encoder / decoder can configure an MPM list containing five or six MPMs.

[0099] To construct an MPM list, three types of modes can be considered: default intra modes, neighbor intra modes, and derived intra modes.

[0100] For the adjacent intra-mode described above, two adjacent blocks, namely the left adjacent block and the upper adjacent block, can be considered.

[0101] As mentioned above, if the MPM list is configured not to include planar mode, then planar mode is excluded from the list, and the number of candidates for the MPM list can be set to 5.

[0102] Furthermore, among the intra-prediction modes, non-directional modes (or non-angle modes) may include average-based DC modes or interpolation-based planar modes of the neighboring reference samples of the current block.

[0103] When interpretation is applied, the prediction unit of the encoding / decoding device can perform interpretation on a block-by-block basis to derive predicted samples. Interpretation can indicate predictions derived in a way that depends on data elements of pictures other than the current picture (e.g., sample values ​​or motion information). When interpretation is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the index of the reference picture. In this case, in order to reduce the amount of motion information transmitted in interpretation mode, the motion information of the current block can be predicted on a block, subblock, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The above motion information may include motion vectors and the index of the reference picture. The above motion information may further include information on the interpretation type (L0 prediction, L1 prediction, BI prediction, etc.). When interpretation is applied, neighboring blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the above-mentioned reference block and the reference picture containing the above-mentioned time-adjacent block may be the same or different. The above-mentioned time-adjacent block may be called a collocated reference block (colCU), and the reference picture containing the above-mentioned time-adjacent block may be called a collocated picture (colPic). For example, a candidate list of motion information may be constructed based on the adjacent blocks of the current block, and flags or index information may be signaled to indicate which candidate is selected (used) in order to derive the motion vector and / or the index of the reference picture of the current block.Interpretation is performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be identical to the motion information of the selected adjacent block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In Motion Vector Prediction (MVP) mode, the motion vector of the selected adjacent block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0104] The above motion information may include L0 motion information and / or L1 motion information, depending on the interpretation type (L0 prediction, L1 prediction, BI prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both the above L0 motion vector and the above L1 motion vector may be called a bi(Bi) prediction. Here, an L0 motion vector may represent a motion vector associated with the reference picture list L0(L0), and an L1 motion vector may represent a motion vector associated with the reference picture list L1(L1). The reference picture list L0 may include pictures earlier in the output order than the above current picture, and the reference picture list L1 may include pictures later in the output order than the above current picture. Pictures prior to the above may be called forward (reference) pictures, and pictures after the above may be called reverse (reference) pictures. The above reference picture list L0 may include pictures that are later in the output order than the above current picture as reference pictures. In this case, the above earlier pictures may be indexed first in the above reference picture list L0, and the above later pictures may be indexed afterward. The above reference picture list L1 may include pictures that are earlier in the output order than the above current picture as reference pictures. In this case, the above later pictures may be indexed first in the above reference picture list L1, and the above earlier pictures may be indexed afterward. Here, the output order may correspond to the POC (Picture Order Count) order.

[0105] The video / image encoding procedure based on interpretation broadly includes, for example, the following:

[0106] Figure 4 shows an example of an interpretation-based video / image encoding method.

[0107] The encoding device performs interpretation for the current block (S400). The encoding device derives the interpretation mode and motion information of the current block and generates prediction samples for the block. Here, the procedures for determining the interpretation mode, deriving motion information, and generating prediction samples may be performed simultaneously, or one procedure may be performed before the others. For example, the interpretation unit of the encoding device includes a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block, the motion information derivation unit derives the motion information of the current block, and the prediction sample derivation unit derives prediction samples for the current block. For example, the interpretation unit of the encoding device searches for blocks similar to the current block within a certain area (search area) of the reference picture by motion estimation and derives reference blocks whose difference from the current block is the minimum or below a certain standard. Based on this, a reference picture index pointing to the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device determines which of the various prediction modes is applied to the current block. The encoding device can compare the RD costs for the various prediction modes and determine the optimal prediction mode for the current block.

[0108] For example, when skip mode or merge mode is applied to the current block, the encoding device can configure a merge candidate list, as described later, and derive a reference block from among the reference blocks pointed to by the merge candidates included in the merge candidate list whose difference from the current block is the minimum or below a certain standard. In this case, a merge candidate associated with the derived reference block is selected, and merge index information pointing to the selected merge candidate is generated and signaled to the decoding device. The movement information of the current block can be derived using the movement information of the selected merge candidate.

[0109] As another example, when the (A)MVP mode is applied to the current block, the encoding device can configure the (A)MVP candidate list described later, and use the motion vector of the selected mvp candidate from among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the motion estimation described above can be used as the motion vector of the current block, and the mvp candidate with the smallest difference from the motion vector of the current block among the mvp candidates can become the selected mvp candidate. The MVD (Motion Vector Difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In that case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index is composed of reference picture index information and is separately signaled to the decoding device.

[0110] The encoding device derives a residual sample based on the predicted sample (S410). The encoding device can derive the residual sample by comparing the original sample of the current block with the predicted sample.

[0111] The encoding device encodes image information including prediction information and residual information (S420). The encoding device outputs the encoded image information in bitstream format. The prediction information is information related to the prediction procedure and includes prediction mode information (e.g., skip flag, merge flag, or mode index) and motion information. The motion information includes candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving the motion vector. The motion information also includes the aforementioned MVD information and / or reference picture index information. The motion information also includes information indicating whether L0 prediction, L1 prediction, or bi prediction is applied. The residual information is information about the residual sample. The residual information includes information about the quantized transformation coefficients for the residual sample.

[0112] The output bitstream may be stored in a (digital) storage medium and transmitted to a decoding device, or it may be transmitted to a decoding device via a network.

[0113] On the other hand, as mentioned above, the encoding device generates a reconstructed picture (including the reconstructed sample and the reconstructed block) based on the reference sample and the residual sample. This is to derive the same prediction results from the encoding device as those performed by the decoding device, thereby improving coding efficiency. Therefore, the encoding device can store the reconstructed picture (or the reconstructed sample and the reconstructed block) in memory and use it as a reference picture for interpretation. As mentioned above, in-loop filtering procedures and the like can be further applied to the reconstructed picture.

[0114] The video / image decoding procedure based on interpretation broadly includes, for example, the following:

[0115] Figure 5 shows an example of an interpretation-based video / image decoding method.

[0116] As shown in Figure 5, the decoding device performs operations corresponding to those performed by the encoding device. Based on the received prediction information, the decoding device can make predictions in the current block and derive prediction samples.

[0117] Specifically, the decoding device determines the prediction mode for the current block based on the received prediction information (S500). The decoding device can then determine which inter-prediction mode is applied to the current block based on the prediction mode information within the prediction information.

[0118] For example, based on the merge flag, it is possible to determine whether the merge mode is applied to the current block, or whether the (A)MVP mode is determined. Alternatively, one of the various inter-prediction mode candidates can be selected based on the mode index. The inter-prediction mode candidates include skip mode, merge mode and / or (A)MVP mode, or include the various inter-prediction modes described below.

[0119] The decoding device derives motion information for the current block based on the determined inter prediction mode (S510). For example, if a skip mode or merge mode is applied to the current block, the decoding device configures a merge candidate list, which will be described later, and selects one of the merge candidates included in the merge candidate list. The above selection is made based on the selection information (merge index) described above. The motion information for the current block can be derived using the motion information for the selected merge candidate. The motion information for the selected merge candidate can be used as the motion information for the current block.

[0120] As another example, when the (A)MVP mode is applied to the current block, the decoding device can configure the (A)MVP candidate list described below, and use the motion vector of the selected mvp candidate from among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. The above selection is made based on the selection information (mvp flag or mvp index) mentioned above. In this case, the MVD of the current block can be derived based on the information about the MVD, and the motion vector of the current block can be derived based on the mvp of the current block and the MVD. In addition, the reference picture index of the current block can be derived based on the reference picture index information. In the reference picture list for the current block, the picture pointed to by the reference picture index can be derived as the reference picture referenced for inter prediction of the current block.

[0121] On the other hand, as described later, the motion information of the current block can be derived without constructing a candidate list. In this case, the motion information of the current block can be derived according to the procedure disclosed in the prediction mode described later. In this case, the candidate list configuration described above may be omitted.

[0122] The decoding device generates predicted samples for the current block based on the motion information of the current block (S520). In this case, the reference picture can be derived based on the reference picture index of the current block, and the predicted samples for the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, a further procedure of filtering all or some of the predicted samples for the current block may be performed.

[0123] For example, the interpretation unit of the decoding device includes a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block based on the prediction mode information received, the motion information derivation unit derives motion information (such as motion vectors and / or reference picture indices) for the current block based on the motion information received, and the prediction sample derivation unit derives prediction samples for the current block.

[0124] The decoding device generates residual samples for the current block based on the received residual information (S530). The decoding device generates reconstructed samples for the current block based on the predicted samples and residual samples, and generates a reconstructed picture based on these (S540). As mentioned above, in-loop filtering procedures and the like can be further applied to the reconstructed picture thereafter.

[0125] Figure 6 illustrates the interpretation prediction procedure.

[0126] Referring to Figure 6, as described above, the interpretation procedure includes an interpretation mode determination step, a motion information derivation step corresponding to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The above interpretation procedure is performed in the encoding device and the decoding device, as described above. In this document, the coding device includes the encoding device and / or the decoding device.

[0127] As shown in Figure 6, the coding device determines the interpretation mode for the current block (S600). Various interpretation modes can be used to predict the current block in the picture. For example, various modes such as merge mode, skip mode, MVP (Motion Vector Prediction) mode, affine mode, subblock merge mode, and MMVD (Merge with MVD) mode can be used. DMVR (Decoder side Motion Vector Refinement) mode, AMVR (Adaptive Motion Vector Resolution) mode, Bi-prediction with CU-level Weight (BCW), and Bi-Directional Optical Flow (BDOF) can be used as additional or alternative modes. The affine mode may also be called affine motion prediction mode. The MVP mode may also be called the AMVP (Advanced Motion Vector Prediction) mode. In this document, motion information candidates derived from some modes and / or some modes may be included as one of the motion information related candidates from other modes. For example, an HMVP candidate may be added as a merge candidate in the merge / skip mode described above, or as an mvp candidate in the MVP mode described above. When the HMVP candidate is used as a motion information candidate in the merge mode or skip mode described above, the HMVP candidate may also be called an HMVP merge candidate.

[0128] Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information includes index information indicating one of a number of candidate modes. Alternatively, the inter-prediction mode can be indicated via hierarchical signaling of flag information. In this case, the prediction mode information includes one or more flags. For example, a skip flag may be signaled to indicate whether a skip mode is applied, a merge flag may be signaled if the skip mode is not applied to indicate whether the merge mode is applied, and if the merge mode is not applied, it may indicate that the MVP mode is applied, or flags for additional distinctions may be further signaled. Affine modes may be signaled as independent modes, or as modes dependent on the merge mode or MVP mode, etc. For example, affine modes include affine merge mode and affine MVP mode.

[0129] On the other hand, the current block may be signaled with information indicating whether the aforementioned List0(L0) prediction, List1(L1) prediction, or bi-prediction is used in the current block (current coding unit). This information may also be called motion prediction direction information, inter-prediction direction information, or inter-prediction instruction information, and can be constructed / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether the aforementioned List0(L0) prediction, List1(L1) prediction, or bi-prediction is used in the current block (current coding unit). For the sake of explanation, in this document, the inter-prediction type (L0 prediction, L1 prediction, or BI prediction) pointed to by the inter_pred_idc syntax element may be expressed as motion prediction direction. L0 prediction may be expressed as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, the following prediction types can be indicated by the value of the inter_pred_idc syntax element:

[0130] [Table 1]

[0131] As mentioned above, a single picture contains one or more slices. A slice can have one of the following slice types: I (Intra) slices, P (Predictive) slices, and B (Bi-Predictive) slices. The slice type is indicated based on the slice type information. For blocks in an I slice, only intra prediction is used for prediction, and inter-predictive prediction is not used. Of course, in this case as well, it is possible to code and signal the original sample values ​​without prediction. For blocks in a P slice, either intra prediction or inter-predictive prediction is used, and if inter-predictive prediction is used, only uni prediction may be used. On the other hand, for blocks in a B slice, either intra prediction or inter-predictive prediction is used, and if inter-predictive prediction is used, up to bi-predictive prediction may be used.

[0132] L0 and L1 contain reference pictures that were encoded / decoded before the current picture. For example, L0 contains reference pictures that are earlier and / or later than the current picture in the POC order, and L1 contains reference pictures that are later and / or earlier than the current picture in the POC order. In this case, L0 is assigned a reference picture index that is lower relative to the reference pictures that are earlier than the current picture in the POC order, and L1 is assigned a reference picture index that is lower relative to the reference pictures that are later than the current picture in the POC order. For B slices, biprediction is applied, and in this case, unidirectional biprediction may be applied, or bidirectional biprediction may be applied. Bidirectional biprediction is also called true biprediction.

[0133] The following table shows the syntax for a coding unit according to one embodiment of this document.

[0134] [Table 2-1]

[0135] [Table 2-2]

[0136] [Table 2-3]

[0137] [Table 2-4]

[0138] [Table 2-5]

[0139] The coding device derives motion information for the current block (S610). The derivation of the motion information can be done based on the inter-prediction mode.

[0140] The coding device can perform interpretation using motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block in the original picture for the current block to search for highly correlated similar reference blocks in fractional pixel units within a defined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference in phase-based sample values. For example, block similarity can be calculated based on the Sum of Absolute Differences (SAD) between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information is signaled to the decoding device in various ways based on the interpretation prediction mode.

[0141] The coding device performs inter prediction based on the motion information for the current block (S620). The coding device can derive one or more predicted samples for the current block based on the motion information. The current block containing the predicted samples may be called a predicted block.

[0142] When merge mode is applied, the movement information of the current predicted block is not transmitted directly, but rather derived using the movement information of surrounding predicted blocks. Therefore, the movement information of the current predicted block can be specified by transmitting flag information indicating that merge mode was used and a merge index indicating which surrounding predicted block was used. This merge mode may also be called regular merge mode.

[0143] To perform a merge, the encoder must search for merge candidate blocks to be used to derive motion information for the current predicted block. For example, up to five merge candidate blocks may be used, but the embodiments described in this document are not limited to this. The maximum number of merge candidate blocks is transmitted in the slice header or tile group header. After finding the merge candidate blocks, the encoder can generate a merge candidate list and select the merge candidate block with the lowest cost among them as the final merge candidate block.

[0144] The above merge candidate list can utilize, for example, five merge candidate blocks. For example, it can utilize four spatial merge candidates and one temporal merge candidate. Hereafter, the above spatial merge candidate, or the spatial MVP candidate described later, may be called SMVP, and the above temporal merge candidate, or the temporal MVP candidate described later, may be called TMVP.

[0145] Figure 7 schematically shows how to construct the merge candidate list related to this document.

[0146] The coding device (encoder / decoder) searches for spatially surrounding blocks of the current block and inserts the derived spatial merge candidates into the merge candidate list (S700). For example, the spatially surrounding blocks include the block around the lower left corner of the current block, the left side surrounding block, the upper right corner surrounding block, the upper side surrounding block, and the upper left corner surrounding block. However, this is an example, and additional surrounding blocks such as the right side surrounding block, the lower side surrounding block, and the lower right side surrounding block can also be used as spatially surrounding blocks. The coding device searches for the spatially surrounding blocks based on priority to detect usable blocks and can derive the movement information of the detected blocks as spatial merge candidates.

[0147] The coding device inserts the time merge candidates derived by searching for the time neighboring blocks of the current block into the merge candidate list (S710). The time neighboring blocks may be located on a reference picture that is a picture different from the current picture in which the current block is located. The reference picture on which the time neighboring blocks are located may be referred to as a collocated picture or a col picture. The time neighboring blocks can be searched in the order of the peripheral blocks of the lower right corner and the lower right center block of the co-located block with respect to the current block on the col picture. On the other hand, when motion data compression is applied, specific motion information is stored as representative motion information for each fixed storage unit in the col picture. In this case, it is not necessary to store the motion information for all blocks within the fixed storage unit, and thus the effect of motion data compression can be obtained. In this case, the fixed storage unit may be predetermined, for example, in units of 16×16 samples or 8×8 samples, or the size information regarding the fixed storage unit may be signaled from the encoder to the decoder. When motion data compression is applied, the motion information of the time neighboring blocks can be replaced with the representative motion information of the fixed storage unit in which the time neighboring blocks are located. That is, in this case, from the aspect of implementation, instead of the prediction block located at the coordinates of the time neighboring blocks, based on the coordinates (the upper left sample position (position)) of the time neighboring blocks, after arithmetic right shift by a certain value, the time merge candidate is derived based on the motion information of the prediction block covering the position after arithmetic left shift. For example, when the fixed storage unit is a 2n×2n sample unit, if the coordinates of the time neighboring blocks are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>n)<<n), (yTnb>>n)<<n)) is used for the time merge candidate.Specifically, for example, if the above-mentioned fixed memory unit is 16x16 samples, and the coordinates of the time-related block are (xTnb, yTnb), then the movement information of the predicted block located at the corrected position ((xTnb>>4)<<4), (yTnb>>4)<<4)) is used for the time merge candidate. Alternatively, for example, if the above-mentioned fixed memory unit is 8x8 samples, and the coordinates of the time-related block are (xTnb, yTnb), then the movement information of the predicted block located at the corrected position ((xTnb>>3)<<3), (yTnb>>3)<<3)) is used for the time merge candidate.

[0148] The coding device can check whether the current number of merge candidates is less than the maximum number of merge candidates (S720). The maximum number of merge candidates can be predefined or signaled from the encoder to the decoder. For example, the encoder generates information about the maximum number of merge candidates, encodes it, and transmits it to the decoder in bitstream form. Once the maximum number of merge candidates is reached (filled), the subsequent candidate addition process does not need to be performed.

[0149] If, as a result of the above check, the current number of merge candidates is less than the maximum number of merge candidates, the coding device inserts additional merge candidates into the merge candidate list (S730).

[0150] If, as a result of the above check, the current number of merge candidates is not less than the maximum number of merge candidates, the coding device terminates the configuration of the merge candidate list (S740). In this case, the encoder can select the optimal merge candidate from among the merge candidates that make up the merge candidate list based on the RD (Rate-Distortion) cost, and can signal selection information (e.g., merge index) pointing to the selected merge candidate to the decoder. The decoder selects the optimal merge candidate based on the merge candidate list and the selection information.

[0151] As previously mentioned, the motion information of the selected merge candidate can be used as the motion information of the current block, and predicted samples of the current block can be derived based on the motion information of the current block. The encoder can derive residual samples of the current block based on the predicted samples and signal residual information regarding the residual samples to the decoder. As previously mentioned, the decoder can generate reconstructed samples based on the residual samples derived from the residual information and the predicted samples, and generate a reconstructed picture based on these.

[0152] When skip mode is applied, the motion information of the current block can be derived in the same way as when merge mode is applied. However, when skip mode is applied, the residual signal for the block in question is omitted, and therefore, the predicted sample can be immediately used as the reconstructed sample.

[0153] When MVP mode is applied, a motion vector predictor (MVP) candidate list is generated using the motion vectors of the restored spatial surrounding blocks and / or the motion vectors corresponding to the temporal surrounding blocks (or Col blocks). That is, the motion vectors of the restored spatial surrounding blocks and / or the motion vectors corresponding to the temporal surrounding blocks can be used as motion vector predictor candidates. When dual prediction is applied, an MVP candidate list for L0 motion information derivation and an MVP candidate list for L1 motion information derivation can be generated and used separately. The aforementioned prediction information (or information about prediction) includes selection information (e.g., an MVP flag or MVP index) that indicates the selected optimal motion vector predictor candidate from among the motion vector predictor candidates included in the above list. Here, the prediction unit can use the above selection information to select the motion vector predictor for the current block from among the motion vector predictor candidates included in the motion vector candidate list. The prediction unit of the encoding device can calculate the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode it, and output it in bitstream format. In other words, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, the prediction unit of the decoding device can obtain the motion vector difference included in the prediction information and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction unit of the decoding device can obtain or derive the reference picture index that indicates the reference picture from the prediction information.

[0154] Figure 8 is a flowchart (sequence diagram) showing how to construct the motion vector predictor candidate list.

[0155] As shown in Figure 8, in one embodiment, first, spatial candidate blocks for motion vector prediction are searched for and inserted into the prediction candidate list (S800). Subsequently, in one embodiment, it is determined whether the number of spatial candidate blocks is less than 2 (S810). For example, in one embodiment, if the number of spatial candidate blocks is less than 2, time candidate blocks are searched for and added to the prediction candidate list (S820), and if time candidate blocks are unavailable, zero motion vectors are used. That is, zero motion vectors can be added to the prediction candidate list (S830). Subsequently, in one embodiment, the configuration of the preliminary candidate list is terminated (S840). Alternatively, in one embodiment, if the number of spatial candidate blocks is not less than 2, the configuration of the preliminary candidate list is terminated (S840). Here, the preliminary candidate list refers to the MVP candidate list.

[0156] On the other hand, when MVP mode is applied, the reference picture index is explicitly signaled. In this case, the reference picture index for L0 prediction (refidxL0) and the reference picture index for L1 prediction (refidxL1) can be signaled separately. For example, when MVP mode is applied and bi-indicator prediction (BI prediction) is applied, both the information regarding refidxL0 and the information regarding refidxL1 can be signaled.

[0157] When MVP mode is applied, as described above, information about the MVD derived from the encoding device is signaled to the decoding device. The information about the MVD may include, for example, information indicating the x and y components for the absolute value and sign of the MVD. In this case, information indicating whether the absolute value of the MVD is greater than 0, whether it is greater than 1, and the remainder of the MVD can be signaled stepwise. For example, information indicating whether the absolute value of the MVD is greater than 1 can be signaled only if the value of the flag information indicating whether the absolute value of the MVD is greater than 0 is 1.

[0158] For example, information regarding MVD is composed of the syntax shown in the table below, encoded in the encoding device, and then signaled to the decoding device.

[0159] [Table 3]

[0160] For example, in Table 3, the abs_mvd_greater0_flag syntax element indicates whether the difference (MVD) is greater than 0, and the abs_mvd_greater1_flag syntax element indicates whether the difference (MVD) is greater than 1. Furthermore, the abs_mvd_minus2 syntax element indicates the value obtained by subtracting 2 from the difference (MVD), and the mvd_sign_flag syntax element indicates the sign of the difference (MVD). Also in Table 3, [0] for each syntax element indicates information about L0, and [1] indicates information about L1.

[0161] For example, MVD[compIdx] is derived based on abs_mvd_greater0_flag[compIdx]*(abs_mvd_minus2[compIdx]+2)*(1-2*mvd_sign_flag[compIdx]). Here, compIdx (or cpIdx) represents the index of each component and can have a value of 0 or 1. compIdx0 represents the x component, and compIdx1 represents the y component. However, this is illustrative, and values ​​can also be represented for each component using other coordinate systems instead of the x,y coordinate system.

[0162] On the other hand, MVDs for L0 prediction (MVD L0) and MVDs for L1 prediction (MVD L1) may be signaled separately, and the information regarding the MVD may include information regarding MVD L0 and / or information regarding MVD L1. For example, if MVP mode is applied to the current block and BI prediction is applied, both the information regarding MVD L0 and the information regarding MVD L1 will be signaled.

[0163] Figure 9 illustrates SMVD (Symmetric Motion Vector Differences).

[0164] When BI prediction is applied, Symmetric MVD (SMVD) may be used to improve coding efficiency. In this case, some of the signaling of motion information may be omitted. For example, when SMVD is applied to the current block, information regarding refidxL0, refidxL1, and MVD L1 can be derived internally without being signaled from the encoding device to the decoding device. For example, when MVP mode and BI prediction are applied to the current block, flag information (e.g., SMVD flag information or sym_mvd_flag syntax element) indicating whether SMVD can be applied is signaled, and if the value of the above flag information is 1, the decoding device determines that SMVD is applied to the above block.

[0165] When SMVD mode is applied (i.e., when the value of the SMVD flag information is 1), information regarding mvp_l0_flag, mvp_l1_flag, and MVD L0 (motion vector difference L0) is explicitly signaled, and as mentioned above, signaling for information regarding refidxL0, refidx1, and MVD L1 (motion vector difference L1) is omitted and can be derived internally. For example, refidxL0 can be derived in the POC procedure within reference picture list 0 (which may be called List0 or L0) as an index pointing to the closest previously referenced picture to the current picture. refidxL1 can be derived in the POC procedure within reference picture list 1 (which may be called List1 or L1) as an index pointing to the closest subsequent referenced picture to the current picture. Alternatively, for example, both refidxL0 and refidxL1 can be derived as 0. Alternatively, for example, refidxL0 and refidxL1 can be derived as the smallest index having the same POC difference in relation to the current picture. Specifically, for example, if "[POC of the current picture] - [POC of the first referenced picture indicated by refidxL0]" is called the first POC difference, and "[POC of the current picture] - [POC of the second referenced picture indicated by refidxL1]" is called the second POC difference, then only if the first POC difference and the second POC difference are the same, the value of refidxL0 pointing to the first referenced picture may be derived as refidxL0 of the current block, and the value of refidxL1 pointing to the second referenced picture may be derived as refidxL1 of the current block. Furthermore, for example, if there are multiple sets where the first POC difference and the second POC difference are identical, the refidxL0 and refidxL1 of the set with the smallest difference can be derived as the refidxL0 and refidxL1 of the current block.

[0166] As shown in Figure 9, reference picture list 0, reference picture list 1, and MVD L0 and MVD L1 are shown. Here, MVD L1 is symmetrical to MVD L0.

[0167] MVD L1 can be derived as minus (-)MVD L0. For example, the final (improved or modified) motion information (motion vector: MV) for the current block is derived based on the following formula:

[0168] <Formula 1>

number

[0169] In Equation 1, mvx0 and mvy0 represent the x and y components of the motion vector for L0 motion information or L0 prediction, and mvx1 and mvy1 represent the x and y components of the motion vector for L1 motion information or L1 prediction. Furthermore, mvpx0 and mvpy0 represent the x and y components of the motion vector predictor for L0 prediction, and mvpx1 and mvpy1 represent the x and y components of the motion vector predictor for L1 prediction. In addition, mvdx0 and mvdy0 represent the x and y components of the motion vector difference for L0 prediction.

[0170] On the other hand, MMVD mode allows motion information, which is directly used to generate predicted samples for the current block (i.e., the current CU), to be implicitly derived as a way to apply MVD (Motion Vector Difference) to merge mode. For example, an MMVD flag (e.g., mmvd_flag) indicating whether or not to use MMVD for the current block (i.e., the current CU) is signaled, and MMVD can be performed based on this MMVD flag. If MMVD is applied to the current block (e.g., mmvd_flag is 1), additional information regarding MMVD can be signaled.

[0171] Here, additional information regarding MMVD includes a merge candidate flag (e.g., mmvd_cand_flag) indicating whether the first or second candidate in the merge candidate list is used with the MVD, a distance index (e.g., mmvd_distance_idx) indicating the motion magnitude, and a direction index (mmvd_direction_idx) indicating the motion direction.

[0172] In MMVD mode, two candidates located in the first and second entries of the merge candidate list (i.e., the first candidate or the second candidate) can be used, and either of these two candidates (i.e., the first candidate or the second candidate) can be used as the base MV. For example, a merge candidate flag (e.g., mmvd_cand_flag) can be signaled to indicate either of the two candidates (i.e., the first candidate or the second candidate) in the merge candidate list.

[0173] Furthermore, the distance index (e.g., mmvd_distance_idx) indicates information about the magnitude of the motion and can indicate a predetermined offset from the starting point. This offset may be added to the horizontal or vertical component of the starting motion vector. The relationship between the distance index and the predetermined offset can be shown in the following table.

[0174] [Table 4]

[0175] Referring to Table 4 above, the MVD distance (e.g., MmvdDistance) is determined by the value of the distance index (e.g., mmvd_distance_idx), and the MVD distance (e.g., MmvdDistance) can be derived using integer sample precision or fractional sample precision based on the value of tile_group_fpel_mmvd_enabled_flag. For example, if tile_group_fpel_mmvd_enabled_flag is 1, it indicates that the MVD distance is derived using integer sample precision in the current tile group (or picture header), and if tile_group_fpel_mmvd_enabled_flag is 0, it indicates that the MVD distance is derived using fractional sample precision in the tile group (or picture header). In Table 1, the information (flags) for tile groups can be replaced with the information for picture headers. For example, tile_group_fpel_mmvd_enabled_flag can be replaced with ph_fpel_mmvd_enabled_flag (or ph_mmvd_fullpel_only_flag).

[0176] Furthermore, the direction index (e.g., mmvd_direction_idx) indicates the direction of the MVD relative to the starting point, and shows four directions as shown in Table 5 below. Here, the direction of the MVD can also indicate the sign of the MVD. The relationship between the direction index and the MVD sign is shown in the table below.

[0177] [Table 5]

[0178] Referring to Table 5 above, the sign of the MVD (e.g., MmvdSign) is determined by the value of the direction index (e.g., mmvd_direction_idx), and the sign of the MVD (e.g., MmvdSign) is derived for L0 and L1 reference pictures.

[0179] Based on the distance index (e.g., mmvd_distance_idx) and direction index (e.g., mmvd_direction_idx) as described above, the MVD offset can be calculated using the following formula.

[0180] <Formula 2>

number

[0181] <Formula 3>

number

[0182] In equations 2 and 3, the MMVD distance (MmvdDistance[x0][y0]) and MMVD sign (MmvdSign[x0][y0][0], MmvdSign[x0][y0][1]) are derived based on Tables 4 and / or 5. In summary, in MMVD mode, a merge candidate indicated by a merge candidate flag (e.g., mmvd_cand_flag) is selected from the merge candidate children of the merge candidate list derived based on the surrounding blocks, and the selected merge candidate can be used as a base candidate (e.g., MVP). Then, the MVD derived using the distance index (e.g., mmvd_distance_idx) and direction index (e.g., mmvd_direction_idx) based on the base candidate can be added to derive the motion information (i.e., motion vector) of the current block.

[0183] Based on the motion information derived by the prediction mode, a predicted block can be derived for the current block. The predicted block includes the predicted sample (predicted sample array) of the current block. If the motion vector of the current block points to fractional sample units, an interpolation procedure can be performed, thereby deriving the predicted sample of the current block based on the reference sample in fractional sample units within the reference picture. When biprediction is applied, the predicted sample derived by a weighted sum or weighted average (depending on the phase) of the predicted sample derived based on the L0 prediction (i.e., prediction using the reference picture in the reference picture list L0 and MVL0) and the predicted sample derived based on the L1 prediction (i.e., prediction using the reference picture in the reference picture list L1 and MVL1) can be used as the predicted sample of the current block. When biprediction is applied, if the reference picture used for the L0 prediction and the reference picture used for the L1 prediction are located in different time directions relative to the current picture (i.e., it is biprediction but corresponds to bidirectional prediction), this may be called true biprediction.

[0184] As mentioned above, reconstructed samples and pictures are generated based on the derived predicted samples, and then procedures such as in-loop filtering can be performed.

[0185] As mentioned above, according to this document, when dual prediction is applied to the current block, the predicted samples can be derived based on a weighted average. Traditionally, the dual prediction signal (i.e., the dual prediction sample) was derived by a simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, the dual prediction sample was derived as the average of the L0 prediction sample based on the L0 reference picture and MVL0 and the L1 prediction sample based on the L1 reference picture and MVL1. However, according to this document, when dual prediction is applied, the dual prediction signal (dual prediction sample) can be derived by a weighted average of the L0 prediction signal and the L1 prediction signal, as follows.

[0186] In the embodiments related to MMVD described above, a method can be proposed that considers long-term reference pictures in the MVD derivation process of MMVD, thereby enabling the maintenance and increase of compression efficiency in various applications. Furthermore, the method proposed in the embodiments of this document can be applied not only to the MMVD technology used in MERGE, but also to SMVD, a symmetric MVD technology used in intermode (MVP mode).

[0187] Figure 10 illustrates the method for deriving motion vectors in interpretation.

[0188] In one embodiment of this document, an MV derivation method is used that considers a long-term reference picture during the motion vector scaling (MV scaling) process of a temporal motion candidate (temporal merge candidate, or temporal mvp candidate). The temporal motion candidate can correspond to an mvCol (mvLXCol). The temporal motion candidate may also be called "TMVP".

[0189] The following table explains the definition of a long-term reference picture.

[0190] [Table 6]

[0191] Referring to Table 6 above, if LongTermRefPic(aPic, aPb, refIdx, LX) is 1 (true), the corresponding reference picture is marked as used for long-term reference. For example, a reference picture that is not marked as used for long-term reference may be a reference picture marked as used for short-term reference. In other examples, a reference picture that is neither marked as used for long-term reference nor as unused may be a reference picture marked as used for short-term reference. Hereinafter, a reference picture marked as used for long-term reference may be referred to as a long-term reference picture, and a reference picture marked as used for short-term reference may be referred to as a short-term reference picture.

[0192] The following table explains the derivation of TMVP(mvLXCol).

[0193] [Table 7]

[0194] Referring to Figure 10 and Table 7, the time motion vector (mvLXCol) is not used if the reference picture type pointed to by the current picture (for example, whether it is a long-term reference picture (LTRP) or a short-term reference picture (STRP)) and the type of the collocated reference picture pointed to by the collocated picture are not the same. That is, colMV is derived if all are long-term reference pictures or all are short-term reference pictures, and not if they are of other types. Also, the collocated motion vector can be used directly without scaling if all are long-term reference pictures, or if the POC difference between the current picture and its reference picture is the same as the POC difference between the collocated picture and its reference picture. If they are short-term reference pictures and the POC differences are different, the scaled motion vector of the collocated block is used.

[0195] In the embodiments of this document, the MMVD used in MERGE / SKIP mode signals the base motion vector index (base MV index), distance index, and direction index to a single coding block as information for deriving MVD information. When performing unidirectional prediction, the MVD is derived from motion information, and when performing bidirectional prediction, symmetric MVD information is generated using mirroring and scaling methods.

[0196] When performing bidirectional prediction, MVD information for L0 or L1 is scaled to generate an MVD for L1 or L0, but if long-term reference pictures are referenced, modifications are required in the MVD derivation process.

[0197] Figure 11 shows the MVD derivation process of MMVD according to one embodiment of this document. The method shown in Figure 11 can be for blocks to which bidirectional prediction is applied.

[0198] Referring to Figure 11, if the distance to the L0 reference picture and the distance to the L1 reference picture are the same, the derived MmvdOffset can be used directly as the MVD. If the POC difference (the POC difference between the L0 reference picture and the current picture and the POC difference between the L1 reference picture and the current picture) is different, the MVD can be derived by scaling or simple mirroring (i.e., -1*MmvdOffset) according to the POC difference and whether it is a long-term or short-term reference picture.

[0199] As an example, the method of deriving a symmetric MVD using MMVD for blocks to which bidirectional prediction is applied is not suitable for blocks that use long-term reference pictures, and in particular, when the reference picture types in each direction are different, it is difficult to expect performance improvement when using MMVD. Therefore, the following figures and embodiments show an example in which MMVD is not applied when the reference picture types of L0 and L1 are different.

[0200] Figure 12 shows the MVD derivation process for MMVD according to another embodiment of this document. The method shown in Figure 12 may be for blocks to which bidirectional prediction is applied.

[0201] Referring to Figure 12, different MVD derivation methods are applied depending on whether the reference picture referenced by the current picture (or current slice, current block) is an LTRP (Long-Term Reference Picture) or a STRP (Short-Term Reference Picture). In one example, when the method of the embodiment shown in Figure 12 is applied, a portion of the standard documentation according to this embodiment is described as follows:

[0202] [Table 8-1]

[0203] [Table 8-2]

[0204] Figure 13 shows the MVD derivation process for MMVD according to another embodiment of this document. The method shown in Figure 13 may be for blocks to which bidirectional prediction is applied.

[0205] Referring to Figure 13 above, different MVD derivation methods are applied depending on whether the reference picture referenced by the current picture (or current slice, current block) is an LTRP (Long-Term Reference Picture) or a STRP (Short-Term Reference Picture). In one example, when the method of the embodiment shown in Figure 13 is applied, a portion of the standard documentation according to this embodiment is described as follows:

[0206] [Table 9-1]

[0207] [Table 9-2]

[0208] In summary, the MMVD derivation process for MMVD is explained, which does not derive MVD when the reference picture types in each direction are different.

[0209] In one embodiment of this document, the MVD is not derived in all cases where a long-term reference picture is referenced. That is, if at least one of the L0 and L1 reference pictures is a long-term reference picture, the MVD is set to 0, and the MVD can only be derived when there is a short-term reference picture. This will be specifically explained in the following drawings and tables.

[0210] Figure 14 shows the MVD derivation process of MMVD according to one embodiment of this document. The method shown in Figure 14 may be for blocks to which bidirectional prediction is applied.

[0211] Referring to Figure 14 above, the MVD for MMVD can be derived when the current picture (or current slice, current block) references only short-term reference pictures based on the highest priority condition (RefPicL0!=LTRP&&RefPicL1!=STRP). In one example, when the method of the embodiment shown in Figure 14 is applied, a portion of the standard documentation according to this embodiment is described as follows:

[0212] [Table 10-1]

[0213] [Table 10-2]

[0214] In one embodiment of this document, if the reference picture types in each direction are different, the MVD is derived if there is a short-term reference picture, and if there is a long-term reference picture, the MVD is derived to 0. This will be specifically illustrated in the following drawings and tables.

[0215] Figure 15 shows the MVD derivation process of an MMVD according to one embodiment of this document. The method shown in Figure 15 may be for blocks to which bidirectional prediction is applied.

[0216] Referring to Figure 15 above, when the reference picture types in each direction are different, MmvdOffset is applied when referencing a reference picture that is close to the current picture (short-term reference picture), and MVD has a value of 0 when referencing a reference picture that is far from the current picture (long-term reference picture). Here, a picture close to the current picture can be considered to have a short-term reference picture, but if the close picture is a long-term reference picture, mmvdOffset can be applied to the motion vector of the list pointing to the short-term reference picture.

[0217] [Table 11]

[0218] For example, the four paragraphs included in Table 11 above can sequentially replace the bottom block (content) of the flowchart included in Figure 15 above.

[0219] In one example, when the method of the embodiment shown in Figure 15 is applied, a portion of the standard documentation according to this embodiment is described as follows:

[0220] [Table 12-1]

[0221] [Table 12-2]

[0222] The following table shows a comparison table of the embodiments included in this document.

[0223] [Table 13]

[0224] Referring to Table 13, a comparison is shown between the methods of applying an offset considering the reference picture type for MVD derivation of MMVD as described in the embodiments of Figures 11 to 15. In Table 13, Embodiment A relates to an existing MMVD, Embodiment B shows the embodiments of Figures 11 to 13, Embodiment C shows the embodiment of Figure 14, and Embodiment D shows the embodiment of Figure 15.

[0225] Specifically, the embodiments shown in Figures 11, 12, and 13 describe a method for deriving the MVD only when the reference picture types in both directions are the same, while the embodiment shown in Figure 14 describes a method for deriving the MVD only when both directions are short-term reference pictures. In the embodiment shown in Figure 14, the MVD is set to 0 if the reference picture is a long-term reference picture for a unidirectional prediction. Furthermore, the embodiment shown in Figure 15 describes a method for deriving the MVD in only one direction when the reference picture types in both directions are different. These differences between embodiments represent various features of the technology described herein, and it will be understood by those ordinary skill in the art to which this specification belongs that the effects that the embodiments described herein aim to achieve can be realized based on these features.

[0226] In the embodiments described herein, a separate process is involved when the reference picture type is a long-term reference picture. When long-term reference pictures are included, POCDiff-based scaling or mirroring does not affect performance improvement, so the MVD in the direction with short-term reference pictures is assigned an MmvdOffset value, and the MVD in the direction with long-term reference pictures is assigned a value of 0. In one example, when this embodiment is applied, a portion of the standard documentation according to this embodiment is described as follows:

[0227] [Table 14-1]

[0228] [Table 14-2]

[0229] In other examples, parts of Table 14 above can be replaced with the following table. Referring to Table 15, the Offset is applied based on the reference picture type, not the POCDiff.

[0230] [Table 15]

[0231] In further examples, parts of Table 14 above can be replaced with the following table. Referring to Table 16, it is always possible to set MmvdOffset to L0 and -MmvdOffset to L1, without considering the reference picture type.

[0232] [Table 16]

[0233] According to one embodiment of this document, similar to the MMVD used in the aforementioned MERGE mode, the SMVD in the inter mode can be performed. When performing bidirectional prediction, whether symmetric MVD derivation is possible is signaled from the encoding device to the decoding device. When the related flag (e.g., sym_mvd_flag) is true (or its value is 1), the second-direction MVD (e.g., MVD L1) is derived by mirroring the first-direction MVD (e.g., MVD L0). In this case, scaling for the first-direction MVD may not be performed.

[0234] The following table shows the syntax for a coding unit according to one embodiment of this document.

[0235] [Table 17]

[0236] [Table 18]

[0237] Referring to Table 17 and Table 18 above, when inter_pred_idc == PRED_BI and the reference pictures of L0 and L1 are available (e.g., RefIdxSymL0 > -1 && RefIdxSymL1 > -1), sym_mvd_flag is signaled.

[0238] The following table shows the decoding procedure for the MMVD reference index by way of an example.

[0239] [Table 19]

[0240] Table 19 describes the procedure for deriving the availability of L0 and L1 reference pictures. Specifically, if there are forward-direction reference pictures among the L0 reference pictures, the reference picture index closest to the current picture is set to RefIdxSymL0, and this value is set to the L0 reference index. Similarly, if there are backward-direction reference pictures among the L1 reference pictures, the reference picture index closest to the current picture is set to RefIdxSymL1, and this value is set to the L1 reference index.

[0241] Table 20 below shows the decoding procedure for MMVD reference indexes using other examples.

[0242] [Table 20]

[0243] Referring to Table 20, if the L0 or L1 reference picture types are different, as in the embodiments described with Figures 11, 12, and 13, i.e., if long-term and short-term reference pictures are used, SMVD should not be used after the reference index derivation for SMVD, if the reference picture types of L0 and L1 are different (see the bottom paragraph of Table 20).

[0244] In one embodiment of this document, SMVD can be applied in intermode, similar to MMVD used in merge mode. When long-term reference pictures are used, as in the embodiment described with Figure 14, the long-term reference pictures can be excluded in the process of deriving the reference index for SMVD, as shown in the table below, to prevent SMVD.

[0245] [Table 21]

[0246] The following table according to another example of this embodiment shows an example of processing such that the SMVD is not applied when using a long-term reference picture after the derivation of the reference picture index for the SMVD.

[0247]

Table 22

[0248] In one embodiment of this document, when the reference picture type of the current picture and the reference picture type of the collocated picture are different in the colMV derivation process of TMVP, the motion vector MV is set to 0, but since it is different from the derivation methods in the cases of MMVD and SMVD, it is made to be unified.

[0249] Even when the reference picture type of the current picture is a long-term reference picture and the reference picture type of the collocated picture is a long-term reference picture, the motion vector uses the collocated motion vector value as it is, but in MMVD and SMVD, in this case, the MV is set to 0. Here, TMVP also sets the MV to 0 without additional derivation.

[0250] Also, even if the reference picture types are different, since there may be a long-term reference picture close to the current picture, considering this, instead of setting the MV to 0, the colMV can be used as the MV without scaling.

[0251] FIG. 16 is a diagram for explaining the SMVD according to one embodiment of this document.

[0252] To derive the SMVD, the method shown in Figure 16 can be used. That is, the SMVD can be derived based on STRP (Short-Term Reference Picture) and / or LTRP (Long-Term Reference Picture). When using a mirrored L0 MVD for the L1 MVD, an inaccurate MVD may be derived if the types of reference pictures are different. This is because the ratio of distances (the distance between reference picture 0 and the current picture and the distance between reference picture 1 and the current picture) becomes larger, and the correlation of the motion vectors in each direction decreases.

[0253] According to one embodiment of this document, the availability of a reference picture is checked, and if the conditions are met, sym_mvd_flag may be purged. If sym_mvd_flag is true, the MVD of L1 (MVDL1) can be derived as mirrored MVDL0 (MVD of L0).

[0254] The following table shows a portion of the coding unit syntax according to this embodiment.

[0255] [Table 23]

[0256] Based on Table 23, the procedure for deriving sym_mvd_flag according to this embodiment can be explained.

[0257] In this embodiment, a reference picture index (RefIdxSymLX with X=0,1) for SMVD can be derived. RefIdxSymL0 can indicate the nearest reference picture (or its index) having a POC smaller than the current picture's POC. RefIdxSymL1 can indicate the nearest reference picture (or its index) having a POC larger than the current picture's POC.

[0258] The following table describes, in standard document format, how to derive a reference picture index for SMVD according to this embodiment.

[0259] [Table 24]

[0260] The following table shows the comparison results between embodiments. The accuracy of the MVD in the SMVD can be improved by considering the reference picture type in the embodiments included in Table 25. In Table 25, the MVD can represent MVD 0 (MVD of L0).

[0261] [Table 25]

[0262] Referring to Table 25, Example P shows an existing method for deriving an SMVD. In Example Q, the SMVD may be limited when a mixed reference picture type (e.g., STRP / LTRP or LTRP / STRP) is used in L0 and L1. In Example R, the SMVD may be limited when a Long-Term Reference Picture (LTRP) is referenced.

[0263] The following table describes, in standard document format, how to derive a reference picture index for SMVD using Example Q in Table 25.

[0264] [Table 26]

[0265] The following table describes, in standard document format, how to derive a reference picture index for SMVD using Example Q in Table 25.

[0266] [Table 27]

[0267] [Table 28]

[0268] Referring to Tables 27 and / or 28, SMVD may be restricted when referencing long-term reference pictures (LTRPs). For example, referring to Table 27, long-term reference pictures may be excluded in the reference picture checking process. This allows other reference pictures (e.g., not long-term reference pictures) to be considered for SMVD. Referring to Table 28, SMVD may not be performed if the closest reference picture to the current picture is a long-term reference picture. For example, even if the reference picture list contains short-term reference pictures, SMVD may not be performed if the closest reference picture to the current picture is a long-term reference picture.

[0269] In one example according to one embodiment of this document, if the POC distance of L0 is greater than or equal to the POC of L1 in the MMVD procedure, the L1 MVD can be derived as a scaled or mirrored L0 MVD. If the POC distance of L0 is less than the POC of L1 in the MMVD procedure, the L0 MVD can be derived as a scaled or mirrored L1 MVD in the MMVD procedure.

[0270] Figure 17 is a flowchart showing a method for deriving MMVD according to one embodiment of this document.

[0271] In one embodiment of this document, the MVD can be derived in the MMVD, taking into account the POC difference and / or the reference picture type. Referring to Figure 17, currPocDiffLX can represent the difference between the POC of the current picture and the POC of the reference picture LX. CurrPocDiffL0 and currPocDiffL1 can be compared with each other, and the type of the reference picture can be checked ("refPicList0 != LTRP" or "refPicList1 != LTRP"). Taking the conditions into account, MmvdOffset (derived using mmvd_cand_flag, mmvd_distance_idx, and / or mmvd_direction_idx) can be assigned as the same value as mMvdLX, a mirrored value, or a scaled value.

[0272] The following table shows a portion of the standard documentation according to this embodiment.

[0273] [Table 29-1]

[0274] [Table 29-2]

[0275] When the current picture references one or more long-term reference pictures (LTRPs), a mirroring procedure that considers the point-of-concept (POC) distance may not be necessary. This is because a mirrored MVD obtained from a reference picture that is much farther away than other MVDs will not be effective in terms of accuracy. A solution to this is described below.

[0276] The following table shows the comparison results between the examples.

[0277] [Table 30]

[0278] Referring to Table 30, Example X shows an existing method for deriving MMVD. In Example Y, the MMVD procedure may be restricted when one or more long-term reference pictures are referenced in the current block. That is, in Example Y, the procedure for comparing POC distances for long-term reference pictures may be omitted. In Example Z, the MMVD derivation procedure may be restricted for all cases. That is, in Example Z, the procedure for comparing POC distances may be omitted for all cases. In Table 30, offset may refer to MmvdOffset.

[0279] Figure 18 is a flowchart illustrating a method for deriving MMVD according to one embodiment of this document. The flowchart in Figure 18 illustrates the method for deriving MMVD according to the aforementioned embodiment Y.

[0280] Referring to Figure 18, the condition for comparing POC differences can be removed when the reference picture type is a long-term reference picture, and the anchor MVD used for the mirroring procedure can be fixed to L0 MVD.

[0281] The following table describes, in standard document format, the method for deriving MMVD according to Example Y in Table 30.

[0282] [Table 31-1]

[0283] [Table 31-2]

[0284] Figure 19 is a flowchart illustrating a method for deriving MMVD according to one embodiment of this document. The flowchart in Figure 19 illustrates the method for deriving MMVD according to the aforementioned embodiment Z.

[0285] Referring to Figure 19, in Example Z, the MMVD derivation procedure can be restricted for all cases. For all cases, the condition for comparing POC differences can be removed, and the anchor MVD used for the mirroring or scaling procedure can be fixed to the L0 MVD.

[0286] The following table describes, in standard document format, the method for deriving MMVD according to Example Z in Table 30.

[0287] [Table 32-1]

[0288] [Table 32-2]

[0289] Furthermore, in one example of this embodiment, the condition for comparing POC differences may be removed in all cases, and only the mirroring approach may be used. The following table describes the method for deriving the MMVD in this example in the format of a standard document.

[0290] [Table 33]

[0291] The following drawings have been prepared to illustrate a specific example of this specification. The names of specific devices and signals / messages / fields shown in the drawings are presented illustratively, and the technical features of this specification are not limited to the specific names used in the following drawings.

[0292] Figures 20 and 21 schematically show an example of a video / image encoding method and related components according to the embodiments of this document. The method disclosed in Figure 20 can be performed by the encoding apparatus disclosed in Figure 2. Specifically, for example, S2000 to S2040 in Figure 20 can be performed by the prediction unit 220 of the encoding apparatus, S2050 can be performed by the residual processing unit 230 of the encoding apparatus, and S2060 can be performed by the entropy encoding unit 240 of the encoding apparatus. The method disclosed in Figure 20 may include embodiments described above.

[0293] Referring to Figure 20, the encoding device derives an interpretation mode for the current block in the current picture (S2000). Here, the interpretation mode can include the merge mode, AMVP mode (mode using motion vector predictor candidates), MMVD, and SMVD as described above.

[0294] The encoding device derives a reference picture for the interprediction mode described above (S2010). In one example, the reference picture may be included in reference picture list 0 (or L0, reference picture list L0) or reference picture list 1 (or L1, reference picture list L1). For example, the encoding device may configure a reference picture list for each slice included in the current picture.

[0295] The encoding device derives motion information for predicting the current block based on the above interprediction mode (S2020). The motion information may include a reference picture index and a motion vector. For example, the encoding device can derive a reference index for SMVD (symmetric motion vector reference index). The reference index for SMVD may point to a reference picture for the application of SMVD. The reference index for SMVD may include reference index L0 (RefIdxSumL0) and reference index L1 (RefIdxSumL1).

[0296] The encoding device can construct a list of motion vector predictor candidates and derive motion vector predictors based on the above list. The encoding device can derive motion vectors based on the symmetric MVD and the above motion vector predictors.

[0297] The encoding device generates predicted samples based on the motion information (S2030). The encoding device can generate the predicted samples based on the motion vectors and reference picture index included in the motion information. For example, the predicted samples can be generated based on the blocks (or samples) within the reference picture pointed to by the reference picture index that are indicated by the motion vectors.

[0298] The encoding device generates prediction-related information, including the above-mentioned interpretation mode (S2040). The above-mentioned prediction-related information may include information about MMVD, information about SMVD, and so on.

[0299] The encoding device derives residual information based on the predicted sample (S2050). Specifically, the encoding device can derive residual samples based on the predicted sample and the original sample. The encoding device can derive residual information based on the residual sample. The transformation and quantization processes described above can be performed to derive the residual information.

[0300] The encoding device encodes image / video information including the prediction-related information and the residual information (S2060). The encoded image / video information can be output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or (digital) storage medium.

[0301] The above image / video information may include a variety of information according to the embodiments of this document. For example, the above image / video information may include information disclosed in at least one of the Tables 1 to 33 mentioned above.

[0302] In one embodiment, the prediction-related information may include interpretation type information indicating whether or not biprediction is applied to the current block in the current picture. For example, based on the interpretation type information, the prediction-related information may include symmetric motion vector difference reference flag information indicating whether or not symmetric motion vector difference reference is applicable. The reference picture may also include a short-term reference picture. Based on the symmetric motion vector difference reference flag information, a symmetric motion vector difference reference index can be derived from a reference index pointing to a short-term reference picture. The motion information may include a motion vector for the current block and the symmetric motion vector difference reference index. Based on the motion vector and the symmetric motion vector difference reference index, the prediction sample can be generated.

[0303] In one embodiment, the symmetric motion vector difference reference index can be derived based on the POC difference between each of the short-term reference pictures and the current picture. Here, in one example, the POC difference between the current picture and a previous reference picture from the current picture may be greater than 0. In another example, the POC difference between the current picture and a subsequent reference picture from the current picture may be less than 0. However, this is merely illustrative.

[0304] In one embodiment, the encoding device can configure a reference picture list L0 (or reference picture list 0) for L0 prediction and a reference picture list L1 (or reference picture list 0) for L1 prediction. For example, the short-term reference picture may include a short-term reference picture L0 included in the reference picture list L0, and a short-term reference picture L1 included in the reference picture list L1. The POC difference may include a first POC difference between the short-term reference picture L0 and the current picture, and a second POC difference between the short-term reference picture L1 and the current picture. For example, the symmetric motion vector difference reference index may include a symmetric motion vector difference reference index L0 and a symmetric motion vector difference reference index L1. Based on the first POC difference, the symmetric motion vector difference reference index L0 can be derived. Based on the second POC difference, the symmetric motion vector difference reference index L1 can be derived.

[0305] In one embodiment, the first POC difference may be the same as the second POC difference.

[0306] In one embodiment, the encoding device can configure a reference picture list L0 for L0 prediction. The short-term reference picture may include a first short-term reference picture L0 and a second short-term reference picture L0 included in the reference picture list L0. For example, the POC difference may include a third POC difference between the first short-term reference picture L0 and the current picture, and a fourth POC difference between the second short-term reference picture L0 and the current picture. For example, the symmetric motion vector difference reference index may include a symmetric motion vector difference reference index L0. Based on a comparison between the third and fourth POC differences, a reference picture index pointing to the first short-term reference picture L0 may be used as the symmetric motion vector difference reference index L0.

[0307] In one embodiment, if the third POC difference is even smaller than the fourth POC difference, the reference picture index pointing to the first short-term reference picture L0 can be used as the symmetric motion vector difference reference index L0.

[0308] In one embodiment, the image information may include information about Motion Vector Differences (MVD). The motion information may include Motion Vectors (MV). Based on the information about the MVD, MVDL0 for L0 prediction can be derived. The MV can be derived based on MVDL0 and MVDL1.

[0309] In one embodiment, the size of MVDL1 may be the same as the size of MVDL0. The reference numeral of MVDL1 may be the opposite of the reference numeral of MVDL0.

[0310] Figures 22 and 23 schematically illustrate an example of an image / video decoding method and related components according to the embodiments of this document. The method disclosed in Figure 22 can be performed by the decoding device disclosed in Figure 3. Specifically, for example, S2200 in Figure 22 can be performed by the entropy decoding unit 310 of the decoding device, S2210 to S2230 can be performed by the prediction unit 330 of the decoding device, S2240 can be performed by the residual processing unit 320 of the decoding device, and S2250 can be performed by the addition unit 340 of the decoding device. The method disclosed in Figure 22 may include embodiments described above.

[0311] Referring to Figure 22, the decoding device receives / acquires image / video information (S2200). The decoding device can receive / acquire the above image / video information via a bitstream. The above image / video information may include prediction-related information (including prediction mode information) and residual information. The above prediction-related information may include information related to MMVD, information related to SMVD, etc. Furthermore, the above image / video information may include various types of information according to the embodiments of this document. For example, the above image / video information may include the information described with Figures 1 to 19 and / or the information disclosed in at least one of Tables 1 to 33 mentioned above.

[0312] The decoding device derives an inter-prediction mode for the current block based on the above prediction-related information (S2210). Here, the inter-prediction mode can include the merge mode, AMVP mode (a mode using motion vector predictor candidates), MMVD, and SMVD as described above.

[0313] The decoding device derives motion information for predicting the current block based on the above interprediction mode (S2220). The motion information may include a reference picture index and a motion vector. For example, the decoding device may derive a reference index for SMVD. The reference index for SMVD may point to a reference picture for the application of SMVD. The reference index for SMVD may include reference index L0 (RefIdxSumL0) and reference index L1 (RefIdxSumL1).

[0314] The decoding device can construct a list of motion vector predictor candidates and derive motion vector predictors based on the above list. The decoding device can derive motion vectors based on the symmetric MVD and the above motion vector predictors.

[0315] The decoding device generates a predicted sample based on the motion information (S2230). The decoding device can generate the predicted sample based on the motion vector and the reference picture index included in the motion information. For example, the predicted sample can be generated based on the block (or sample) indicated by the motion vector among the blocks (or samples) in the reference picture pointed to by the reference picture index.

[0316] The decoding device generates residual samples based on the residual information (S2240). Specifically, the decoding device can derive quantized transformation coefficients based on the residual information. The quantized transformation coefficients may take the form of a one-dimensional vector based on the coefficient scan order. The decoding device can derive transformation coefficients based on an inverse quantization procedure for the quantized transformation coefficients. The decoding device can derive residual samples based on an inverse transformation procedure for the transformation coefficients.

[0317] The decoding device generates a reconstructed sample of the current picture based on the predicted sample and the residual sample (S2250). The decoding device may also perform further filtering steps to generate a (corrected) reconstructed sample.

[0318] In one embodiment, the prediction-related information may include interpretation type information indicating whether or not biprediction is applied to the current block in the current picture. For example, based on the interpretation type information, the prediction-related information may include symmetric motion vector difference reference flag information indicating whether or not symmetric motion vector difference reference is applicable. The reference picture may also include a short-term reference picture. Based on the symmetric motion vector difference reference flag information, a symmetric motion vector difference reference index can be derived from a reference index pointing to a short-term reference picture. The motion information may include a motion vector for the current block and the symmetric motion vector difference reference index. Based on the motion vector and the symmetric motion vector difference reference index, the prediction sample can be generated.

[0319] In one embodiment, the symmetric motion vector difference reference index can be derived based on the POC difference between each of the short-term reference pictures and the current picture. Here, in one example, the POC difference between the current picture and a previous reference picture from the current picture may be greater than 0. In another example, the POC difference between the current picture and a subsequent reference picture from the current picture may be less than 0. However, this is merely illustrative.

[0320] In one embodiment, the encoding device can configure a reference picture list L0 (or reference picture list 0) for L0 prediction and a reference picture list L1 (or reference picture list 0) for L1 prediction. For example, the short-term reference picture may include a short-term reference picture L0 included in the reference picture list L0, and a short-term reference picture L1 included in the reference picture list L1. The POC difference may include a first POC difference between the short-term reference picture L0 and the current picture, and a second POC difference between the short-term reference picture L1 and the current picture. For example, the symmetric motion vector difference reference index may include a symmetric motion vector difference reference index L0 and a symmetric motion vector difference reference index L1. Based on the first POC difference, the symmetric motion vector difference reference index L0 can be derived. Based on the second POC difference, the symmetric motion vector difference reference index L1 can be derived.

[0321] In one embodiment, the first POC difference may be the same as the second POC difference.

[0322] In one embodiment, the encoding device can configure a reference picture list L0 for L0 prediction. The short-term reference picture may include a first short-term reference picture L0 and a second short-term reference picture L0 included in the reference picture list L0. For example, the POC difference may include a third POC difference between the first short-term reference picture L0 and the current picture, and a fourth POC difference between the second short-term reference picture L0 and the current picture. For example, the symmetric motion vector difference reference index may include a symmetric motion vector difference reference index L0. Based on a comparison between the third and fourth POC differences, a reference picture index pointing to the first short-term reference picture L0 may be used as the symmetric motion vector difference reference index L0.

[0323] In one embodiment, if the third POC difference is even smaller than the fourth POC difference, the reference picture index pointing to the first short-term reference picture L0 can be used as the symmetric motion vector difference reference index L0.

[0324] In one embodiment, the image information may include information about Motion Vector Differences (MVD). The motion information may include Motion Vectors (MV). Based on the information about the MVD, MVDL0 for L0 prediction can be derived. The MV can be derived based on MVDL0 and MVDL1.

[0325] In one embodiment, the size of MVDL1 may be the same as the size of MVDL0. The reference numeral of MVDL1 may be the opposite of the reference numeral of MVDL0.

[0326] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and different steps may be included, or one or more steps in the flowchart may be omitted without affecting the scope of the embodiments described herein.

[0327] The methods according to the embodiments of this document described above can be implemented in software form, and the encoding and / or decoding devices relating to this document may be included in, for example, image processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.

[0328] In this document, when embodiments are embodied in software, the methods described above can be embodied in modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (Application-Specific Integrated Circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (Read-Only Memory), RAM (Random Access Memory), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described in this document may be embodied on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be embodied on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for the embodiment may be stored in a digital storage medium.

[0329] Furthermore, the decoding and encoding devices to which the embodiments of this document apply may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providers, OTT video (Over The Top video) devices, internet streaming service providers, 3D video devices, VR (Virtual Reality) devices, AR (Augmented (Argumente) Reality) devices, image phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, and may be used to process video signals or data signals. For example, OTT video (Over The Top video) devices may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0330] Furthermore, the processing methods to which the embodiments of this document apply can be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices on which data to be read by a computer is stored. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), Universal Serial Bus (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wireless network.

[0331] Furthermore, the embodiments described herein can be embodied in a computer program product comprising program code, the program code of which can be executed on a computer according to the embodiments described herein. The program code of which can be stored on a computer-readable carrier.

[0332] Figure 24 shows an example of a content streaming system to which the embodiments disclosed in this document may be applied.

[0333] Referring to Figure 24, the content streaming system to which the embodiments described in this document apply can broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0334] The above-mentioned encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the above-mentioned streaming server. As an alternative example, if the multimedia input device such as a smartphone, camera, or camcorder generates the bitstream directly, the above-mentioned encoding server may be omitted.

[0335] The bitstream described above can be generated by an encoding method or a bitstream generation method to which an embodiment of this document applies, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0336] The streaming server transmits multimedia data to user devices based on user requests via the web server, and the web server acts as an intermediary to inform the user about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.

[0337] The above-mentioned streaming server can receive content from a media storage device (storage) and / or an encoding server. For example, when receiving content from the above-mentioned encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the above-mentioned streaming server can store the above-mentioned bitstream for a certain period of time.

[0338] Examples of user devices mentioned above include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation systems, slate PCs, tablet PCs, ultrabooks (ULTRABOOK®), wearable devices (such as smartwatches, smart glasses, and HMDs (Head Mounted Displays)), digital TVs, desktop computers, and digital signatures (Signiji).

[0339] Each server within the above content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0340] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined to embody an apparatus, and the technical features of the apparatus claims herein can be combined to embody a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody a method.

Claims

1. In an image decoding method performed by a decoding device, The steps include receiving image information from a bitstream, including residual information and prediction-related information including information on MVD (motion vector difference), Based on the aforementioned prediction-related information, the steps include: deriving an interpretation mode for the current block in the current picture; The steps include: deriving motion information for the current block based on the inter prediction mode; The steps include generating a predictive sample based on the aforementioned motion information, The steps include generating residual samples based on the residual information, The step of generating a restored sample of the current picture based on the predicted sample and the residual sample, The aforementioned prediction-related information includes interprediction type information indicating whether dual prediction is applied to the current block, Based on the aforementioned interpretation type information, the prediction-related information further includes symmetric motion vector difference flag information indicating whether a symmetric motion vector difference reference is applied, Based on the aforementioned symmetric motion vector difference flag information, a symmetric motion vector difference reference index is derived from the reference index indicating the short-term reference picture. Based on the Picture Order Count (POC) difference between each of the aforementioned short-term reference pictures and the current picture, the symmetric motion vector difference reference index is derived. Based on the information relating to the MVD, MVDL0 for L0 prediction is derived. A method for deriving MVDL1 for L1 prediction based on the aforementioned MVDL0.

2. In an image encoding method performed by an encoding device, The steps include: deriving the interpretation mode for the current block within the current picture, The steps include deriving a reference picture for the aforementioned interpretation mode, Based on the aforementioned inter-prediction mode, the step of deriving motion information for predicting the current block, The steps include generating a predictive sample based on the aforementioned motion information, The steps include generating prediction-related information including information about the interpretation mode and information about MVD (motion vector difference), The steps include generating residual information based on the aforementioned prediction sample, The step of encoding image information including the residual information and the prediction-related information including the information relating to the interpretation mode and the information relating to the MVD, The aforementioned prediction-related information includes interprediction type information indicating whether dual prediction is applied to the current block, Based on the aforementioned interpretation type information, the prediction-related information further includes symmetric motion vector difference flag information indicating whether a symmetric motion vector difference reference is applied, The aforementioned reference picture includes a short-term reference picture. Based on the aforementioned symmetric motion vector difference flag information, a symmetric motion vector difference reference index is derived from the reference index indicating the short-term reference picture. Based on the Picture Order Count (POC) difference between each of the aforementioned short-term reference pictures and the current picture, the symmetric motion vector difference reference index is derived. Based on the information relating to the MVD, MVDL0 for L0 prediction is derived. A method for deriving MVDL1 for L1 prediction based on the aforementioned MVDL0.

3. Regarding methods for transmitting image-related data, A step of obtaining the bitstream of the aforementioned data, wherein the bitstream is The steps include: deriving the interpretation mode for the current block within the current picture, The steps include deriving a reference picture for the aforementioned interpretation mode, Based on the aforementioned inter-prediction mode, the step of deriving motion information for predicting the current block, The steps include generating a predictive sample based on the aforementioned motion information, The steps include generating prediction-related information including information about the interpretation mode and information about MVD (motion vector difference), The steps include generating residual information based on the aforementioned prediction sample, A step of encoding image information including the residual information, the information relating to the interpretation mode, and the prediction-related information including the information relating to the MVD, which is generated based on the steps of: The step of transmitting the data, which includes the bitstream, The aforementioned prediction-related information includes interprediction type information indicating whether dual prediction is applied to the current block, Based on the aforementioned interpretation type information, the prediction-related information further includes symmetric motion vector difference flag information indicating whether a symmetric motion vector difference reference is applied, The aforementioned reference picture includes a short-term reference picture. Based on the aforementioned symmetric motion vector difference flag information, a symmetric motion vector difference reference index is derived from the reference index indicating the short-term reference picture. Based on the Picture Order Count (POC) difference between each of the aforementioned short-term reference pictures and the current picture, the symmetric motion vector difference reference index is derived. Based on the information relating to the MVD, MVDL0 for L0 prediction is derived. A method for deriving MVDL1 for L1 prediction based on the aforementioned MVDL0.