Video coding method and apparatus utilizing motion vector difference

By signaling motion vector differences using symmetrical motion vector differences and deriving a reference index, the method improves video compression efficiency and reduces complexity in inter prediction, addressing the need for efficient compression of high-resolution and immersive media.

JP7866124B2Active Publication Date: 2026-05-26LG ELECTRONICS INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-07-17
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, as well as immersive media, has led to a need for more efficient video compression techniques to reduce transmission and storage costs, particularly in inter prediction using motion vector differences.

Method used

A method and apparatus for signaling information regarding motion vector differences, specifically using a symmetrical motion vector difference (SMVD) and deriving an SMVD reference index, to improve video coding efficiency.

Benefits of technology

This approach enhances video compression efficiency by efficiently signaling motion vector differences and reducing coding system complexity, particularly when dual prediction is applied to a block.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007866124000050
    Figure 0007866124000050
  • Figure 0007866124000051
    Figure 0007866124000051
  • Figure 0007866124000052
    Figure 0007866124000052
Patent Text Reader

Abstract

To provide a method and apparatus for increasing image / video coding efficiency.SOLUTION: According to embodiments of the present document, a prediction procedure can be performed for image / video coding, and the prediction procedure can comprise merge mode motion vector differences (MMVD) and symmetric motion vector differences (SMVD) according to inter prediction. The inter prediction can be performed on the basis of reference pictures of a current picture, and types of the reference pictures (e.g., a long-term reference picture, a short-term reference picture, etc.) can be taken into account for the inter prediction. Accordingly, performance and coding efficiency in the prediction procedure can be increased.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to a video coding method and apparatus using motion vector differences.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or 8K and above UHD (Ultra High Definition) images / videos has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcasting of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] Therefore, a highly efficient image / video compression technique is required to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.

[0005] In particular, inter prediction in video / video coding can utilize motion vector differences. In relation to the above procedure, there has been discussion on deriving motion vector differences based on a reference picture type (e.g., short-term or long-term reference picture).

Summary of the Invention

Means for Solving the Problems

[0006] According to one embodiment of this document, a method and apparatus for improving video coding efficiency are provided.

[0007] According to one embodiment of this document, a method and apparatus for performing efficient inter-prediction in a video coding system are provided.

[0008] According to one embodiment of this document, a method and apparatus for signaling information regarding motion vector differences in interpretation are provided.

[0009] According to one embodiment of this document, a method and apparatus for signaling information regarding the L0 motion vector difference and the L1 motion vector difference when dual prediction is applied to a block is provided.

[0010] The embodiments of this document provide a method and apparatus for signaling the SMVD flag.

[0011] According to one embodiment of this document, a specific reference picture type is used to derive a symmetrical motion vector difference.

[0012] According to one embodiment of this document, the procedure for deriving an SMVD reference index is performed using a shot term reference picture (a picture marked as being used for shot term references).

[0013] According to one embodiment of this document, a video / image decoding method performed by a decoding device is provided.

[0014] According to one embodiment of this document, a decoding device for performing video / image decoding is provided.

[0015] According to one embodiment of this document, a video / image encoding method performed by an encoding device is provided.

[0016] According to one embodiment of this document, an encoding device for performing video / image encoding is provided.

[0017] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded video / image information generated by a video / image encoding method disclosed in at least one of the embodiments of this document.

[0018] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded information or encoded video / image information, causing a decoding device to perform a video / image decoding method disclosed in at least one of the embodiments of this document. [Effects of the Invention]

[0019] According to this document, it is possible to improve the overall video compression efficiency.

[0020] According to this document, information regarding motion vector differences can be efficiently signaled.

[0021] According to this document, when biprediction is applied to the current block, the L1 motion vector difference can be efficiently derived.

[0022] According to this document, the information used to derive the L1 motion vector difference is signaled based on the type of reference picture, thus reducing the complexity of the coding system.

[0023] According to the examples in this document, efficient interpretation can be performed by using a specific reference picture type to derive a reference picture index for SMVD.

[0024] The effects that can be obtained through a specific example of this document are not limited to the effects listed above. For example, there can be various technical effects that can be understood or induced by a person having ordinary skill in the related art from this document. Accordingly, the specific effects of this document can include various effects that can be understood or induced from the technical features of this document, rather than being limited to those explicitly described in this document. Brief Description of the Drawings

[0025] [Figure 1] An example of a video / video coding system applicable to an embodiment of this document is schematically shown. [Figure 2] This is a diagram schematically explaining the configuration of a video / video encoding device applicable to an embodiment of this document. [Figure 3] This is a diagram schematically explaining the configuration of a video / video decoding device applicable to an embodiment of this document. [Figure 4] An example of an inter-prediction-based video / video encoding method is shown. [Figure 5] An example of an inter-prediction-based video / video decoding method is shown. [Figure 6] An example of an inter-prediction procedure is illustratively shown. [Figure 7] This is a diagram for explaining SMVD. [Figure 8] This is a diagram for explaining a method of deriving a motion vector in inter-prediction. [Figure 9] An MVD derivation method of MMVD according to an embodiment of this document is shown. [Figure 10] An MVD derivation method of MMVD according to an embodiment of this document is shown. [Figure 11] An MVD derivation method of MMVD according to an embodiment of this document is shown. [Figure 12] An MVD derivation method of MMVD according to an embodiment of this document is shown. [Figure 13] This document presents an embodiment of a method for inducing MVD from MMVD. [Figure 14] This is a diagram illustrating SMVD according to one embodiment of this document. [Figure 15] This flowchart shows a method for deriving MMVD according to one embodiment of this document. [Figure 16] This flowchart shows a method for deriving MMVD according to one embodiment of this document. [Figure 17] This flowchart shows a method for deriving MMVD according to one embodiment of this document. [Figure 18] An example of a video / image encoding method and related components according to the embodiments described herein is outlined. [Figure 19] An example of a video / image encoding method and related components according to the embodiments described herein is outlined. [Figure 20] An example of an image / video decoding method and related components according to the embodiments of this document is schematically shown. [Figure 21] An example of an image / video decoding method and related components according to the embodiments of this document is schematically shown. [Figure 22] Examples of content streaming systems to which the embodiments disclosed in this document can be applied are shown below. [Modes for carrying out the invention]

[0026] The disclosures in this document can be modified in various ways and may have various embodiments, but specific embodiments are illustrated in the drawings and described in detail. However, this does not mean that the disclosure is limited to any particular embodiment. The terms used in this document are used solely to describe specific embodiments and are not intended to limit the technical ideas of the embodiments described herein. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this document, terms such as “includes” or “has” are intended to indicate the existence of features, figures, stages, operations, components, parts, or combinations thereof described in the document, and should be understood not to preemptively exclude the possibility of the existence or addition of one or more different features, figures, stages, operations, components, parts, or combinations thereof.

[0027] On the other hand, each configuration shown in the drawings described in this document is shown independently for the convenience of explaining its distinct characteristic functions, and does not mean that each configuration is embodied in separate hardware or separate software. For example, two or more configurations may be combined to form a single configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included within the scope of disclosure in this document.

[0028] The embodiments described in this document will be explained below with reference to the attached diagrams. The same reference numerals may be used for the same components in the diagrams, and redundant explanations for the same components may be omitted.

[0029] Figure 1 schematically shows an example of a video / image coding system to which the embodiments described in this document can be applied.

[0030] As shown in Figure 1, a video / image coding system may comprise a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.

[0031] The source device may comprise a video source, an encoding device, and a transmitter. The receiving device may comprise a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be provided in the encoding device. The receiver may be provided in the decoding device. The renderer may comprise a display unit, which may consist of a separate device or external component.

[0032] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or video / image archives containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.

[0033] An encoding device can encode input video / image data. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.

[0034] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0035] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.

[0036] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0037] This document relates to video / image coding. For example, the methods / examples disclosed in this document can be applied to methods disclosed in the VVC (versatile video coding) standard. Furthermore, the methods / examples disclosed in this document can be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., 267 or H.268, etc.).

[0038] This document presents various examples of video / image coding, and unless otherwise noted, these examples may be combined with each other.

[0039] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a unit representing a single image at a specific time point in time, and "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more coding tree units (CTUs). A single picture can consist of one or more slices or tiles. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.

[0040] On the other hand, a single picture can be divided into two or more subpictures. A subpicture can be a rectangular region of one or more slices within a picture.

[0041] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, and can represent only the luma component pixel / pixel value, or only the chroma component pixel / pixel value.

[0042] A unit can represent a basic unit of image processing. A unit can contain at least one of a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block can contain a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0043] In this document, "A or B" may mean "just A," "just B," or "both A and B." In other words, in this document, "A or B" may be interpreted as "A and / or B." For example, in this document, "A, B or C" may mean "just A," "just B," "just C," or "any combination of A, B and C."

[0044] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "just A", "just B", or "both A and B". For example, "A, B, C" can mean "A, B or C".

[0045] In this document, "at least one of A and B" may mean "just A," "just B," or "both A and B." Furthermore, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."

[0046] Furthermore, in this document, "at least one of A, B and C" may mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0047] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."

[0048] Technical features described individually within a single drawing in this document may be embodied individually or simultaneously.

[0049] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments described in this document can be applied. Hereinafter, the term "encoding device" may include an image encoding device and / or a video encoding device.

[0050] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a reconstructor or a reconstructed block generator. The aforementioned video splitting unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0051] The video splitting unit 210 can split the input video (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary structure. Alternatively, the binary-tree structure may be applied first. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. In this case, based on coding efficiency due to video characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further comprise a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving conversion coefficients and / or a unit for deriving a residual signal from conversion coefficients.

[0052] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.

[0053] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input video signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes predicted samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each prediction mode. Prediction information can be encoded by the entropy encoding unit 240 and output in bitstream format.

[0054] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample can be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 can also determine the prediction mode to apply to the current block using the prediction modes applied to adjacent blocks.

[0055] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, BI prediction, etc.). In the case of interprediction, adjacent blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the inter-prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-prediction can be performed based on various prediction modes; for example, in skip mode and merge mode, the inter-prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0056] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content video / video coding such as in games, for example, in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can use at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values ​​within the picture can be signaled based on information about the palette table and palette index.

[0057] The prediction signal generated via the prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can apply a transformation technique to the residual signal to generate transformation coefficients. For example, the transformation technique may include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when relational information between pixels is represented in a graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and based on that. The transformation process may also be applied to pixel blocks of the same size that are square, or to non-square blocks of variable size.

[0058] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information.The video / image information can be encoded via the encoding procedure described above and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network can include broadcast networks and / or communication networks, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.

[0059] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.

[0060] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.

[0061] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream format.

[0062] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.

[0063] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.

[0064] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments described in this document can be applied. Hereinafter, the term "decoding device" may include an image decoding device and / or a video decoding device.

[0065] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-predictor 331 and an inter-predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering device 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may also be configured by a digital storage medium. The aforementioned hardware component may also further include memory 360 as an internal / external component.

[0066] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 3. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed video signal decoded and output via the decoding device 300 can then be played back via a playback device.

[0067] The decoding device 300 can receive the signal output from the encoding device shown in Figure 3 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image restoration and the quantized values ​​of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of adjacent and decoded blocks or symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, performs arithmetic decoding of the bins, and generates symbols corresponding to the values ​​of each syntax element.In this case, the CABAC entropy decoding method can update the context model after determining the context model by utilizing the decoded symbol / bin information for the context model of the next symbol / bin. Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values ​​that have been entropy decoded by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, addition unit 340, filtering unit 350, memory 360, interpretation unit 332, and intraprediction unit 331.

[0068] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0069] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0070] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.

[0071] The prediction unit 330 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its prediction on an intra-block copy (IBC) prediction mode or on a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content video / movie coding such as games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / movie information and signaled.

[0072] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample can be located adjacent to or far from the current block depending on the prediction mode. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to adjacent blocks.

[0073] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, BI prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks that exist in the current picture and temporally adjacent blocks that exist in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.

[0074] The summing unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (which comprises an inter-prediction unit 332 and / or an intra-prediction unit 331). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.

[0075] The summing unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.

[0076] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0077] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0078] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.

[0079] In this specification, the embodiments described for the filtering unit 260, the inter-prediction unit 221, and the intra-prediction unit 222 of the encoding device 200 can also be applied identically or in a corresponding manner to the filtering unit 350, the inter-prediction unit 332, and the intra-prediction unit 331 of the decoding device 300, respectively.

[0080] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived in both the encoding and decoding devices, and the encoding device can improve video coding efficiency by signaling the decoding device information about the residual between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can combine the residual block and the predicted block to generate a restored block containing restored samples, and can generate a restored picture containing the restored block.

[0081] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information (via a bitstream) to a decoding device by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive a residual sample (or residual block) by performing an inverse quantization / inverse transformation procedure based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.

[0082] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If transformation / inverse transformation is omitted, the transformation coefficients may also be called coefficients or residual coefficients, or for consistency of expression, they may still be called transformation coefficients.

[0083] In this document, quantized transformation coefficients and transformation coefficients may be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients, and such information about the transformation coefficients may be signaled via residual coding syntax. Based on the residual information (or information about the transformation coefficients), transformation coefficients may be derived, and scaled transformation coefficients may be derived via inverse transformation (scaling) of the transformation coefficients. Based on inverse transformation (transformation) of the scaled transformation coefficients, residual samples may be derived. This may be applied / expressed similarly in other parts of this document.

[0084] Intra prediction can represent a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived. The adjacent reference samples of the current block may include a total of 2 × nH samples adjacent to the left boundary and bottom-left of the current block of size nW × nH, a total of 2 × nW samples adjacent to the top boundary and top-right of the current block, and one sample adjacent to the top-left of the current block. Alternatively, the adjacent reference samples of the current block may include multiple columns of top-adjacent samples and multiple rows of left-adjacent samples. Furthermore, the adjacent reference samples of the current block may also include a total of nH samples adjacent to the right boundary of the current block, which is of size nW × nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right side of the current block.

[0085] However, some of the adjacent reference samples in the current block may not yet be decoded or available. In this case, the decoder can construct adjacent reference samples to use for prediction by substituting the unavailable samples as available samples, or by constructing adjacent reference samples to use for prediction through interpolation of available samples.

[0086] If neighboring reference samples are derived, (i) predicted samples can be derived based on the average or interpolation of neighboring reference samples in the current block, or (ii) predicted samples can be derived based on neighboring reference samples in the current block that are located in a specific (predicted) direction relative to the predicted sample. Case (i) may be called non-directional mode or non-angular mode, and case (ii) may be called directional mode or angular mode.

[0087] Furthermore, the predicted sample can also be generated by interpolation between a first adjacent sample located in the prediction direction of the current block's intra-prediction mode and a second adjacent sample located in the opposite direction of the prediction direction, using the predicted sample of the current block from the adjacent reference samples as a reference. In the above case, it can be called linear interpolation intra-prediction (LIP). Alternatively, a linear model can be used to generate chroma predicted samples based on chroma samples. In this case, it can be called LM mode.

[0088] Alternatively, temporary predicted samples for the current block can be derived based on filtered adjacent reference samples, and the predicted samples for the current block can be derived by performing a weighted sum on the temporary predicted samples and at least one reference sample derived by the intra-prediction mode from the existing adjacent reference samples, i.e., unfiltered adjacent reference samples. In the above case, it can be called PDPC (Position dependent intra-prediction).

[0089] Furthermore, intra-predictive coding can be performed by selecting the reference sample line with the highest prediction accuracy from among the adjacent multi-reference sample lines in the current block, deriving a predicted sample using the reference sample located in the prediction direction on that line, and then instructing (signaling) the decoding device with the reference sample line used at that time. In the case described above, this can be called multi-reference line intra-prediction or MRL-based intra-prediction.

[0090] Furthermore, the current block can be divided into vertical or horizontal subpartitions, and intra-prediction can be performed based on the same intra-prediction mode. Adjacent reference samples can then be derived and used on a subpartition-by-subpartition basis. In other words, in this case, the intra-prediction mode for the current block is also applied to the subpartition, and by deriving and using adjacent reference samples on a subpartition-by-subpartition basis, intra-prediction performance can be improved in some cases. Such a prediction method can be called ISP (intra sub-partitions) based intra-prediction.

[0091] The intra-prediction method described above can be called an intra-prediction type, distinct from the intra-prediction mode. The intra-prediction type can be referred to by various terms, such as intra-prediction technique or additional intra-prediction mode. For example, the intra-prediction type (or additional intra-prediction mode, etc.) can include at least one of the aforementioned LIP, PDPC, MRL, and ISP. A general intra-prediction method that excludes specific intra-prediction types such as LIP, PDPC, MRL, and ISP can be called a normal intra-prediction type. The normal intra-prediction type can be applied generally when the aforementioned specific intra-prediction types are not applicable, and predictions can be performed based on the aforementioned intra-prediction modes. On the other hand, post-processing filtering can also be performed on the derived prediction samples as needed.

[0092] Specifically, the intra-prediction procedure may include an intra-prediction mode / type determination step, an adjacent reference sample derivation step, and an intra-prediction mode / type-based predictive sample derivation step. Additionally, a post-filtering step may be performed on the derived predictive samples as needed.

[0093] When intraprediction is applied, the intraprediction mode applied to the current block can be determined by utilizing the intraprediction modes of adjacent blocks. For example, the decoding device can select one of the MPM candidates in the MPM (most probable mode) list derived based on the intraprediction modes of the current block's adjacent blocks (e.g., left and / or upper adjacent blocks) and additional candidate modes, based on the received MPM index, or it can select one of the remaining intraprediction modes not included in the MPM candidates (and planar modes), based on the remaining intraprediction mode information. The MPM list may or may not include planar modes as candidates. For example, if the MPM list includes planar modes as candidates, it may have six candidates; if the MPM list does not include planar modes as candidates, it may have five candidates. If the MPM list does not include planar mode as a candidate, a not-planar flag (e.g., intra_luma_not_planar_flag) can be signaled to indicate that the current intra-prediction mode of the block is not planar mode. For example, the MPM flag may be signaled first, and the MPM index and not-planar flag may be signaled if the value of the MPM flag is 1. The MPM index may also be signaled if the value of the not-planar flag is 1. Here, the reason why the MPM list is configured not to include planar mode as a candidate is not because the planar mode is not an MPM, but because planar mode is always considered as an MPM, so the flag (not-planar flag) is signaled first to check whether it is planar mode or not.

[0094] For example, whether the intra-prediction mode applied to the current block is in the MPM candidate (and planar mode) or in the remaining mode can be indicated based on the MPM flag (e.g., intra_luma_mpm_flag). A value of 1 for the MPM flag indicates that the intra-prediction mode for the current block is in the MPM candidate (and planar mode), and a value of 0 for the MPM flag indicates that the intra-prediction mode for the current block is not in the MPM candidate (and planar mode). A value of 0 for the not-planar flag (e.g., intra_luma_not_planar_flag) indicates that the intra-prediction mode for the current block is planar mode, and a value of 1 for the not-planar flag indicates that the intra-prediction mode for the current block is not planar mode. The MPM index can be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntex element, and the remaining intra-prediction mode information can be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntex element. For example, the remaining intra-prediction mode information can be indexed in order of prediction mode number among the overall intra-prediction modes that are not included in the MPM candidate (and planar mode) and point to one of them. The intra-prediction mode is an intra-prediction mode for a luma component (sample). The intra prediction mode information may include at least one of the following: the MPM flag (e.g., intra_luma_mpm_flag), the not planar flag (e.g., intra_luma_not_planar_flag), the MPM index (e.g., mpm_idx or intra_luma_mpm_idx), or the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder).In this document, the MPM list may be referred to by various terms, such as the MPM candidate list or candModeList. When an MIP is applied to the current block, a separate mpm flag (e.g., intra_mip_mpm_flag), an mpm index (e.g., intra_mip_mpm_idx), and remaining intra predictive mode information (e.g., intra_mip_mpm_remainder) for the MIP may be signaled, while the not planar flag is not signaled.

[0095] In other words, when video is generally divided into blocks, the current block to be coded and its neighboring blocks will have similar video characteristics. Therefore, the current block and its neighboring blocks have a high probability of having the same or similar intra-prediction modes. Thus, the encoder can utilize the intra-prediction mode of the neighboring block to encode the intra-prediction mode of the current block.

[0096] For example, an encoder / decoder can configure an MPM (most probable modes) list for the current block. This MPM list can also be referred to as an MPM candidate list. Here, MPM can mean a mode used to improve coding efficiency by considering the similarity between the current block and adjacent blocks during intra predictive mode coding. As mentioned above, the MPM list can be configured to include planar modes or to exclude planar modes. For example, if the MPM list includes planar modes, the number of candidates in the MPM list is 6. If the MPM list does not include planar modes, the number of candidates in the MPM list is 5.

[0097] The encoder / decoder can configure an MPM list containing five or six MPMs.

[0098] To construct an MPM list, three types of modes can be considered: default intra modes, neighbor intra modes, and derived intra modes.

[0099] For the aforementioned adjacent intra-mode, two adjacent blocks, namely the left adjacent block and the upper adjacent block, can be considered.

[0100] As mentioned above, if the MPM list is configured not to include planar mode, planar mode is excluded from the list, and the number of MPM list candidates can be set to 5.

[0101] Furthermore, among the intra-prediction modes, non-directional modes (or non-angle modes) may include average-based DC modes or interpolation-based planar modes of the neighboring reference samples of the current block.

[0102] When interpretation is applied, the prediction unit of the encoding / decoding device can perform interpretation on a block-by-block basis to derive predicted samples. Interpretation can indicate predictions derived in a manner dependent on data elements of pictures other than the current picture (e.g., sample values ​​or motion information). When interpretation is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the index of the reference picture. In this case, in order to reduce the amount of motion information transmitted in interpretation mode, the motion information of the current block can be predicted on a block, subblock, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include motion vectors and the index of the reference picture. The motion information may further include information on the interpretation type (L0 prediction, L1 prediction, BI prediction, etc.). When interpretation is applied, neighboring blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the aforementioned reference block and the reference picture containing the aforementioned temporally adjacent block may be the same or different. The temporally adjacent block may be referred to by names such as collocated reference block, colCU, etc., and the reference picture containing the aforementioned temporally adjacent block may be referred to as collocated picture (colPic). For example, a candidate list of motion information can be constructed based on the adjacent blocks of the current block, and flags or index information can be signaled to indicate which candidate is selected (used) in order to derive the motion vector and / or index of the reference picture of the current block.Interpretation is performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be identical to the motion information of the selected adjacent block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected adjacent block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0103] The motion information may include L0 motion information and / or L1 motion information depending on the interpretation type (L0 prediction, L1 prediction, BI prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a bi(Bi) prediction. Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures earlier in the output order than the current picture, and the reference picture list L1 may include pictures later in the output order than the current picture. The aforementioned earlier picture may be called a forward (reference) picture, and the aforementioned later picture may be called a reverse (reference) picture. The reference picture list L0 may include further reference pictures that are later in the output order than the current picture. In this case, the earlier picture may be indexed first in the reference picture list L0, and the later picture may be indexed afterward. The reference picture list L1 may include further reference pictures that are earlier in the output order than the current picture. In this case, the later picture may be indexed first in the reference picture list L1, and the earlier picture may be indexed afterward. Here, the output order may correspond to the POC (picture order count) order.

[0104] The video / image encoding procedure based on interpretation generally includes, for example, the following:

[0105] Figure 4 shows an example of an interpretation-based video / image encoding method.

[0106] The encoding device performs interpretation for the current block (S400). The encoding device derives the interpretation mode and motion information of the current block and generates a prediction sample for the block. Here, the procedures of determining the interpretation mode, deriving motion information, and generating a prediction sample may be performed simultaneously, or one procedure may be performed before the others. For example, the interpretation unit of the encoding device includes a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block, the motion information derivation unit derives the motion information of the current block, and the prediction sample derivation unit derives a prediction sample for the current block. For example, the interpretation unit of the encoding device searches for blocks similar to the current block within a certain area (search area) of the reference picture by motion estimation and derives a reference block whose difference from the current block is the minimum or below a certain standard. Based on this, a reference picture index pointing to the reference picture in which the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device determines which of the various prediction modes is to be applied to the current block. The encoding device can compare the RD costs for the various prediction modes and determine the optimal prediction mode for the current block.

[0107] For example, when skip mode or merge mode is applied to the current block, the encoding device can configure a merge candidate list (described later) and derive a reference block from among the reference blocks pointed to by the merge candidates included in the merge candidate list whose difference from the current block is the minimum or below a certain standard. In this case, a merge candidate related to the derived reference block is selected, and merge index information pointing to the selected merge candidate is generated and signaled to the decoding device. The movement information of the current block can be derived using the movement information of the selected merge candidate.

[0108] As another example, when the (A)MVP mode is applied to the current block, the encoding device can configure the (A)MVP candidate list described later, and use the motion vector of the selected mvp candidate from among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the motion estimation described above can be used as the motion vector of the current block, and the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block among the mvp candidates can become the selected mvp candidate. The MVD (motion vector difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In that case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index is composed of reference picture index information and is separately signaled to the decoding device.

[0109] The encoding device derives a residual sample based on the predicted sample (S410). The encoding device can derive the residual sample by comparing the original sample of the current block with the predicted sample.

[0110] The encoding device encodes video information including prediction information and residual information (S420). The encoding device outputs the encoded video information in bitstream format. The prediction information is information related to the prediction procedure and includes prediction mode information (e.g., skip flag, merge flag, or mode index) and motion information. The motion information includes candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. The motion information also includes the aforementioned MVD information and / or reference picture index information. The motion information also includes information indicating whether L0 prediction, L1 prediction, or bi prediction is applied. The residual information is information about the residual sample. The residual information includes information about the quantized conversion coefficients for the residual sample.

[0111] The output bitstream may be stored in a (digital) storage medium and transmitted to a decoding device, or it may be transmitted to a decoding device via a network.

[0112] On the other hand, as mentioned above, the encoding device generates a reconstructed picture (including a reconstructed sample and a reconstructed block) based on the reference sample and the residual sample. This is to derive the same prediction results from the encoding device as those performed by the decoding device, thereby improving coding efficiency. Therefore, the encoding device can store the reconstructed picture (or reconstructed sample, reconstructed block) in memory and use it as a reference picture for interpretation. As mentioned above, in-loop filtering procedures and the like can be further applied to the reconstructed picture.

[0113] The video / image decoding procedure based on interpretation broadly includes, for example, the following:

[0114] Figure 5 shows an example of an interpretation-based video / image decoding method.

[0115] As shown in Figure 5, the decoding device performs operations corresponding to those performed by the encoding device. Based on the received prediction information, the decoding device can make predictions in the current block and derive prediction samples.

[0116] Specifically, the decoding device determines the prediction mode for the current block based on the received prediction information (S500). The decoding device can determine which inter-prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0117] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block, or whether the (A)MVP mode is determined. Alternatively, one of several inter-prediction mode candidates can be selected based on the mode index. The inter-prediction mode candidates include skip mode, merge mode and / or (A)MVP mode, or include various inter-prediction modes as described later.

[0118] The decoding device derives motion information of the current block based on the determined inter prediction mode (S510). For example, if a skip mode or merge mode is applied to the current block, the decoding device configures a merge candidate list, which will be described later, and selects one of the merge candidates included in the merge candidate list. The selection is made based on the selection information (merge index) described above. The motion information of the selected merge candidate can be used to derive motion information of the current block. The motion information of the selected merge candidate can be used as motion information of the current block.

[0119] As another example, when the (A)MVP mode is applied to the current block, the decoding device can configure the (A)MVP candidate list described below, and use the motion vector of the selected mvp candidate from among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. The selection is made based on the selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the mvp of the current block and the MVD. Furthermore, the reference picture index of the current block can be derived based on the reference picture index information. Within the reference picture list for the current block, the picture pointed to by the reference picture index can be derived as the reference picture referenced for interpretation of the current block.

[0120] On the other hand, as described later, the motion information of the current block can be derived without constructing a candidate list, in which case the motion information of the current block can be derived according to the procedure disclosed in the prediction mode described later. In this case, the candidate list configuration described above may be omitted.

[0121] The decoding device generates predicted samples for the current block based on the motion information of the current block (S520). In this case, the reference picture can be derived based on the reference picture index of the current block, and the predicted samples for the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, a further procedure of predictive sample filtering may be performed on all or some of the predicted samples for the current block.

[0122] For example, the interpretation unit of the decoding device includes a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block based on the prediction mode information received, the motion information derivation unit derives motion information (such as motion vectors and / or reference picture indices) for the current block based on the motion information received, and the prediction sample derivation unit derives prediction samples for the current block.

[0123] The decoding device generates a residual sample for the current block based on the received residual information (S530). The decoding device generates a restored sample for the current block based on the predicted sample and the residual sample, and generates a restored picture based on this (S540). As mentioned above, further procedures such as in-loop filtering can be applied to the restored picture thereafter.

[0124] Figure 6 illustrates the interpretation prediction procedure.

[0125] Referring to Figure 6, as described above, the interpretation procedure includes an interpretation mode determination step, a motion information derivation step corresponding to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The interpretation procedure is performed in the encoding device and the decoding device, as described above. In this document, the coding device includes the encoding device and / or the decoding device.

[0126] As shown in Figure 6, the coding device determines the interpretation mode for the current block (S600). Various interpretation modes can be used for predicting the current block in the picture. For example, various modes such as merge mode, skip mode, MVP (motion vector prediction) mode, affine mode, subblock merge mode, and MMVD (merge with MVD) mode can be used. Additional modes such as DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, Bi-prediction with CU-level weight (BCW), and Bi-directional optical flow (BDOF) can be used or replaced. The affine mode may also be called affine motion prediction mode. The MVP mode may also be called the AMVP (advanced motion vector prediction) mode. In this document, motion information candidates derived from some modes and / or some modes may be included as one of the motion information related candidates for other modes. For example, an HMVP candidate may be added as a merge candidate for the merge / skip mode, or as an mvp candidate for the MVP mode. When the HMVP candidate is used as a motion information candidate for the merge mode or skip mode, the HMVP candidate may be called an HMVP merge candidate.

[0127] Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be received by the decoding device in the bitstream. The prediction mode information includes index information indicating one of a number of candidate modes. Alternatively, the inter-prediction mode can be indicated via hierarchical signaling of flag information. In this case, the prediction mode information includes one or more flags. For example, a skip flag may be signaled to indicate whether a skip mode is applied, a merge flag may be signaled if the skip mode is not applied to indicate whether the merge mode is applied, and if the merge mode is not applied, an MVP mode may be applied, or flags for additional distinctions may be further signaled. Affine modes may be signaled as independent modes, or as modes dependent on the merge mode or MVP mode, etc. For example, affine modes include affine merge mode and affine MVP mode.

[0128] On the other hand, the current block may be signaled with information indicating whether the aforementioned list0(L0) prediction, list1(L1) prediction, or bi-prediction is used in the current block (current coding unit). This information may also be called motion prediction direction information, inter-prediction direction information, or inter-prediction instruction information, and can be constructed / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether the aforementioned list0(L0) prediction, list1(L1) prediction, or bi-prediction is used in the current block (current coding unit). For the sake of clarity, in this document, the inter-prediction type (L0 prediction, L1 prediction, or BI prediction) pointed to by the inter_pred_idc syntax element may be represented as motion prediction direction. L0 prediction may be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, the following prediction types can be indicated by the value of the inter_pred_idc syntax element:

[0129] [Table 1]

[0130] As mentioned above, a single picture contains one or more slices. A slice can have one of the following slice types: I (intra) slices, P (predictive) slices, and B (bi-predictive) slices. The slice type is indicated based on the slice type information. For blocks in an I slice, only intra prediction is used for prediction, and inter-predictive prediction is not used. Of course, in this case as well, it is possible to code and signal the original sample values ​​without prediction. For blocks in a P slice, intra or inter-predictive prediction is used, and if inter-predictive prediction is used, only uni prediction may be used. On the other hand, for blocks in a B slice, intra or inter-predictive prediction is used, and if inter-predictive prediction is used, up to bi-predictive prediction may be used.

[0131] L0 and L1 contain reference pictures that were encoded / decoded before the current picture. For example, L0 contains reference pictures that are earlier and / or later than the current picture in the POC order, and L1 contains reference pictures that are later and / or earlier than the current picture in the POC order. In this case, L0 is assigned a reference picture index that is lower relative to the reference pictures that are earlier than the current picture in the POC order, and L1 is assigned a reference picture index that is lower relative to the reference pictures that are later than the current picture in the POC order. For B slices, biprediction is applied, and in this case, unidirectional biprediction may be applied, or bidirectional biprediction may be applied. Bidirectional biprediction is also called true biprediction.

[0132] The following table shows the syntax for a coding unit according to one embodiment of this document.

[0133] [Table 2-1]

[0134] [Table 2-2]

[0135] [Table 2-3]

[0136] [Table 2-4]

[0137] [Table 2-5]

[0138] Referring to Table 2 above, general_merge_flag indicates that general merge is available, and when the value of general_merge_flag is 1, regular merge mode, mmvd mode, and merge subblock mode (subblock merge mode) are available. For example, when the value of general_merge_flag is 1, merge data syntax can be parsed from encoded video / image information (or bitstream), and merge data syntax is structured / coded to include information as shown in the following table.

[0139] [Table 3]

[0140] The coding device derives motion information for the current block (S610). The derivation of the motion information can be based on the inter-prediction mode.

[0141] The coding device can perform interpretation using the motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block in the original picture for the current block to search for a highly correlated similar reference block in fractional pixel units within a defined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference in phase-based sample values. For example, block similarity can be calculated based on the sum of absolute differences (SAD) between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information is signaled to the decoding device in various ways based on the interpretation mode.

[0142] The coding device performs inter prediction based on motion information for the current block (S620). The coding device can derive predicted samples for the current block based on the motion information. The current block containing the predicted samples may be called a predicted block.

[0143] When merge mode is applied, the movement information of the currently predicted block is not transmitted directly, but rather the movement information of the surrounding predicted block is used to guide the movement information of the currently predicted block. Therefore, the movement information of the currently predicted block can be instructed by transmitting flag information indicating that merge mode was used and a merge index indicating which surrounding predicted block was used. This merge mode may also be called regular merge mode.

[0144] To perform a merge, the encoder must search for merge candidate blocks to be used to guide the motion information of the currently predicted block. For example, up to five merge candidate blocks may be available, but the embodiments in this document are not limited to this. The maximum number of merge candidate blocks is transmitted in the slice header or tile group header. After finding the merge candidate blocks, the encoder generates a merge candidate list and can select the merge candidate block with the lowest cost from among them as the final merge candidate block.

[0145] The merge candidate list can, for example, utilize five merge candidate blocks. For example, it can utilize four spatial merge candidates and one temporal merge candidate. Hereinafter, the spatial merge candidate, or the spatial MVP candidate described later, may be called SMVP, and the temporal merge candidate, or the temporal MVP candidate described later, may be called TMVP.

[0146] The following describes how to construct the merge candidate list according to this document.

[0147] The coding device (encoder / decoder) searches for spatially surrounding blocks of the current block and inserts the derived spatial merge candidates into the merge candidate list. For example, the spatially surrounding blocks include the block around the lower left corner, the left side, the upper right corner, the top side, and the upper left corner of the current block. However, this is an example, and additional surrounding blocks such as the right side, bottom side, and lower right side can also be used as spatially surrounding blocks. The coding device can search for the spatially surrounding blocks based on priority to detect available blocks and derive the movement information of the detected blocks as the spatial merge candidates.

[0148] The coding device inserts the temporal merge candidates derived by searching the temporal neighboring blocks of the current block into the merge candidate list. The temporal neighboring blocks may be located on a reference picture that is a picture different from the current picture in which the current block is located. The reference picture on which the temporal neighboring blocks are located may be referred to as a collocated picture or a col picture. The temporal neighboring blocks can be searched in the order of the peripheral blocks of the lower right corner and the lower right center block of the co-located block for the current block on the col picture. On the other hand, when motion data compression is applied, specific motion information is stored as representative motion information for each fixed storage unit in the col picture. In this case, it is not necessary to store the motion information for all blocks within the fixed storage unit, and thus a motion data compression effect can be obtained. In this case, the fixed storage unit may be predetermined, for example, in a 16×16 sample unit or an 8×8 sample unit, or the size information regarding the fixed storage unit may be signaled from the encoder to the decoder. When the motion data compression is applied, the motion information of the temporal neighboring blocks can be replaced with the representative motion information of the fixed storage unit in which the temporal neighboring blocks are located. That is, in this case, from the perspective of implementation, instead of the prediction block located at the coordinates of the temporal neighboring blocks, based on the coordinates (the upper left sample position) of the temporal neighboring blocks, after arithmetic right shift by a certain value, the temporal merge candidate is derived based on the motion information of the prediction block covering the position after arithmetic left shift. For example, when the fixed storage unit is a 2n×2n sample unit, if the coordinates of the temporal neighboring blocks are (xTnb, yTnb), the motion information of the prediction block located at the modified position ((xTnb>n)<<n), (yTnb>n)<<n)) is used for the temporal merge candidate.Specifically, for example, if the constant storage unit is 16 × 16 samples, and the coordinates of the temporally surrounding block are (xTnb, yTnb), then the motion information of the predicted block located at the corrected position ((xTnb>4)<<4), (yTnb>4)<<4)) is used for the temporal merge candidate. Alternatively, for example, if the constant storage unit is 8 × 8 samples, and the coordinates of the temporally surrounding block are (xTnb, yTnb), then the motion information of the predicted block located at the corrected position ((xTnb>3)<<3), (yTnb>3)<<3)) is used for the temporal merge candidate.

[0149] The coding device can check whether the current number of merge candidates is less than the maximum number of merge candidates. The maximum number of merge candidates can be predefined or signaled from the encoder to the decoder. For example, the encoder generates information about the maximum number of merge candidates, encodes it, and transmits it to the decoder in bitstream form. Once the maximum number of merge candidates is filled, the subsequent candidate addition process may not be necessary.

[0150] If, as a result of the above check, the number of current merge candidates is less than the number of maximum merge candidates, the coding device inserts additional merge candidates into the merge candidate list.

[0151] If, as a result of the above check, the number of current merge candidates is not less than the number of maximum merge candidates, the coding device terminates the configuration of the merge candidate list. In this case, the encoder can select the optimal merge candidate from among the merge candidates constituting the merge candidate list based on the RD (rate-distortion) cost, and can signal selection information (e.g., merge index) pointing to the selected merge candidate to the decoder. The decoder selects the optimal merge candidate based on the merge candidate list and the selection information.

[0152] As previously mentioned, the motion information of the selected merge candidate can be used as the motion information of the current block, and predicted samples of the current block can be derived based on the motion information of the current block. The encoder can derive the residual samples of the current block based on the predicted samples and signal the decoder with residual information regarding the residual samples. As previously mentioned, the decoder can generate restored samples based on the residual samples derived based on the residual information and the predicted samples, and generate a restored picture based on these.

[0153] When skip mode is applied, the motion information of the current block can be derived in the same manner as when merge mode is applied. However, when skip mode is applied, the residual signal for the block in question is omitted, and therefore, the predicted sample can be immediately used as the restored sample.

[0154] When MVP mode is applied, a motion vector predictor (MVP) candidate list is generated using the motion vectors of the restored spatially surrounding blocks and / or the motion vectors corresponding to the temporally surrounding blocks (or Col blocks). That is, the motion vectors of the restored spatially surrounding blocks and / or the motion vectors corresponding to the temporally surrounding blocks can be used as motion vector predictor candidates. When dual prediction is applied, an MVP candidate list for L0 motion information derivation and an MVP candidate list for L1 motion information derivation can be generated and used separately. The aforementioned prediction information (or information related to prediction) includes selection information (e.g., an MVP flag or MVP index) that indicates the optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the list. Here, the prediction unit can use the selection information to select the motion vector predictor for the current block from among the motion vector predictor candidates included in the motion vector candidate list. The prediction unit of the encoding device can calculate the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode it, and output it in bitstream format. In other words, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, the prediction unit of the decoding device can obtain the motion vector difference included in the prediction information and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction unit of the decoding device can obtain or derive a reference picture index that indicates a reference picture from the prediction information.

[0155] The following describes how to construct the motion vector predictor candidate list according to this document.

[0156] One embodiment first searches for spatial candidate blocks for motion vector prediction and inserts them into the prediction candidate list. Subsequently, the embodiment determines whether the number of spatial candidate blocks is less than 2. For example, if the number of spatial candidate blocks is less than 2, the embodiment searches for temporal candidate blocks and adds them to the prediction candidate list. If temporal candidate blocks are unavailable, it uses a zero motion vector. That is, a zero motion vector can be added to the prediction candidate list. Subsequently, the embodiment finishes constructing the preliminary candidate list. Alternatively, if the number of spatial candidate blocks is not less than 2, the embodiment finishes constructing the preliminary candidate list. Here, the preliminary candidate list refers to the MVP candidate list.

[0157] On the other hand, when MVP mode is applied, the reference picture index is explicitly signaled. In this case, the reference picture index can be divided and signaled separately for L0 prediction (refidxL0) and L1 prediction (refidxL1). For example, when MVP mode is applied and bi-indication prediction (BI prediction) is applied, both the information regarding refidxL0 and the information regarding refidxL1 can be signaled.

[0158] When MVP mode is applied, as described above, information about the MVD derived from the encoding device is signaled to the decoding device. The information about the MVD may include, for example, information indicating the x and y components for the MVD absolute value and sign. In this case, information indicating whether the MVD absolute value is greater than 0, whether it is greater than 1, and the rest of the MVD can be signaled stepwise. For example, information indicating whether the MVD absolute value is greater than 1 can be signaled only if the value of the flag information indicating whether the MVD absolute value is greater than 0 is 1.

[0159] For example, information regarding MVD is composed of the syntax shown in the table below, encoded in an encoding device, and then signaled to a decoding device.

[0160] [Table 4]

[0161] For example, in Table 4, the abs_mvd_greater0_flag syntax element indicates whether the difference (MVD) is greater than 0, and the abs_mvd_greater1_flag syntax element indicates whether the difference (MVD) is greater than 1. Furthermore, the abs_mvd_minus2 syntax element indicates information about the value obtained by subtracting 2 from the difference (MVD), and the mvd_sign_flag syntax element indicates information about the sign of the difference (MVD). Also in Table 4, [0] for each syntax element indicates information about L0, and [1] indicates information about L1.

[0162] For example, MVD[compIdx] is derived based on abs_mvd_greater0_flag[compIdx]*(abs_mvd_minus2[compIdx]+2)*(1-2*mvd_sign_flag[compIdx]). Here, compIdx (or cpIdx) indicates the index of each component and can have a value of 0 or 1. For example, compIdx 0 indicates the x component, and compIdx1 indicates the 7 component. However, this is an example, and values ​​can be represented for each component using other coordinate systems instead of the x,y coordinate system.

[0163] On the other hand, MVDs for L0 prediction (MVD L0) and MVDs for L1 prediction (MVD L1) may be signaled separately, and the information regarding the MVD may include information regarding MVD L0 and / or information regarding MVD L1. For example, if MVP mode is applied to the current block and BI prediction is applied, both the information regarding MVD L0 and the information regarding MVD L1 are signaled.

[0164] Figure 7 is a diagram illustrating SMVD (symmetric motion vector differences).

[0165] When BI prediction is applied, SMVD (symmetric MVD) may be used to improve coding efficiency. In this case, some signaling of motion information may be omitted. For example, when SMVD is applied to the current block, information regarding refidxL0, refidxL1, and MVD L1 can be derived internally without being signaled from the encoding device to the decoding device. For example, when MVP mode and BI prediction are applied to the current block, flag information indicating whether SMVD is applicable (e.g., SMVD flag information or sym_mvd_flag syntax element) is signaled, and if the value of the flag information is 1, the decoding device determines that SMVD is applied to the current block.

[0166] When SMVD mode is applied (i.e., when the value of the SMVD flag information is 1), information regarding mvp_l0_flag, mvp_l1_flag, and MVD L0 (motion vector difference L0) is explicitly signaled, and the signaling of information regarding refidxL0, refidx1, and MVD L1 (motion vector difference L1) is omitted and can be derived internally, as described above. For example, refidxL0 can be derived in the POC procedure within reference picture list 0 (which may be called list0 or L0) as an index pointing to the closest previously referenced picture to the current picture. refidxL1 can be derived in the POC procedure within reference picture list 1 (which may be called list1 or L1) as an index pointing to the closest subsequent referenced picture to the current picture. Alternatively, for example, both refidxL0 and refidxL1 can be derived as 0. Alternatively, for example, refidxL0 and refidxL1 can be derived as the smallest index having the same POC difference in relation to the current picture. Specifically, for example, if "[POC of the current picture] - [POC of the first referenced picture indicated by refidxL0]" is called the first POC difference, and "[POC of the current picture] - [POC of the second referenced picture indicated by refidxL1]" is called the second POC difference, then only if the first POC difference and the second POC difference are the same, the value of refidxL0 pointing to the first referenced picture may be derived as refidxL0 of the current block, and the value of refidxL1 pointing to the second referenced picture may be derived as refidxL1 of the current block. Furthermore, for example, if there are multiple sets where the first POC difference and the second POC difference are the same, the refidxL0 and refidxL1 of the set with the smallest difference can be derived as the refidxL0 and refidxL1 of the current block.

[0167] As shown in Figure 7, reference picture list 0, reference picture list 1, and MVD L0 and MVD L1 are shown. Here, MVD L1 is symmetrical to MVD L0.

[0168] MVD L1 can be derived as minus (-)MVD L0. For example, the final (improved or modified) motion information (motion vector: MV) for the current block is derived based on the following formula:

[0169]

number

[0170] In Equation 1, mvx0 and mvy0 represent the x and y components of the motion vector for L0 motion information or L0 prediction, and mvx1 and mvy1 represent the x and y components of the motion vector for L1 motion information or L1 prediction. Furthermore, mvpx0 and mvpy0 represent the x and y components of the motion vector predictor for L0 prediction, and mvpx1 and mvpy1 represent the x and y components of the motion vector predictor for L1 prediction. In addition, mvdx0 and mvdy0 represent the x and y components of the motion vector difference for L0 prediction.

[0171] On the other hand, in MMVD mode, motion information used directly for generating prediction samples of the current block (i.e., current CU) can be implicitly derived as a way to apply MVD (motion vector difference) to merge mode. For example, an MMVD flag (e.g., mmvd_flag) indicating whether or not to use MMVD on the current block (i.e., current CU) can be signaled, and MMVD can be performed based on this MMVD flag. If MMVD is applied to the current block (e.g., mmvd_flag is 1), additional information for MMVD can be signaled.

[0172] Here, additional information for MMVD includes a merge candidate flag (e.g., mmvd_cand_flag) indicating whether the first or second candidate in the merge candidate list is used with the MVD, a distance index (e.g., mmvd_distance_idx) indicating the motion magnitude, and a direction index (mmvd_direction_idx) indicating the motion direction.

[0173] In MMVD mode, two candidates located in the first and second entries of the merge candidate list (i.e., the first candidate or the second candidate) can be used, and either of these two candidates (i.e., the first candidate or the second candidate) can be used as the base MV. For example, a merge candidate flag (e.g., mmvd_cand_flag) can be signaled to indicate either of the two candidates (i.e., the first candidate or the second candidate) in the merge candidate list.

[0174] Furthermore, the distance index (e.g., mmvd_distance_idx) indicates the magnitude of the motion and can specify a predetermined offset from the starting point. This offset may be added to the horizontal or vertical component of the starting motion vector. The relationship between the distance index and the predetermined offset can be shown in the following table.

[0175] [Table 5]

[0176] Referring to Table 5 above, the MVD distance (e.g., MmvdDistance) is determined by the value of the distance index (e.g., mmvd_distance_idx), and the MVD distance (e.g., MmvdDistance) can be derived using integer sample precision or fractional sample precision based on the value of tile_group_fpel_mmvd_enabled_flag. For example, if tile_group_fpel_mmvd_enabled_flag is 1, it indicates that the MVD distance is currently derived using integer sample precision in the tile group (or picture header), and if tile_group_fpel_mmvd_enabled_flag is 0, it indicates that the MVD distance is derived using fractional sample precision in the tile group (or picture header). In Table 1, the information (flags) for tile groups can be replaced with the information for picture headers; for example, tile_group_fpel_mmvd_enabled_flag can be replaced with ph_fpel_mmvd_enabled_flag (or ph_mmvd_fullpel_only_flag).

[0177] Furthermore, the direction index (e.g., mmvd_direction_idx) indicates the direction of the MVD relative to the starting point, and shows four directions as shown in Table 5 below. Here, the direction of the MVD can also indicate the sign of the MVD. The relationship between the direction index and the MVD sign is shown in the table below.

[0178] [Table 6]

[0179] Referring to Table 6 above, the sign of the MVD (e.g., MmvdSign) is determined by the value of the direction index (e.g., mmvd_direction_idx), and the sign of the MVD (e.g., MmvdSign) is derived for the L0 reference picture and the L1 reference picture.

[0180] Based on the distance index (e.g., mmvd_distance_idx) and direction index (e.g., mmvd_direction_idx) as described above, the MVD offset can be calculated using the following formula.

[0181]

number

[0182]

number

[0183] In equations 2 and 3, the MMVD distance (MmvdDistance[x0][y0]) and MMVD sign (MmvdSign[x0][y0][0], MmvdSign[x0][y0][1]) are derived based on Table 5 and / or Table 6. In summary, in MMVD mode, a merge candidate indicated by a merge candidate flag (e.g., mmvd_cand_flag) is selected from the merge candidate children of the merge candidate list derived based on the surrounding blocks, and the selected merge candidate can be used as a base candidate (e.g., MVP). Then, the motion information (i.e., motion vector) of the current block can be derived by adding the MVD derived using the distance index (e.g., mmvd_distance_idx) and direction index (e.g., mmvd_direction_idx) based on the base candidate.

[0184] Based on the motion information derived by the prediction mode, a predicted block can be derived for the current block. The predicted block includes a predicted sample (predicted sample array) of the current block. If the motion vector of the current block points to fractional sample units, an interpolation procedure can be performed, thereby deriving a predicted sample of the current block based on a reference sample in fractional sample units within the reference picture. When biprediction is applied, a predicted sample derived by a weighted sum or weighted average (with respect to phase) of the predicted sample derived based on the L0 prediction (i.e., prediction using the reference picture in the reference picture list L0 and MVL0) and the predicted sample derived based on the L1 prediction (i.e., prediction using the reference picture in the reference picture list L1 and MVL1) can be used as the predicted sample of the current block. When biprediction is applied, if the reference picture used for the L0 prediction and the reference picture used for the L1 prediction are located in different temporal directions relative to the current picture (i.e., it is biprediction but corresponds to bidirectional prediction), this may be called true biprediction.

[0185] As mentioned above, reconstructed samples and pictures are generated based on the derived predicted samples, and then procedures such as in-loop filtering can be performed.

[0186] As mentioned above, according to this document, when dual prediction is applied to a block, the predicted samples can be derived based on a weighted average. Previously, the dual prediction signal (i.e., the dual prediction sample) was derived by the simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, the dual prediction sample was derived as the average of the L0 prediction sample based on the L0 reference picture and MVL0 and the L1 prediction sample based on the L1 reference picture and MVL1. However, according to this document, when dual prediction is applied, the dual prediction signal (dual prediction sample) can be derived by the weighted average of the L0 prediction signal and the L1 prediction signal, as follows.

[0187] In the aforementioned embodiment related to MMVD, a method can be proposed that considers long-term reference pictures in the MVD induction process of MMVD, thereby maintaining and increasing compression efficiency in various applications. Furthermore, the method proposed in the embodiment of this document can be applied not only to the MMVD technology used in MERGE, but also to SMVD, a symmetric MVD technology used in intermode (MVP mode).

[0188] Figure 8 illustrates the method for deriving motion vectors in interpretation.

[0189] In one embodiment of this document, an MV induction method that considers long-term reference pictures is used in the motion vector scaling (MV scaling) process of a temporal motion candidate (temporal merge candidate, or temporal mvp candidate). The temporal motion candidate can correspond to an mvCol (mvLXCol). The temporal motion candidate may also be called "TMVP".

[0190] The following table explains the definition of a long-term reference picture.

[0191] [Table 7]

[0192] Referring to Table 7 above, if LongTermRefPic(aPic, aPb, refIdx, LX) is 1 (true), the corresponding reference picture is marked as used for long-term reference. For example, a reference picture that is not marked as used for long-term reference may be a reference picture marked as used for short-term reference. In other examples, a reference picture that is neither marked as used for long-term reference nor as unused may be a reference picture marked as used for short-term reference. Hereinafter, a reference picture marked as used for long-term reference may be referred to as a long-term reference picture, and a reference picture marked as used for short-term reference may be referred to as a short-term reference picture.

[0193] The following table explains the derivation of TMVP(mvLXCol).

[0194] [Table 8]

[0195] Referring to Figure 8 and Table 8, the time motion vector (mvLXCol) is not used unless the type of reference picture pointed to by the current picture (for example, whether it is a long-term reference picture (LTRP) or a short-term reference picture (STRP)) and the type of collocated reference picture pointed to by the collocated picture are the same. That is, if all are long-term reference pictures or all are short-term reference pictures, colMV is induced; if they are of other types, colMV is not induced. Also, if all are long-term reference pictures, or if the POC difference between the current picture and its reference picture is the same as the POC difference between the collocated picture and its reference picture, the collocated motion vector can be used directly without scaling. If they are short-term reference pictures and the POC differences are different, the scaled motion vector of the collocated block is used.

[0196] In the embodiments described in this document, the MMVD used in MERGE / SKIP mode signals the base motion vector index (base MV index), distance index, and direction index for a single coding block as information to guide the MVD information. For unidirectional prediction, the MVD is guided from motion information, and for bidirectional prediction, symmetric MVD information is generated using mirroring and scaling methods.

[0197] When performing bidirectional prediction, MVD information for L0 or L1 is scaled to generate MVDs for L1 or L0. However, when referencing long-term reference pictures, modifications are required during the MVD induction process.

[0198] Figures 9 to 13 illustrate the MVD induction method for MMVD according to the embodiments of this document. The methods shown in Figures 9 to 13 may apply to blocks to which bidirectional prediction is applied.

[0199] In one embodiment shown in Figure 9, if the distance to the L0 reference picture and the distance to the L1 reference picture are the same, the induced MmvdOffset can be used directly as the MVD. When the POC differences (POC difference between the L0 reference picture and the current picture and the POC difference between the L1 reference picture and the current picture) are different, the MVD can be induced by scaling or simple mirroring (i.e., -1 *MmvdOffset) depending on the POC difference and whether it is a long-term or short-term reference picture.

[0200] As an example, the method of using MMVD to induce symmetric MVD for blocks to which bidirectional prediction is applied is not suitable for blocks that use long-term reference pictures, and in particular, when the reference picture types in each direction are different, performance improvement when using MMVD is unlikely to be expected. Therefore, the following figure and embodiment show an example in which MMVD is not applied when the reference picture types of L0 and L1 are different.

[0201] In one embodiment shown in Figure 10, different MVD induction methods are applied depending on whether the reference picture referenced by the current picture (or current slice, current block) is an LTRP (long-term reference picture) or a STRP (short-term reference picture). In one example, when the method of the embodiment shown in Figure 10 is applied, a portion of the standard documentation for this embodiment is described as follows:

[0202] [Table 9-1]

[0203] [Table 9-2]

[0204] In one embodiment shown in Figure 11, different MVD induction methods are applied depending on whether the reference picture referenced by the current picture (or current slice, current block) is an LTRP (long-term reference picture) or a STRP (short-term reference picture). In one example, when the method of the embodiment shown in Figure 11 is applied, a portion of the standard documentation for this embodiment is described as follows:

[0205] [Table 10-1]

[0206] [Table 10-2]

[0207] In summary, the MVD induction process of MMVD, which does not induce MVD when the reference picture types in each direction are different, is explained.

[0208] In one embodiment shown in Figure 12, MVD is not induced in all cases where a long-term reference picture is referenced. That is, if at least one of the L0 and L1 reference pictures is a long-term reference picture, MVD is set to 0, and MVD can be induced only when there is a short-term reference picture.

[0209] In one example, the MVD for the MMVD can be derived when the current picture (or current slice, current block) refers only to a shot-term reference picture, based on the highest priority condition (RefPicL0!=LTRP&&RefPicL1!=STRP). In one example, when the method of the embodiment shown in Figure 12 is applied, a portion of the standard documentation according to this embodiment is described as follows:

[0210] [Table 11-1]

[0211] [Table 11-2]

[0212] In one embodiment shown in Figure 13, if the reference picture types in each direction are different, the MVD is induced when there is a short-term reference picture, and the MVD is induced to 0 when there is a long-term reference picture.

[0213] In one example, if the reference picture types differ in each direction, MmvdOffset is applied when referencing a reference picture that is close to the current picture (short-term reference picture), and MVD has a value of 0 when referencing a reference picture that is far from the current picture (long-term reference picture). Here, a picture close to the current picture can be considered to have a short-term reference picture, but if the close picture is a long-term reference picture, mmvdOffset can be applied to the motion vector of the list pointing to the short-term reference picture.

[0214] [Table 12]

[0215] For example, the four paragraphs included in Table 12 can sequentially replace the bottom block (content) of the sequence diagram included in Figure 13.

[0216] In one example, when the method of the embodiment shown in Figure 13 is applied, a portion of the standard documentation according to this embodiment is described as follows:

[0217] [Table 13-1]

[0218] [Table 13-2]

[0219] The following table shows a comparison table between the examples included in this document.

[0220] [Table 14]

[0221] Referring to Table 14, a comparison is shown between methods for applying an offset considering the reference picture type for MVD derivation of MMVD as described in the embodiments shown in Figures 9 to 13. In Table 14, Embodiment A relates to an existing MMVD, Embodiment B shows the embodiment shown in Figures 9 to 11, Embodiment C shows the embodiment shown in Figure 12, and Embodiment D shows the embodiment shown in Figure 13.

[0222] Specifically, the embodiments shown in Figures 9, 10, and 11 describe a method for inducing MVD only when the reference picture types in both directions are the same, while the embodiment shown in Figure 12 describes a method for inducing MVD only when both directions use short-term reference pictures. In the embodiment shown in Figure 12, MVD is set to 0 if the reference picture is a long-term reference picture for unidirectional prediction. Furthermore, the embodiment shown in Figure 13 describes a method for inducing MVD in only one direction when the reference picture types in both directions are different. These differences between embodiments represent various features of the technology described in this document, and it should be understandable to a person with ordinary skill in the art to which this specification belongs that the effects to be achieved by the embodiments described in this document can be realized based on these features.

[0223] In the embodiments described herein, a separate process is required when the reference picture type is a long-term reference picture. When long-term reference pictures are included, POCDiff-based scaling or mirroring does not affect performance improvement, so the MVD in the direction with short-term reference pictures is assigned an MmvdOffset value, and the MVD in the direction with long-term reference pictures is assigned a value of 0. In one example, when this embodiment is applied, a portion of the standard documentation according to this embodiment is described as follows:

[0224] [Table 15-1]

[0225] [Table 15-2]

[0226] In other examples, portions of Table 15 can be replaced with the following table. Referring to Table 16, the Offset is applied based on the reference picture type, not the POCDiff.

[0227] [Table 16]

[0228] In other examples, parts of Table 15 can be replaced with the following table. Referring to Table 17, it is possible to always set MmvdOffset to L0 and -MmvdOffset to L1 without considering the reference picture type.

[0229] [Table 17]

[0230] According to one embodiment of this document, intermode SMVD can be performed in a manner similar to the MMVD used in the MERGE mode described above. When bidirectional prediction is performed, the possibility of symmetric MVD derivation is signaled from the encoding device to the decoding device, and when the relevant flag (e.g., sym_mvd_flag) is true (or its value is 1), the second direction MVD (e.g., MVD L1) is induced by mirroring the first direction MVD (e.g., MVD L0). In this case, scaling of the first direction MVD may not be performed.

[0231] The following table shows the syntax for a coding unit according to one embodiment of this document.

[0232] [Table 18]

[0233] [Table 19]

[0234] Referring to Tables 18 and 19 above, if inter_pred_idc == PRED_BI and the reference pictures for L0 and L1 are available (for example, RefIdxSymL0 > -1 &&RefIdxSymL1 > -1), then sym_mvd_flag is signaled.

[0235] The following table shows an example of a decoding procedure for an MMVD reference index.

[0236] [Table 20]

[0237] Referring to Table 20, the procedure for deriving the availability of the reference pictures of L0 and L1 is described. That is, if there is a reference picture in the forward direction among the L0 reference pictures, the reference picture index closest to the current picture is set to RefIdxSymL0, and the corresponding value is set to the reference index of L0. Also, if there is a reference picture in the backward direction among the L1 reference pictures, the reference picture index closest to the current picture is set to RefIdxSymL1, and the corresponding value is set to the reference index of L1.

[0238] The following Table 21 shows the decoding procedure for the MMVD reference index according to another example.

[0239]

Table 21

[0240] Referring to Table 21, when the L0 or L1 reference picture types are different as in the embodiments described with FIGS. 9, 10, and 11, that is, when long-term reference pictures and short-term reference pictures are used, after the reference index derivation for SMVD to prevent SMVD, if the L0 and L1 reference picture types are different, do not use SMVD (see the bottom paragraph of Table 20).

[0241] In one embodiment of this document, similar to MMVD used in merge mode, SMVD can be applied in inter mode. As in the embodiment described with FIG. 12, when long-term reference pictures are used, in order to prevent SMVD, long-term reference pictures can be excluded in the process of deriving the reference index for SMVD as shown in the following table.

[0242]

Table 22

[0243] The following table, based on another example of this embodiment, shows an example of how to avoid applying SMVD when using a long-term reference picture after the reference picture index induction for SMVD.

[0244] [Table 23]

[0245] In one embodiment of this document, when the reference picture type of the current picture and the reference picture type of the collated picture are different during the colMV induction process of TMVP, the motion vector MV is set to 0. However, this differs from the induction method in the case of MMVD and SMVD, so this should be standardized.

[0246] Currently, even when the reference picture type of a picture is a long-term reference picture and the reference picture type of a collated picture is also a long-term reference picture, the motion vector uses the collated motion vector value as is. However, in MMVD and SMVD, in this case, MV is set to 0. Here, TMVP is also set to 0 without any additional guidance.

[0247] Furthermore, even if the reference picture type is different, there may be long-term reference pictures that are close to the current picture. Taking this into consideration, instead of setting MV to 0, colMV can be used as MV without scaling.

[0248] Figure 14 is a diagram illustrating SMVD according to one embodiment of this document.

[0249] The method shown in Figure 14 can be used to derive the SMVD. That is, the SMVD can be derived based on STRP (Short-Term Reference Picture) and / or LTRP (Long-Term Reference Picture). When using a mirrored L0 MVD for the L1 MVD, an inaccurate MVD may be derived if the types of reference pictures are different. This is because the ratio of distances (the distance between reference picture 0 and the current picture and the distance between reference picture 1 and the current picture) becomes larger, and the correlation of the motion vectors in each direction decreases.

[0250] According to one embodiment of this document, the availability of a reference picture is checked, and if the conditions are met, sym_mvd_flag can be parsed. If sym_mvd_flag is true, the MVD of L1 (MVDL1) can be mirrored to MVDL0 (MVD of L0).

[0251] The following table shows a portion of the coding unit syntax according to this embodiment.

[0252] [Table 24]

[0253] Based on Table 24, the procedure for deriving sym_mvd_flag according to this embodiment can be explained.

[0254] In this embodiment, a reference picture index (RefIdxSymLX with X=0,1) for SMVD can be derived. RefIdxSymL0 can indicate the nearest reference picture (or its index) that has a POC smaller than the current picture's POC. RefIdxSymL1 can indicate the nearest reference picture (or its index) that has a POC larger than the current picture's POC.

[0255] The following table describes, in the format of a standard document, the method for deriving the reference picture index for SMVD according to this embodiment.

[0256]

Table 25

[0257] The following table shows the comparison results among the embodiments. By considering the reference picture type in the embodiments included in Table 26, the accuracy of MVD in SMVD can be improved. In Table 26, MVD can indicate MVD 0 (the MVD of L0).

[0258]

Table 26

[0259] Referring to Table 26, Example P shows the existing method for deriving SMVD. In Example Q, SMVD may be restricted when using a mixed reference picture type (ex. STRP / LTRP or LTRP / STRP) at L0 and L1. In Example R, SMVD may be restricted when referring to a long-term reference picture (LTRP).

[0260] The following table describes, in the format of a standard document, the method for deriving the reference picture index for SMVD according to Example Q of Table 26.

[0261]

Table 27

[0262] The following table describes, in the format of a standard document, the method for deriving the reference picture index for SMVD according to Example Q of Table 26.

[0263]

Table 28

[0264] [Table 29]

[0265] Referring to Tables 28 and / or 29, SMVD may be restricted when referencing long-term reference pictures (LTRPs). For example, referring to Table 28, long-term reference pictures can be excluded in the reference picture checking process. This allows other reference pictures (e.g., not long-term reference pictures) to be considered for SMVD. Referring to Table 29, SMVD may not be performed if the closest reference picture to the current picture is a long-term reference picture. For example, even if the reference picture list contains short-term reference pictures, SMVD may not be performed if the closest reference picture to the current picture is a long-term reference picture.

[0266] In one example according to one embodiment of this document, if the POC distance of L0 is greater than or equal to the POC of L1 in the MMVD procedure, the L1 MVD can be derived as a scaled or mirrored L0 MVD. If the POC distance of L0 is less than the POC of L1 in the MMVD procedure, the L0 MVD can be derived as a scaled or mirrored L1 MVD in the MMVD procedure.

[0267] Figure 15 is a flowchart showing a method for deriving MMVD according to one embodiment of this document.

[0268] In one embodiment of this document, the MVD can be derived in MMVD, taking into account the POC difference and / or the reference picture type. Referring to Figure 15, currPocDiffLX can represent the difference between the POC of the current picture and the POC of the reference picture LX. CurrPocDiffL0 and currPocDiffL1 can be compared with each other, and the type of the reference picture can be checked ("refPicList0 != LTRP" or "refPicList1 != LTRP"). Taking the conditions into account, MmvdOffset (derived using mmvd_cand_flag, mmvd_distance_idx, and / or mmvd_direction_idx) can be assigned as the same value as mMvdLX, a mirrored value, or a scaled value.

[0269] The following table shows a portion of the standard documentation according to this embodiment.

[0270] [Table 30-1]

[0271] [Table 30-2]

[0272] Currently, when a picture references one or more Long-Term Reference Pictures (LTRPs), a mirroring procedure that considers the Point of Constraint (POC) distance may not be necessary. This is because a mirrored MVD obtained from a reference picture that is much farther away than other MVDs is ineffective in terms of accuracy. A solution to this is described below.

[0273] The following table shows the comparison results between the examples.

[0274] [Table 31]

[0275] Referring to Table 31, Example X shows an existing method for deriving MMVD. In Example Y, the MMVD procedure may be restricted to cases where one or more long-term reference pictures are referenced in the current block. That is, in Example Y, the step of comparing POC distances for long-term reference pictures may be omitted. In Example Z, the MMVD derivation procedure may be restricted for all cases. That is, in Example Z, the step of comparing POC distances may be omitted for all cases. In Table 31, offset may refer to MmvdOffset.

[0276] Figure 16 is a flowchart illustrating a method for deriving MMVD according to one embodiment of this document. The flowchart in Figure 16 illustrates the method for deriving MMVD according to the aforementioned embodiment Y.

[0277] Referring to Figure 16, the condition for comparing POC differences can be removed when the reference picture type is a long-term reference picture, and the anchor MVD used for the mirroring procedure can be fixed to the L0 MVD.

[0278] The following table describes, in standard document format, the method for deriving MMVD according to Example Y in Table 31.

[0279] [Table 32-1]

[0280] [Table 32-2]

[0281] Figure 17 is a flowchart illustrating a method for deriving MMVD according to one embodiment of this document. The flowchart in Figure 17 illustrates the method for deriving MMVD according to the aforementioned embodiment Z.

[0282] Referring to Figure 17, in Example Z, the MMVD derivation procedure can be restricted for all cases. For all cases, the condition for comparing POC differences can be removed, and the anchor MVD used for the mirroring or scaling procedure can be fixed to the L0 MVD.

[0283] The following table describes, in standard document format, the method for deriving MMVD using Example Z in Table 31.

[0284] [Table 33-1]

[0285] [Table 33-2]

[0286] Furthermore, in one example of this embodiment, the condition for comparing POC differences may be removed in all cases, and only the mirroring approach may be used. The following table describes the method for deriving the MMVD in this example in the format of a standard document.

[0287] [Table 34]

[0288] The following drawings have been prepared to illustrate a specific example of this specification. The names of specific devices and signal message fields shown in the drawings are presented illustratively, and the technical features of this specification are not limited to the specific names used in the following drawings.

[0289] Figures 18 and 19 schematically show an example of a video / image encoding method and related components according to the embodiments of this document. The method disclosed in Figure 18 can be performed by the encoding apparatus disclosed in Figure 2. Specifically, for example, steps S1800 to S1850 in Figure 18 can be performed by the prediction unit 220 of the encoding apparatus, and S1860 can be performed by the residual processing unit 230 of the encoding apparatus. S1870 can be performed by the entropy encoding unit 240 of the encoding apparatus. The method disclosed in Figure 18 may include embodiments described above in this document.

[0290] Referring to Figure 18, the encoding device derives an interprediction mode for the current block in the current picture (S1800). Here, the interprediction mode can include the merge mode, AMVP mode (mode using motion vector predictor candidates), MMVD, and SMVD as described above.

[0291] The encoding device can derive a reference picture for the interpretation mode. The encoding device can configure a reference picture list for deriving the reference picture. In one example, the reference picture list may include reference picture list 0 (or L0, reference picture list L0) or reference picture list 1 (or L1, reference picture list L1). For example, the encoding device can configure a reference picture list for each slice currently contained in the picture.

[0292] The encoding device constructs an MVP candidate list for the current block based on the surrounding blocks of the current block (S1810). The MVP candidate list may include MVP candidate list L0 and MVP candidate list L1. In one example, the surrounding blocks may be included in the current picture containing the current block. In another example, the surrounding blocks may be included in a previous (reference) picture or a subsequent (reference) picture from the current picture. Here, the POC of the previous picture may be smaller than the POC of the current picture, and the POC of the subsequent picture may be larger than the POC of the current picture. In one example, the POC difference between the current picture and a previous (reference) picture from the current picture may be greater than 0. In another example, the POC difference between the current picture and a subsequent (reference) picture from the current picture may be less than 0. However, this is only illustrative.

[0293] The encoding device can derive an MVP for the current block based on the MVP candidate list (S1820). The MVP may include MVPL0 and MVPL1. MVPL0 can be derived from the MVP candidate list L0, and MVPL1 can be derived from the MVP candidate list L1. The encoding device can derive the optimal motion vector predictor candidate from among the motion vector predictor candidates included in the MVP candidate list. The encoding device can generate selection information (e.g., MVP flag or MVP index) indicating the optimal motion vector predictor candidate.

[0294] The encoding device generates prediction-related information including the interpretation mode (S1830). In one example, the prediction-related information may include information regarding the MVD (motion vector difference) for the current block. The prediction-related information may also include information regarding MMVD, SMVD, etc.

[0295] The encoding device derives motion information for predicting the current block based on the information regarding the MVP and the MVD (S1840). For example, the motion information may include a reference index for the SMVD (symmetric motion vector difference reference index). The reference index for the SMVD may point to a reference picture for applying the SMVD. The reference index for the SMVD may include reference index L0 (RefIdxSumL0) and reference index L1 (RefIdxSumL1).

[0296] The encoding device generates predicted samples based on the motion information (S1850). The encoding device can generate the predicted samples based on the motion vectors and reference picture index included in the motion information. For example, the predicted samples can be generated based on the blocks (or samples) within the reference picture pointed to by the reference picture index that are indicated by the motion vectors.

[0297] The encoding device derives residual information based on the predicted sample (S1860). Specifically, the encoding device can derive a residual sample based on the predicted sample and the original sample. The encoding device can derive residual information based on the residual sample. The transformation and quantization processes described above can be performed to derive the residual information.

[0298] The encoding device encodes image / video information including the prediction-related information and the residual information (S1870). The encoded image / video information can be output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or (digital) storage medium.

[0299] The aforementioned image / video information may include a variety of information according to the embodiments of this document. For example, the image / video information may include information disclosed in at least one of the tables 1 to 34 described above.

[0300] In one embodiment, the motion information may include a motion vector (MV) and a symmetric motion vector difference reference index. The MV may include MVL0 for the L0 prediction and MVL1 for the L1 prediction. The symmetric motion vector difference reference index may include a symmetric motion vector difference reference index L0 for the L0 prediction and a symmetric motion vector difference reference index L1 for the L1 prediction. The information regarding the MVD may include information regarding MVDL0 for the L0 prediction. In one example, information regarding MVDL1 for the L1 prediction can be derived based on the information regarding MVDL0. In another example, information regarding MVDL0 and / or MVDL1 can be derived based on the surrounding blocks for the prediction of the current block. The encoding device may exclude information regarding MVDL1 when encoding image / video information. The MVL0 may be derived based on the information regarding MVDL0, and the MVL1 may be derived based on the information regarding MVDL1. The aforementioned symmetric motion vector difference reference index L0 and the aforementioned symmetric motion vector difference reference index L1 can be derived based on the short-term reference picture among the reference pictures included in the reference picture list.

[0301] In one embodiment, the MVP may include an MVPL0 for the L0 prediction and an MVPL1 for the L1 prediction. The MVL0 can be derived based on the sum of the MVDL0 and the MVPL0. The MVL1 can be derived based on the sum of the MVDL1 and the MVPL1.

[0302] In one embodiment, the prediction-related information may include information regarding the symmetric motion vector difference (SMVD information or SMVD flag information). If the symmetric motion vector difference reference index L0 and the symmetric motion vector difference reference index L1 are derived based on the picture order count (POC) difference between the short-term reference picture and the current picture including the current block, the value of the information regarding the symmetric motion vector difference may be 1.

[0303] In one embodiment, the size of MVDL1 may be the same as the size of MVDL0. The reference numeral of MVDL1 may be opposite to the reference numeral of MVDL0.

[0304] In one embodiment, the short-term reference picture may include short-term reference picture L0 and short-term reference picture L1. For example, the symmetric motion vector difference reference index L0 may point to the short-term reference picture L0. Also, the symmetric motion vector difference reference index L1 may point to the short-term reference picture L1.

[0305] In one embodiment, the reference picture list may include reference picture list 0. The reference picture list 0 may include the short-term reference picture L0. Based on the picture order count (POC) difference between each of the short-term reference pictures included in reference picture list 0 and the current picture including the current block, the symmetric motion vector difference reference index L0 can be derived.

[0306] In one embodiment, the symmetric motion vector difference reference index L0 can be derived based on a comparison between the POC differences.

[0307] In one embodiment, the reference picture list 0 may further include other short-term reference pictures L0. The POC difference may include a first POC difference between the short-term reference picture L0 and the current picture, and a second POC difference between the other short-term reference picture L0 and the current picture. The first POC difference may be smaller than the second POC difference.

[0308] Figures 20 and 21 schematically show an example of an image / video decoding method and related components according to the embodiments of this document. The method disclosed in Figure 20 can be performed by the decoding device disclosed in Figure 3. Specifically, for example, S2000 in Figure 20 can be performed by the entropy decoding unit 310 of the decoding device, and S2010 to S2050 can be performed by the prediction unit 330 of the decoding device. The method disclosed in Figure 20 may include embodiments described above in this document.

[0309] Referring to Figure 20, the decoding device receives / acquires image / video information (S2000). The decoding device can receive / acquire the image / video information via a bitstream. The image / video information may include prediction-related information (including prediction mode information), information regarding MVD, and / or residual information. The prediction-related information may include information regarding MMVD, information regarding SMVD, etc. Furthermore, the image / video information may include various types of information according to the embodiments of this document. For example, the image / video information may include the information described with Figures 1 to 17 and / or the information disclosed in at least one of Tables 1 to 34 mentioned above.

[0310] The decoding device derives an inter-prediction mode for the current block based on the prediction-related information (S2010). Here, the inter-prediction mode can include the merge mode, AMVP mode (a mode using motion vector predictor candidates), MMVD, and SMVD.

[0311] The decoding device constructs an MVP candidate list for the current block based on the surrounding blocks of the current block (S2020). The MVP candidate list may include MVP candidate list L0 and MVP candidate list L1. In one example, the surrounding blocks may be contained within the current picture containing the current block. In another example, the surrounding blocks may be contained within a previous (reference) picture or a subsequent (reference) picture from the current picture, where the POC of the previous picture may be smaller than the POC of the current picture, and the POC of the subsequent picture may be larger than the POC of the current picture. In one example, the POC difference between the current picture and a previous (reference) picture from the current picture may be greater than 0. In another example, the POC difference between the current picture and a subsequent (reference) picture from the current picture may be less than 0. However, this is illustrative only.

[0312] The decoding device can derive an MVP for the current block based on the MVP candidate list (S2030). The MVP may include MVPL0 and MVPL1. MVPL0 can be derived from the MVP candidate list L0, and MVPL1 can be derived from the MVP candidate list L1. The decoding device can derive the optimal motion vector predictor candidate from among the motion vector predictor candidates included in the MVP candidate list. The encoding device can generate selection information (e.g., MVP flag or MVP index) indicating the optimal motion vector predictor candidate.

[0313] The decoding device derives motion information for the current block based on the information regarding the MVD and the MVP (S2040). For example, the motion information may include a reference index for the SMVD. The reference index for the SMVD may point to a reference picture for the application of the SMVD. The reference index for the SMVD may include reference index L0 (RefIdxSumL0) and reference index L1 (RefIdxSumL1).

[0314] The decoding device generates a predicted sample based on the motion information (S2050). The decoding device can generate the predicted sample based on the motion vector and reference picture index included in the motion information. For example, the predicted sample can be generated based on the block (or sample) indicated by the motion vector among the blocks (or samples) in the reference picture pointed to by the reference picture index.

[0315] The decoding device can generate a residual sample based on the residual information. Specifically, the decoding device can derive quantized transformation coefficients based on the residual information. The quantized transformation coefficients may take the form of a one-dimensional vector based on the coefficient scan order. The decoding device can derive transformation coefficients based on an inverse quantization procedure on the quantized transformation coefficients. The decoding device can derive a residual sample based on an inverse transformation procedure on the transformation coefficients.

[0316] The decoding device can generate a restored sample of the current picture based on the predicted sample and the residual sample. The decoding device can also perform further filtering steps to generate a (corrected) restored sample.

[0317] In one embodiment, the motion information may include a motion vector (MV) and a symmetric motion vector difference reference index. The MV may include MVL0 for the L0 prediction and MVL1 for the L1 prediction. The symmetric motion vector difference reference index may include a symmetric motion vector difference reference index L0 for the L0 prediction and a symmetric motion vector difference reference index L1 for the L1 prediction. The information regarding the MVD may include information regarding MVDL0 for the L0 prediction. In one example, information regarding MVDL1 for the L1 prediction can be derived based on the information regarding MVDL0. In another example, information regarding MVDL0 and / or MVDL1 can be derived based on surrounding blocks for the prediction of the current block. The encoding device may exclude information regarding MVDL1 when encoding image / video information. The MVL0 may be derived based on the information regarding MVDL0, and the MVL1 may be derived based on the information regarding MVDL1. The aforementioned symmetric motion vector difference reference index L0 and the aforementioned symmetric motion vector difference reference index L1 can be derived based on the short-term reference picture among the reference pictures included in the reference picture list.

[0318] In one embodiment, the MVP may include an MVPL0 for the L0 prediction and an MVPL1 for the L1 prediction. The MVL0 can be derived based on the sum of the MVDL0 and the MVPL0. The MVL1 can be derived based on the sum of the MVDL1 and the MVPL1.

[0319] In one embodiment, the prediction-related information may include information regarding symmetric motion vector differences (information for SMVD or SMVD flag information). For example, if the value of the information regarding symmetric motion vector differences is 1, the symmetric motion vector difference reference index L0 and the symmetric motion vector difference reference index L1 can be derived based on the POC difference between the short-term reference picture and the current picture containing the current block.

[0320] In one embodiment, the size of MVDL1 may be the same as the size of MVDL0. The reference numeral of MVDL1 may be opposite to the reference numeral of MVDL0.

[0321] In one embodiment, the short-term reference picture may include short-term reference picture L0 and short-term reference picture L1. For example, the symmetric motion vector difference reference index L0 may point to the short-term reference picture L0. Also, the symmetric motion vector difference reference index L1 may point to the short-term reference picture L1.

[0322] In one embodiment, the reference picture list may include reference picture list 0. The reference picture list 0 may include the short-term reference picture L0. Based on the picture order count (POC) difference between each of the short-term reference pictures included in reference picture list 0 and the current picture including the current block, the symmetric motion vector difference reference index L0 can be derived.

[0323] In one embodiment, the symmetric motion vector difference reference index L0 can be derived based on a comparison between the POC differences.

[0324] In one embodiment, the reference picture list 0 may further include other short-term reference pictures L0. The POC difference may include a first POC difference between the short-term reference picture L0 and the current picture, and a second POC difference between the other short-term reference picture L0 and the current picture. The first POC difference may be smaller than the second POC difference.

[0325] In one embodiment, the symmetric motion vector difference reference index L0 and the symmetric motion vector difference reference index L1 (ex. ref_idx_l1[x0][y0], ref_idx_l1[x0][y0]) are not directly signaled and can be derived based on information regarding the symmetric motion vector difference (ex. sym_mvd_flag).

[0326] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and that different steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments described herein.

[0327] The methods relating to the embodiments described in this document above can be implemented in software form, and the encoding and / or decoding devices relating to this document may be included in, for example, video processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.

[0328] In this document, when embodiments are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0329] Furthermore, the decoding and encoding devices to which the embodiments described in this document apply may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, customized video (VoD) service providers, OTT video (Over the Top Video) devices, internet streaming service providers, 3D video devices, VR (virtual reality) devices, AR (argumente reality) devices, video telephone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, and may be used to process video signals or data signals. For example, OTT video (Over the Top Video) devices may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0330] Furthermore, the processing methods to which the embodiments of this document apply can be produced in the form of programs executed by a computer and stored on a computer-readable recording medium. Multimedia data having the data structure relating to the embodiments of this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data to be read by a computer. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wired wireless network.

[0331] Furthermore, the embodiments described in this document can be embodied in a computer program product using program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable carrier.

[0332] Figure 22 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0333] Referring to Figure 22, the content streaming system to which the embodiments of this document apply can broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0334] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. As an alternative, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.

[0335] The bitstream can be generated by an encoding method or bitstream generation method to which the embodiments of this document apply, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0336] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.

[0337] The streaming server can receive content from a media storage and / or encoding server. For example, if it begins receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0338] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage.

[0339] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0340] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined to embody an apparatus, and the technical features of the apparatus claims herein can be combined to embody a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody a method.

Claims

1. In an image decoding method performed by a decoding device, A step of obtaining image information from a bitstream, including prediction-related information and information regarding MVD (motion vector difference), The steps include: deriving an inter-prediction mode for the current block based on the aforementioned prediction-related information; The steps include: constructing a list of MVP (motion vector predictor) candidates for the current block based on the surrounding blocks of the current block; The steps include: deriving the MVP for the current block based on the MVP candidate list; The steps include: deriving the MV (motion vector) for the current block based on the MVD and the MVP; The step of generating a predicted sample for the current block based on motion information including the MV and the symmetric motion vector difference reference index, The current block is now subject to dual prediction, The aforementioned MV includes MVL0 for L0 prediction and MVL1 for L1 prediction, The symmetric motion vector difference reference index includes the symmetric motion vector difference reference index L0 for the L0 prediction and the symmetric motion vector difference reference index L1 for the L1 prediction. The information relating to the MVD includes information relating to MVDL0 for the L0 prediction, The MVD includes the MVDL0 for the L0 prediction and the MVDL1 for the L1 prediction, The aforementioned MVDL0 is derived based on the information relating to the aforementioned MVDL0, The aforementioned MVDL1 is derived based on the aforementioned MVDL0, The MVP includes MVPL0 for the L0 prediction and MVPL1 for the L1 prediction, The aforementioned MVL0 is derived based on the sum of the aforementioned MVDL0 and the aforementioned MVPL0, The aforementioned MVL1 is derived based on the sum of the aforementioned MVDL1 and the aforementioned MVPL1. The method by which the symmetric motion vector difference reference index L0 and the symmetric motion vector difference reference index L1 are derived based on the picture order count (POC) difference between a short-term reference picture among the reference pictures included in the reference picture list and the current picture including the current block.

2. In an image encoding method performed by an encoding device, The current step is to derive the interprediction mode for the block, The steps include: constructing a list of MVP (motion vector predictor) candidates for the current block based on the surrounding blocks of the current block; The steps include: deriving the MVP for the current block based on the MVP candidate list; A step of deriving motion information for the current block, including MV (motion vector) and symmetric motion vector difference reference index, The steps include generating prediction-related information including information about the interpretation mode and information about the motion vector difference (MVD) for the current block, The steps include generating a predicted sample for the current block based on the aforementioned motion information, The steps include generating residual information based on the aforementioned prediction samples, The step of encoding image information including the prediction-related information and the residual information, The current block is now subject to dual prediction, The aforementioned MVP includes MVPL0 for L0 prediction and MVPL1 for L1 prediction, The MV includes MVL0 for the L0 prediction and MVL1 for the L1 prediction, The symmetric motion vector difference reference index includes the symmetric motion vector difference reference index L0 for the L0 prediction and the symmetric motion vector difference reference index L1 for the L1 prediction. The information relating to the MVD includes information relating to MVDL0 for the L0 prediction, By subtracting MVPL0 from MVL0, MVDL0 is derived. MVDL1 is derived by subtracting MVPL1 from MVL1. The method by which the symmetric motion vector difference reference index L0 and the symmetric motion vector difference reference index L1 are derived based on the picture order count (POC) difference between a short-term reference picture among the reference pictures included in the reference picture list and the current picture including the current block.

3. A method for transmitting data for images, A step of obtaining a bitstream for the aforementioned image, wherein the bitstream is The current step is to derive the interprediction mode for the block, The steps include: constructing a list of MVP (motion vector predictor) candidates for the current block based on the surrounding blocks of the current block; The steps include: deriving the MVP for the current block based on the MVP candidate list; A step of deriving motion information for the current block, including MV (motion vector) and symmetric motion vector difference reference index, The steps include generating prediction-related information including information about the interpretation mode and information about the motion vector difference (MVD) for the current block, The steps include generating a predicted sample for the current block based on the aforementioned motion information, The steps include generating residual information based on the aforementioned prediction samples, A step of encoding image information including the prediction-related information and the residual information, and a step of generating based on, The step of transmitting the data, which includes the bitstream, The current block is now subject to dual prediction, The aforementioned MVP includes MVPL0 for L0 prediction and MVPL1 for L1 prediction, The MV includes MVL0 for the L0 prediction and MVL1 for the L1 prediction, The symmetric motion vector difference reference index includes the symmetric motion vector difference reference index L0 for the L0 prediction and the symmetric motion vector difference reference index L1 for the L1 prediction. The information relating to the MVD includes information relating to MVDL0 for the L0 prediction, By subtracting MVPL0 from MVL0, MVDL0 is derived. MVDL1 is derived by subtracting MVPL1 from MVL1. The method by which the symmetric motion vector difference reference index L0 and the symmetric motion vector difference reference index L1 are derived based on the picture order count (POC) difference between a short-term reference picture among the reference pictures included in the reference picture list and the current picture including the current block.