Image coding method and apparatus based on inter prediction

The method improves image/video coding efficiency by employing efficient inter prediction techniques, such as signaling motion vector differences and utilizing SMVD flags, to address the challenges of high-resolution and diverse image characteristics, thereby reducing data transmission and storage costs.

JP7691561B2Active Publication Date: 2025-06-11LG ELECTRONICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024147564
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-24
Filing Date
2024-08-29
Publication Date
2025-06-11
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, such as 4K or UHD, has led to higher data transmission and storage costs due to increased bit rates. Additionally, emerging immersive media technologies like VR and AR require efficient image/video compression solutions that can handle diverse image characteristics.

Method used

The proposed method and apparatus enhance image/video coding efficiency through efficient inter prediction, particularly by signaling information related to motion vector differences, dual prediction, and Symmetric Motion Vector Difference (SMVD) flags. This includes deriving L1 motion vector differences based on SMVD flags and utilizing short-term reference pictures for efficient inter prediction.

Benefits of technology

The solution increases overall image/video compression efficiency, enables efficient signaling of motion vector differences, and reduces the complexity of the coding system, thereby addressing the challenges of high-resolution and diverse image characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691561000044
    Figure 0007691561000044
  • Figure 0007691561000045
    Figure 0007691561000045
  • Figure 0007691561000046
    Figure 0007691561000046
Patent Text Reader

Abstract

To provide an image decoding method.SOLUTION: An image decoding method according to the present document may comprise steps of: deriving an inter-prediction mode from encoded information; configuring reference picture lists; deriving motion information having reference picture indexes for symmetric motion vector differences (SMVDs) on the basis of reference pictures included in the reference picture lists; and generating prediction samples on the basis of the motion information, where the reference picture indexes for the SMVDs can be derived on the basis of short-term reference pictures included in the reference picture lists.SELECTED DRAWING: Figure 21
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to an image coding method and apparatus based on inter prediction.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing (conventional) image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] Thus, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce information of high-resolution and high-quality images / videos having various characteristics as described above.

[0005] Also, inter prediction in image / video coding can include procedures for a SMVD (Symmetric Motion Vector Difference) reference index and / or procedures for MMVD (Merge Motion Vector Difference). In the procedure for the SMVD reference index, discussions are underway to apply related technologies considering reference picture marking (e.g., short-term or long-term reference).

Summary of the Invention

Means for Solving the Problems

[0006] According to one embodiment of this document, a method and an apparatus for improving image / video coding efficiency are provided.

[0007] According to one embodiment of this document, a method and an apparatus for performing efficient inter prediction in an image / video coding system are provided.

[0008] According to one embodiment of this document, a method and an apparatus for signaling information related to motion vector differences in inter prediction are provided.

[0009] According to one embodiment of this document, when dual prediction is applied to a current block, a method and an apparatus for signaling information related to L0 motion vector difference and L1 motion vector difference are provided.

[0010] According to an embodiment of this document, a method and an apparatus for signaling an SMVD flag are provided.

[0011] According to one embodiment of this document, a method and an apparatus for deriving an L1 motion vector difference based on an SMVD flag are provided.

[0012] According to one embodiment of this document, procedures related to an SMVD reference index can be performed based on reference picture marking.

[0013] According to one embodiment of this document, procedures related to an SMVD reference index can be performed using a short-term reference picture (a picture marked as being used for short-term reference).

[0014] According to one embodiment of this document, a video / image decoding method executed by a decoding device is provided.

[0015] According to one embodiment of the present document, a decoding device that executes video / image decoding is provided.

[0016] According to one embodiment of the present document, a video / image encoding method executed by an encoding device is provided.

[0017] According to one embodiment of the present document, an encoding device that executes video / image encoding is provided.

[0018] According to one embodiment of the present document, a computer-readable digital storage medium storing encoded video / image information generated by the video / image encoding method disclosed in at least one of the embodiments of the present document is provided.

[0019] According to one embodiment of the present document, a computer-readable digital storage medium storing encoded information or encoded video / image information for executing the video / image decoding method disclosed in at least one of the embodiments of the present document by a decoding device is provided.

Advantages of the Invention

[0020] According to the present document, the overall image / video compression efficiency can be increased.

[0021] According to the present document, information regarding motion vector differences can be efficiently signaled.

[0022] According to the present document, when dual prediction is applied to a current block, the L1 motion vector difference can be efficiently derived.

[0023] According to the present document, the information used to derive the L1 motion vector difference can be efficiently signaled to reduce the complexity of the coding system.

[0024] According to an embodiment of this document, for the derivation of a reference picture index for SMVD, efficient inter prediction can be performed by using short-term reference pictures.

[0025] The effects that can be obtained through a specific example of this document are not limited to the effects listed above. For example, there can be various technical effects that a person having ordinary skill in the related art can understand or derive from this document. Accordingly, the specific effects of this document can include various effects that can be understood or derived from the technical features of this document, rather than being limited to those explicitly described in this document.

Brief Description of the Drawings

[0026]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

DETAILED DESCRIPTION OF THE INVENTION

[0027] The disclosure of this document can be modified in various ways and can have various embodiments. However, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present disclosure to specific embodiments. The terms used in this document are merely used to describe specific embodiments and are not intended to limit the technical idea of the embodiments in this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this document, terms such as "including" or "having" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the document, and should be understood not to preclude in advance the possibility of the presence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.

[0028] On the other hand, each configuration on the drawings described in this document is shown independently for the convenience of explaining different characteristic functions, and does not mean that each configuration is implemented by separate hardware or separate software. For example, among each configuration, two or more configurations may be combined to form one configuration, or one configuration may be divided into a plurality of configurations. Embodiments in which each configuration is integrated and / or separated are also included in the disclosure scope of this document.

[0029] Hereinafter, embodiments of this document will be described with reference to the accompanying drawings. Hereinafter, the same reference numerals may be used for the same components on the drawings, and redundant descriptions for the same components may be omitted.

[0030] FIG. 1 schematically shows an example of a video / image coding system to which the embodiments of this document can be applied.

[0031] As shown in FIG. 1, a video / image coding system can include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in file or streaming form.

[0032] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be referred to as a video / image encoding device, and the decoding device can be referred to as a video / image decoding device. A transmitter can be provided in the encoding device. A receiver can be provided in the decoding device. The renderer can include a display unit, and the display unit can also be composed of a separate device or an external component.

[0033] The video source can obtain video / images through processes such as capture, synthesis, or generation of video / images. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer or the like, and in this case, the video / image capture process can be replaced by the process of generating related data.

[0034] The encoding device can encode the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0035] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the above bitstream and transmit it to the decoding device.

[0036] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0037] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0038] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (Versatile Video Coding) standard. Also, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the EVC (Essential Video Coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of Audio Video coding Standard), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).

[0039] This document presents various embodiments related to video / image coding, and unless otherwise mentioned, the above embodiments may be combined with each other.

[0040] In this document, video can mean a collection of a series of images over the passage of time. A picture generally means a unit indicating one image in a particular time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs may be ordered consecutively in a CTU raster scan within a tile, and tiles within a picture may be ordered consecutively in a raster scan of the tiles of the picture. A slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.

[0041] On the other hand, one picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular region of one or more slices within a picture.

[0042] A pixel or pel can mean the smallest unit that makes up one picture (or image). Also, as a term corresponding to a pixel, "sample" can be used. A sample can generally indicate a pixel or the value of a pixel, and can indicate only the pixel / pixel value of the luma component, or can also indicate only the pixel / pixel value of the chroma component.

[0043] A unit can indicate the basic unit of image processing. A unit can include at least one of a specific area of a picture and information related to that area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0044] In this document, "A or B" can mean "only A", "only B", or "both A and B". In other words, in this document, "A or B" can be interpreted as "A and / or B". For example, in this document, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".

[0045] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".

[0046] In this document, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this document, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0047] Also, in this document, "at least one of A, B and C" can mean "only A", "only B", "only C", or "any combination of A, B and C". Also, "at least one of A, B or C" and "at least one of A, B and / or C" can mean "at least one of A, B and C".

[0048] Also, the parentheses used in this document can mean "for example". Specifically, when it is shown as "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, "prediction" in this document is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". Also, when it is shown as "prediction (that is, intra prediction)", "intra prediction" can be proposed as an example of "prediction".

[0049] In this document, the technical features separately described within one drawing may be embodied separately or simultaneously.

[0050] FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding apparatus to which the embodiments of this document can be applied. Hereinafter, the encoding apparatus can include an image encoding apparatus and / or a video encoding apparatus.

[0051] As shown in FIG. 2, the encoding apparatus 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor (231). The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to the embodiments. Also, the memory 270 can include a DPB (Decoded Picture Buffer) and can also be configured by a digital storage medium. The above hardware components can further include the memory 270 as an internal / external component.

[0052] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied thereafter. Alternatively, the binary tree structure can also be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that a coding unit of an optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the final coding unit described above.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0053] The term "unit" can, in some cases, be used interchangeably with terms such as "block" or "area". In a general case, an M×N block can represent a set such as samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or can also represent only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).

[0054] The encoding device 200 can subtract a prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0055] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block or remotely located depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the level of detail of the prediction direction. However, this is an example, and a greater or lesser number of directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the adjacent blocks.

[0056] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The above motion information can include a motion vector and a reference picture index. The above motion information can further include inter prediction direction (L0 prediction, L1 prediction, BI prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the above reference block and the reference picture including the above temporal neighboring block can be the same or different. The above temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the above temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal may not be transmitted.In the case of the motion information prediction (Motion Vector Prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and by signaling the motion vector difference, the motion vector of the current block can be indicated.

[0057] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, for the prediction of one block, the prediction unit can apply not only intra prediction or inter prediction, but also apply intra prediction and inter prediction simultaneously. This can be called Combined Inter and Intra Prediction (CIIP). Also, the prediction unit can be based on the Intra Block Copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values in the picture can be signaled based on information regarding the palette table and the palette index.

[0058] The prediction signal generated through the above prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), and CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when expressing the relationship information between pixels in a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process may be applied to a pixel block having the same size of a square or may be applied to a block of variable size that is not square.

[0059] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the block-shaped quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding). In addition to the quantized transform coefficients, the entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (for example, values of syntax elements). The encoded information (for example, encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (Network Abstraction Layer) units. The video / image information can further include information regarding various parameter sets such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / image information. The video / image information can be encoded through the encoding procedure described above and included in the bitstream.The above bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured such that a transmitting unit (not shown) for transmitting and / or a storage unit (not shown) for storing are internal / external elements of the encoding device 200, or the transmitting unit can also be included in the entropy encoding unit 240.

[0060] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and as will be described later, can also be used for inter prediction of the next picture after passing through filtering.

[0061] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied during the picture encoding and / or restoration process.

[0062] The filtering unit 260 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240 as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0063] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve the encoding efficiency.

[0064] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 to be utilized as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0065] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which the embodiments of this document can be applied. Hereinafter, the decoding apparatus can include an image decoding apparatus and / or a video decoding apparatus.

[0066] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an intra predictor 331 and an inter predictor 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The above-described entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 can be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (Decoded Picture Buffer) and can be configured by a digital storage medium. The above hardware component can further include the memory 360 as an internal / external component.

[0067] When a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information was processed by the encoding device of FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block splitting related information obtained from the above bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be split from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a reproducing device.

[0068] The decoding device 300 can receive the signal output from the encoding device in FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the above bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The above video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the above video / image information can further include general constraint information. The decoding device can further decode the picture based on the information regarding the above parameter sets and / or the above general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the above decoding procedure and obtained from the above bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the information of the adjacent and decoded blocks of the block to be decoded or the symbol / bin information decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, performs arithmetic decoding of the bin, and can generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the symbol / bin information decoded for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction units (inter prediction unit 332 and intra prediction unit 331), and the residual values for which entropy decoding has been performed by the entropy decoding unit 310, that is, the quantized transform coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can also be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / image / picture decoding device, and the above decoding device can also be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The above information decoder can include the above entropy decoding unit 310, and the above sample decoder can include at least one of the above inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.

[0069] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the above reordering can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.

[0070] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).

[0071] The prediction unit can perform a prediction on the current block and generate a predicted block that includes a prediction sample for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0072] The prediction unit 330 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called Combined Inter and Intra Prediction (CIIP). Also, the prediction unit can be based on the Intra Block Copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / image information.

[0073] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to above can be located adjacent to the current block or at a distance therefrom depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to an adjacent block.

[0074] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between an adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, BI prediction, etc.) information. In the case of inter prediction, the adjacent block can include a spatial neighboring block existing within the current picture and a temporal neighboring block existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0075] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0076] The adder 340 can be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed within the current picture, and as will be described later, can be output after filtering, or can also be used for inter prediction of the next picture.

[0077] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied during the picture decoding process.

[0078] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.

[0079] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 so as to be utilized as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.

[0080] In this specification, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 200 can be applied to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300 in the same or corresponding manner, respectively.

[0081] As described above, when performing video coding, prediction is executed to improve the compression efficiency. Through this, a predicted block including prediction samples for the current block which is the block to be coded can be generated. Here, the predicted block includes prediction samples in the spatial domain (domain) (or pixel domain). The predicted block is also derived in the encoding device and the decoding device, and the encoding device can improve the image coding efficiency by signaling to the decoding device information (residual information) regarding the residual between the original block and the predicted block, which is not the original sample value of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, and can generate a restored block including restored samples by combining the residual block and the predicted block, and can generate a restored picture including the restored block.

[0082] The residual information can be generated through the conversion and quantization procedures. For example, the encoding device derives a residual block between the original block and the predicted block, executes a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, and executes a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, so that the relevant residual information can be signaled to the decoding device (via a bitstream). Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter-prediction of subsequent pictures to derive a residual block, and generate a restored picture based on this.

[0083] In this document, at least one of quantization / inverse quantization and / or conversion / inverse conversion can be omitted. When the quantization / inverse quantization is omitted, the quantized conversion coefficients can be called conversion coefficients. When the conversion / inverse conversion is omitted, the conversion coefficients can also be called coefficients or residual coefficients, or, for the sake of uniformity of expression, can still be called conversion coefficients.

[0084] In this document, the quantized transform coefficients and the transform coefficients can each be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information can include information regarding the transform coefficient(s), and the information regarding the transform coefficient(s) can be signaled via a residual coding syntax. The transform coefficient can be derived based on the residual information (or the information regarding the transform coefficient(s)), and the scaled transform coefficient can be derived via an inverse transform (scaling) for the transform coefficient. The residual sample can be derived based on an inverse transform (transformation) for the scaled transform coefficient. This can be applied / expressed similarly in other parts of this document.

[0085] Intra prediction can indicate a prediction that generates a prediction sample for a current block based on reference samples within a picture (hereinafter referred to as the current picture) to which the current block belongs. When intra prediction is applied to the current block, adjacent reference samples used for the intra prediction of the current block can be derived. The adjacent reference samples of the current block can include a total of 2×nH samples adjacent to the left boundary of the current block of size nW×nH and adjacent to the bottom-left, samples adjacent to the top boundary of the current block and a total of 2×nW samples adjacent to the top-right, and 1 sample adjacent to the top-left of the current block. Alternatively, the adjacent reference samples of the current block can also include a plurality of columns of upper adjacent samples and a plurality of rows of left adjacent samples. Also, the adjacent reference samples of the current block can include a total of nH samples adjacent to the right boundary of the current block of size nW×nH, a total of nW samples adjacent to the bottom boundary of the current block, and 1 sample adjacent to the bottom-right of the current block.

[0086] However, some of the adjacent reference samples of the current block may not yet be decoded or may not be available. In this case, the decoder can substitute the samples that are not available with available samples to form adjacent reference samples for use in prediction. Alternatively, adjacent reference samples for use in prediction can be formed through interpolation of the available samples.

[0087] When the adjacent reference samples are derived, (i) a predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can also be derived based on the reference samples that exist in a specific (predicted) direction with respect to the predicted sample among the adjacent reference samples of the current block. In the case of (i), it is called a non-directional mode or a non-angular mode, and in the case of (ii), it can be called a directional mode or an angular mode.

[0088] Also, among the above adjacent reference samples, based on the predicted sample of the current block, the predicted sample can also be generated through interpolation between a first adjacent sample located in the prediction direction of the intra prediction mode of the current block and a second adjacent sample located in the direction opposite to the prediction direction. In the case described above, it can be called Linear Interpolation Intra Prediction (LIP). Also, chroma predicted samples can be generated based on luma samples using a linear model. In this case, it can be called the LM mode.

[0089] Also, a temporary prediction sample of the current block is derived based on the filtered adjacent reference samples, and a weighted sum is performed on at least one reference sample derived by the intra prediction mode among the existing adjacent reference samples, i.e., the non-filtered adjacent reference samples, and the temporary prediction sample to derive a prediction sample of the current block. In the case described above, it can be called PDPC (Position Dependent Intra Prediction).

[0090] Also, the reference sample line with the highest prediction accuracy is selected from among the adjacent multiple reference sample lines of the current block, and a prediction sample is derived using the reference sample located in the prediction direction on the corresponding line, and the reference sample line used at this time is signaled to the decoding device to perform intra prediction coding. In the case described above, it can be called multi-reference line intra prediction or MRL-based intra prediction.

[0091] Also, the current block is divided into vertical or horizontal sub-partitions, and intra prediction is performed based on the same intra prediction mode, and adjacent reference samples can be derived and used in units of the sub-partitions. That is, in this case, the intra prediction mode for the current block is also applied to the sub-partitions, and by deriving and using adjacent reference samples in units of the sub-partitions, in some cases, the intra prediction performance can be improved. Such a prediction method can be called ISP (Intra Sub-Partitions)-based intra prediction.

[0092] The above intra prediction method can be called an intra prediction type, distinguished from the intra prediction mode. The above intra prediction type can be called by various terms such as an intra prediction technique or an additional intra prediction mode. For example, the above intra prediction type (or an additional intra prediction mode, etc.) can include at least one of the above-mentioned LIP, PDPC, MRL, and ISP. A general intra prediction method excluding specific intra prediction types such as the above LIP, PDPC, MRL, and ISP can be called a normal intra prediction type. The normal intra prediction type can be generally applied when the above specific intra prediction types are not applicable, and prediction can be performed based on the above-mentioned intra prediction mode. On the other hand, if necessary, post-processing filtering for the derived prediction sample can also be performed.

[0093] Specifically, the intra prediction procedure can include an intra prediction mode / type determination step, an adjacent reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. Also, if necessary, a post-processing filtering step for the derived prediction sample can also be performed.

[0094] When intra prediction is applied, the intra prediction mode applied to the current block can be determined using the intra prediction modes of adjacent blocks. For example, the decoding device can select one of the MPM (Most Probable Mode) candidates in the MPM list derived based on the intra prediction modes of the adjacent blocks (e.g., left and / or upper adjacent blocks) of the current block and additional candidate modes based on the received MPM index, or can select one of the remaining intra prediction modes not included in the above MPM candidates (and the planar mode) based on the remaining (remaining) intra prediction mode information. The above MPM list can be configured with or without including the planar mode as a candidate. For example, when the above MPM list includes the planar mode as a candidate, the above MPM list can have 6 candidates, and when the above MPM list does not include the planar mode as a candidate, the above MPM list can have 5 candidates. When the above MPM list does not include the planar mode as a candidate, a not planar flag (e.g., intra_luma_not_planar_flag) indicating that the intra prediction mode of the current block is not the planar mode can be signaled. For example, the MPM flag can be signaled first, and the MPM index and the not planar flag can be signaled when the value of the MPM flag is 1. Also, the above MPM index can be signaled when the value of the above not planar flag is 1. Here, the configuration such that the above MPM list does not include the planar mode as a candidate is for signaling a flag (not planar flag) first to confirm whether it is the planar mode first because the planar mode is always considered as an MPM rather than meaning that the above planar mode is not an MPM.

[0095] For example, whether the intra prediction mode applied to the current block is within the MPM candidates (and planar mode), or within the remaining modes, can be indicated based on an MPM flag (e.g., intra_luma_mpm_flag). A value of 1 for the MPM flag can indicate that the intra prediction mode for the current block is within the MPM candidates (and planar mode), and a value of 0 for the MPM flag can indicate that the intra prediction mode for the current block is not within the MPM candidates (and planar mode). A value of 0 for the above not planar flag (e.g., intra_luma_not_planar_flag) can indicate that the intra prediction mode for the current block is the planar mode, and a value of 1 for the not planar flag can indicate that the intra prediction mode for the current block is not the planar mode. The above MPM index can be signaled in the form of the mpm_idx or intra_luma_mpm_idx syntax element, and the above remaining intra prediction mode information can be signaled in the form of the rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the above remaining intra prediction mode information can index the remaining intra prediction modes not included in the above MPM candidates (and planar mode) among all intra prediction modes in ascending order of prediction mode numbers and point to one of them. The above intra prediction mode is the intra prediction mode for the luma component (samples). Hereinafter, the intra prediction mode information can include at least one of the above MPM flag (e.g., intra_luma_mpm_flag), the above not planar flag (e.g., intra_luma_not_planar_flag), the above MPM index (e.g., mpm_idx or intra_luma_mpm_idx), and the above remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list can be referred to by various terms such as the MPM candidate list, candModeList, etc.When MIP is applied to the current block, separate mpm flags (e.g., intra_mip_mpm_flag), mpm indices (e.g., intra_mip_mpm_idx), and the remaining intra prediction mode information (e.g., intra_mip_mpm_remainder) for MIP can be signaled, and the above not planar flag is not signaled.

[0096] That is, generally when it comes to block partitioning for an image, the current block to be coded and neighboring blocks tend to have similar image characteristics. Therefore, the current block and neighboring blocks are highly likely to have the same or similar intra prediction modes to each other. Thus, the encoder can utilize the intra prediction mode of neighboring blocks to encode the intra prediction mode of the current block.

[0097] For example, the encoder / decoder can construct a MPM (Most Probable Modes) list for the current block. The above MPM list can also be referred to as a MPM candidate list. Here, MPM can be defined as the mode used to improve coding efficiency by considering the similarity between the current block and neighboring blocks during intra prediction mode coding. As described above, the MPM list can be constructed to include the planar mode or exclude the planar mode. For example, when the MPM list includes the planar mode, the number of candidates in the MPM list is 6. And when the MPM list does not include the planar mode, the number of candidates in the MPM list is 5.

[0098] The encoder / decoder can construct a MPM list including 5 or 6 MPMs.

[0099] To construct the MPM list, three types of modes can be considered: Default intra modes, Neighbour intra modes, and Derived intra modes.

[0100] For the above Neighbour intra modes, two adjacent blocks, namely the left adjacent block and the upper adjacent block, can be considered.

[0101] As described above, when the MPM list is configured not to include the planar mode, the planar mode is excluded from the list, and the number of the above MPM list candidates can be set to 5.

[0102] Also, among the intra prediction modes, the non-directional mode (or non-angle mode) can include the average-based DC mode of the neighboring reference samples of the current block or the interpolation-based planar mode.

[0103] When inter prediction is applied, the prediction unit of the encoding device / decoding device can perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction can indicate a prediction derived in a way that depends on data elements (e.g., sample values, or motion information) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by the index of the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The above motion information can include a motion vector and an index of a reference picture. The above motion information can further include information on the inter prediction type (L0 prediction, L1 prediction, BI prediction, etc.). When inter prediction is applied, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the above reference block and the reference picture including the above temporal neighboring block may be the same or different. The above temporal neighboring block may be called by names such as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the above temporal neighboring block may also be called a collocated picture (colPic). For example, a candidate list of motion information can be configured based on the adjacent blocks of the current block, and a flag or index information indicating which candidate is selected (used) can be signaled in order to derive the motion vector and / or the index of the reference picture of the current block.Inter prediction is performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block may be the same as that of the selected adjacent block. In the case of the skip mode, unlike the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected adjacent block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0104] The above motion information can include L0 motion information and / or L1 motion information according to the inter prediction type (such as L0 prediction, L1 prediction, BI prediction, etc.). The motion vector in the L0 direction can be called the L0 motion vector or MVL0, and the motion vector in the L1 direction can be called the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be called L0 prediction, the prediction based on the L1 motion vector can be called L1 prediction, and the prediction based on both the above L0 motion vector and the above L1 motion vector can be called bi (Bi) prediction. Here, the L0 motion vector can indicate a motion vector related to the reference picture list L0 (L0), and the L1 motion vector can indicate a motion vector related to the reference picture list L1 (L1). The reference picture list L0 can include, as reference pictures, pictures that are earlier than the current picture in the output order, and the reference picture list L1 can include pictures that are later than the current picture in the output order. The above earlier pictures can be called forward (reference) pictures, and the above later pictures can be called backward (reference) pictures. The reference picture list L0 can further include, as reference pictures, pictures that are later than the current picture in the output order. In this case, the above earlier pictures can be indexed first within the reference picture list L0, and the above later pictures can be indexed thereafter. The reference picture list L1 can further include, as reference pictures, pictures that are earlier than the current picture in the output order. In this case, the above later pictures can be indexed first within the reference picture list 1, and the above earlier pictures can be indexed thereafter. Here, the output order can correspond to the POC (Picture Order Count) order (order).

[0105] The video / image encoding procedure based on inter prediction generally includes, for example, the following.

[0106] FIG. 4 shows an example of an inter prediction-based video / image encoding method.

[0107] The encoding device performs inter prediction on the current block (S400). The encoding device derives the inter prediction mode and motion information of the current block, and generates prediction samples for the above-mentioned block. Here, the procedures of inter prediction mode determination, motion information derivation, and prediction sample generation may be performed simultaneously, or one procedure may be performed prior to the other procedures. For example, the inter prediction unit of the encoding device includes a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block, the motion information derivation unit derives the motion information of the current block, and the prediction sample derivation unit derives the prediction samples of the current block. For example, the inter prediction unit of the encoding device searches for a block similar to the current block within a certain area (search area) of the reference picture by motion estimation, and derives a reference block whose difference from the current block is the smallest or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device determines the mode to be applied to the current block among various prediction modes. The encoding device can compare the RD costs for the various prediction modes and determine the optimal prediction mode for the current block.

[0108] For example, when the skip mode or the merge mode is applied to the current block, the encoding device constructs a merge candidate list described later, and among the reference blocks pointed to by the merge candidates included in the merge candidate list, a reference block whose difference from the current block is the smallest or below a certain criterion can be derived. In this case, the merge candidate related to the derived reference block is selected, and merge index information indicating the selected merge candidate is generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0109] As another example, when the (A)MVP mode is applied to the current block, the encoding device constructs an (A)MVP candidate list to be described later, and among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list, the motion vector of the selected mvp candidate can be used as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the above-mentioned motion estimation can be used as the motion vector of the current block, and among the mvp candidates, the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block can be the selected mvp candidate. An MVD (Motion Vector Difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In that case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index is composed of reference picture index information and is separately signaled to the decoding device.

[0110] The encoding device derives a residual sample based on the prediction sample (S410). The encoding device can derive the residual sample by comparing the original sample of the current block with the prediction sample.

[0111] The encoding device encodes image information including prediction information and residual information (S420). The encoding device outputs the encoded image information in the form of a bitstream. The prediction information is information related to the prediction procedure and includes prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information related to motion information. The information related to the motion information includes candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. Also, the information related to the motion information includes information related to the aforementioned MVD and / or reference picture index information. Also, the information related to the motion information includes information indicating whether L0 prediction, L1 prediction, or bi prediction is applied (The information on the motion information may include information indicating whether L0 prediction, L1 prediction, or bi prediction is applied). The residual information is information related to the residual samples. The residual information includes information related to the quantized transform coefficients for the residual samples.

[0112] The output bitstream may be stored in a (digital) storage medium and transmitted to the decoding device, or may be transmitted to the decoding device via a network.

[0113] On the one hand, as described above, the encoding device generates a restored picture (including restored samples and restored blocks) based on the above reference samples and the above residual samples. This is to derive the same prediction result from the encoding device as that performed by the decoding device, thereby improving the coding efficiency. Therefore, the encoding device can store the restored picture (or restored samples, restored blocks) in the memory and utilize it as a reference picture for inter prediction. As described above, loop filtering procedures and the like can be further applied to the above restored picture.

[0114] Video / image decoding procedures based on inter prediction generally include, for example, the following.

[0115] FIG. 5 shows an example of a video / image decoding method based on inter prediction.

[0116] As shown in FIG. 5, the decoding device performs operations corresponding to the operations performed by the above encoding device. The decoding device can perform prediction on the current block based on the received prediction information and derive prediction samples.

[0117] Specifically, the decoding device determines a prediction mode for the current block based on the received prediction information (S500). The decoding device can determine which inter prediction mode is applicable to the current block based on the prediction mode information in the above prediction information.

[0118] For example, based on the above merge flag, it can be determined whether the merge mode is applicable to the current block, or whether the (A)MVP mode is determined. Alternatively, any one of various inter prediction mode candidates can be selected based on the above mode index. The inter prediction mode candidates include the skip mode, the merge mode, and / or the (A)MVP mode, or various inter prediction modes to be described later.

[0119] The decoding device derives the motion information of the current block based on the determined inter prediction mode (S510). For example, when the skip mode or the merge mode is applied to the current block, the decoding device constructs a merge candidate list described later and selects any one of the merge candidates included in the merge candidate list. The selection is performed based on the aforementioned selection information (merge index). The motion information of the current block can be derived using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0120] As another example, when the (A)MVP mode is applied to the current block, the decoding device constructs an (A)MVP candidate list described later, and the motion vector of the selected mvp (motion vector predictor) candidate included in the (A)MVP candidate list can be used as the mvp of the current block. The selection is performed based on the aforementioned selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the mvp of the current block and the MVD. Also, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index in the reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.

[0121] On the other hand, as will be described later, the motion information of the current block can be derived without constructing a candidate list. In this case, the motion information of the current block can be derived according to the procedure disclosed in the prediction mode described later. In this case, the construction of the candidate list as described above may be omitted.

[0122] The decoding device generates a prediction sample for the current block based on the motion information of the current block (S520). In this case, the reference picture is derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the samples of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, in some cases, a prediction sample filtering procedure may be further performed on all or part of the prediction samples of the current block.

[0123] For example, the inter prediction unit of the decoding device includes a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode for the current block is determined based on the prediction mode information received by the prediction mode determination unit, the motion information of the current block (such as a motion vector and / or a reference picture index, etc.) is derived based on the information related to the motion information received by the motion information derivation unit, and the prediction sample of the current block can be derived from the prediction sample derivation unit.

[0124] The decoding device generates a residual sample for the current block based on the received residual information (S530). The decoding device generates a restored sample for the current block based on the prediction sample and the residual sample, and generates a restored picture based on this (S540). As described above, an in-loop filtering procedure or the like can be further applied to the restored picture hereafter.

[0125] FIG. 6 exemplarily shows an inter prediction procedure.

[0126] Referring to FIG. 6, as described above, the inter prediction procedure includes an inter prediction mode determination step, a step of deriving motion information according to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. As described above, the above inter prediction procedure is performed in an encoding device and a decoding device. In this document, the coding device includes an encoding device and / or a decoding device.

[0127] As shown in FIG. 6, the coding device determines an inter prediction mode for the current block (S600). Various inter prediction modes can be used for predicting the current block in the picture. For example, various modes such as a merge mode, a skip mode, an MVP (Motion Vector Prediction) mode, an Affine mode, a sub-block merge mode, an MMVD (Merge with MVD) mode, etc. can be used. DMVR (Decoder side Motion Vector Refinement) mode, AMVR (Adaptive Motion Vector Resolution) mode, Bi-prediction with CU-level Weight (BCW), Bi-Directional Optical Flow (BDOF), etc. can be used additionally or alternatively as accompanying modes. The Affine mode may be called an affine motion prediction mode. The MVP mode may be called an "AMVP (Advanced Motion Vector Prediction) mode. In this document, some modes and / or motion information candidates derived by some modes may be included as one of the motion information related candidates of other modes. For example, the HMVP candidate may be added as a merge candidate for the above merge / skip mode, or may be added as an mvp candidate for the above MVP mode. When the above HMVP candidate is used as a motion information candidate for the above merge mode or skip mode, the above HMVP candidate may be called an HMVP merge candidate.

[0128] Prediction mode information indicating an inter prediction mode of a current block can be signaled from an encoding device to a decoding device. The prediction mode information can be included in a bitstream and received by the decoding device. The prediction mode information includes index information indicating one of a number of candidate modes. Alternatively, the inter prediction mode can be indicated via hierarchical signaling of flag information. In this case, the prediction mode information includes one or more flags. For example, a skip flag is signaled to indicate whether the skip mode is applied, and when the skip mode is not applied, a merge flag is signaled to indicate whether the merge mode is applied, and when the merge mode is not applied, it can be indicated that the MVP mode is applied, or a flag for additional classification can be further signaled. The affine mode may be signaled as an independent mode, or may be signaled as a mode subordinate to the merge mode or the MVP mode, etc. For example, the affine mode includes an affine merge mode and an affine MVP mode.

[0129] On the one hand, information indicating whether the aforementioned List0 (L0) prediction, List1 (L1) prediction, or bi-prediction is used for the current block (current coding unit) can be signaled. The above information may be referred to as motion prediction direction information, inter prediction direction information, or inter prediction indication information, and can be configured / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether the aforementioned List0 (L0) prediction, List1 (L1) prediction, or bi-prediction is used for the current block (current coding unit). In this document, for convenience of explanation, the inter prediction type (L0 prediction, L1 prediction, or BI prediction) indicated by the inter_pred_idc syntax element may be displayed as the motion prediction direction. L0 prediction may be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, the following prediction types can be indicated by the value of the inter_pred_idc syntax element.

[0130]

Table 1

[0131] As described above, one picture includes one or more slices. A slice can have one of the slice types including an I (Intra) slice, a P (Predictive) slice, and a B (Bi-Predictive) slice. The above slice type is indicated based on the slice type information. For blocks within an I slice, only intra prediction is used without using inter prediction for prediction. Of course, in this case, it is also possible to code and signal the original sample value without prediction. For blocks within a P slice, either intra prediction or inter prediction is used, and when inter prediction is used, only uni prediction can be used. On the other hand, for blocks within a B slice, either intra prediction or inter prediction is used, and when inter prediction is used, up to maximum bi prediction can be used.

[0132] L0 and L1 include reference pictures encoded / decoded before the current picture. For example, L0 includes reference pictures before and / or after the current picture in the POC order, and L1 includes reference pictures after and / or before the current picture in the POC order. In this case, a relatively lower reference picture index is assigned to L0 with respect to the reference picture before the current picture in the POC order, and a relatively lower reference picture index is assigned to L1 with respect to the reference picture after the current picture in the POC order. In the case of a B slice, bi prediction is applied, and in this case, either uni-directional bi prediction or bi-directional bi prediction may be applied. Bi-directional bi prediction is also called true bi prediction.

[0133] The coding device derives motion information for the current block (S610). The derivation of the motion information can be derived based on the inter prediction mode.

[0134] The coding device can perform inter prediction using the motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can search for highly correlated similar reference blocks within a determined search range in the reference picture in fractional pixel units using the original block in the original picture for the current block, and thereby derive motion information. The similarity of the blocks can be derived based on the difference in phase-based sample values. For example, the similarity of the blocks can be calculated based on the SAD (Sum of Absolute Differences) between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, the motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information is signaled to the decoding device in various ways based on the inter prediction mode.

[0135] The coding device performs inter prediction based on the motion information for the above-mentioned current block (S620). The coding device can derive one or more predicted samples for the current block based on the motion information. The current block including the predicted sample(s) may be called a predicted block.

[0136] When the merge mode is applied, instead of directly transmitting the motion information of the current predicted block, the motion information of the current predicted block is derived using the motion information of the surrounding predicted blocks. Therefore, by transmitting flag information indicating the use of the merge mode and a merge index indicating which surrounding predicted block is used, the motion information of the current predicted block can be indicated. The above merge mode may be called a regular merge mode.

[0137] The encoder must search for merge candidate blocks used to derive the motion information of the current prediction block in order to perform the merge mode. For example, up to five of the above merge candidate blocks can be used, but the embodiments of this document are not limited to this. And the maximum number of the above merge candidate blocks is transmitted in the slice header or the tile group header. After finding the above merge candidate blocks, the encoder can generate a merge candidate list and select the merge candidate block with the minimum cost among them as the final merge candidate block.

[0138] For example, five merge candidate blocks can be used in the above merge candidate list. For example, four spatial merge candidates and one temporal merge candidate can be used. Hereinafter, the above spatial merge candidate or the spatial MVP candidate described later may be referred to as SMVP, and the above temporal merge candidate or the temporal MVP candidate described later may be referred to as TMVP.

[0139] Hereinafter, a method of constructing a merge candidate list will be described.

[0140] The coding device (encoder / decoder) inserts the spatial merge candidates derived by searching the spatial neighboring blocks of the current block into the merge candidate list. For example, the above spatial neighboring blocks include the lower left corner neighboring blocks, left neighboring blocks, upper right corner neighboring blocks, upper neighboring blocks, and upper left corner neighboring blocks of the current block. However, this is an example, and in addition to the above-mentioned spatial neighboring blocks, additional neighboring blocks such as right neighboring blocks, lower neighboring blocks, and lower right corner neighboring blocks can also be used as the above spatial neighboring blocks. The coding device can search the above spatial neighboring blocks based on the priority to detect available blocks, and derive the motion information of the detected blocks as the above spatial merge candidates.

[0141] The coding device inserts the time merge candidates derived by searching for the time neighboring blocks of the current block into the merge candidate list. The time neighboring blocks may be located on a reference picture that is a picture different from the current picture in which the current block is located. The reference picture on which the time neighboring blocks are located may be called a collocated picture or a col picture. The time neighboring blocks can be searched in the order of the peripheral blocks of the lower right corner and the lower right center block of the co-located block with respect to the current block on the col picture. On the other hand, when motion data compression is applied, specific motion information is stored as representative motion information for each fixed storage unit in the col picture. In this case, it is not necessary to store the motion information for all blocks within the fixed storage unit, and thus the effect of motion data compression can be obtained. In this case, the fixed storage unit may be predetermined, for example, in units of 16×16 samples or 8×8 samples, or size information regarding the fixed storage unit may be signaled from the encoder to the decoder. When motion data compression is applied, the motion information of the time neighboring blocks can be replaced with the representative motion information of the fixed storage unit in which the time neighboring blocks are located. That is, in this case, from the perspective of implementation, instead of the prediction block located at the coordinates of the time neighboring blocks, based on the coordinates (the upper left sample position (position)) of the time neighboring blocks, after arithmetic right shift by a certain value, the time merge candidate is derived based on the motion information of the prediction block covering the position after arithmetic left shift. For example, when the fixed storage unit is a 2n×2n sample unit, if the coordinates of the time neighboring blocks are (xTnb, yTnb), the motion information of the prediction block located at the modified position ((xTnb>>n)<<n), (yTnb>>n)<<n)) is used for the time merge candidate.Specifically, for example, when the above fixed memory unit is a 16×16 sample unit, if the coordinates of the above temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>4)<<4), (yTnb>>4)<<4)) is used for the above temporal merge candidate. Alternatively, for example, when the above fixed memory unit is an 8×8 sample unit, if the coordinates of the above temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>3)<<3), (yTnb>>3)<<3)) is used for the above temporal merge candidate.

[0142] The coding device can check whether the current merge candidate number (the number of current merge candidates) is less than the maximum merge candidate number (the number of maximum merge candidates). The above maximum merge candidate number can be predefined or signaled from the encoder to the decoder. For example, the encoder generates information regarding the above maximum merge candidate number, encodes it, and transmits it to the above decoder in the form of a bitstream. When the above maximum merge candidate number is filled, the subsequent candidate addition process may not be performed.

[0143] As a result of the above check, when the current merge candidate number is less than the maximum merge candidate number, the coding device inserts an additional merge candidate into the above merge candidate list.

[0144] As a result of the above check, when the current merge candidate number is not less than the maximum merge candidate number, the coding device terminates the configuration of the above merge candidate list. In this case, the encoder can select the optimal merge candidate among the merge candidates that make up the above merge candidate list based on the RD (Rate-Distortion) cost, and can signal selection information (for example, merge index) indicating the selected merge candidate to the decoder. The decoder selects the above optimal merge candidate based on the above merge candidate list and the above selection information.

[0145] As described above, the motion information of the selected merge candidate can be used as the motion information of the current block, and the predicted sample of the current block can be derived based on the motion information of the current block. The encoder can derive the residual sample of the current block based on the predicted sample and signal the residual information regarding the residual sample to the decoder. As described above, the decoder can generate a restored sample based on the residual sample derived based on the residual information and the predicted sample, and generate a restored picture based on this.

[0146] When the skip mode is applied, the motion information of the current block can be derived in the same way as when the aforementioned merge mode is applied. However, when the skip mode is applied, the residual signal for the corresponding block is omitted, and thus the predicted sample can be immediately used as the restored sample.

[0147] When the MVP mode is applied, a motion vector predictor (mvp) candidate list is generated using the motion vectors of the restored spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks (or Col blocks). That is, the motion vectors of the restored spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks can be used as motion vector predictor candidates. When dual prediction is applied, an mvp candidate list for L0 motion information derivation and an mvp candidate list for L1 motion information derivation can be individually generated and used. The aforementioned prediction information (or information related to prediction) includes selection information (e.g., an MVP flag or an MVP index) indicating the selected optimal motion vector predictor candidate among the motion vector predictor candidates included in the above list. Here, the prediction unit can use the above selection information to select the motion vector predictor of the current block from among the motion vector predictor candidates included in the motion vector candidate list. The prediction unit of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and encode this and output it in the form of a bitstream. That is, the MVD is obtained as the value obtained by subtracting the above motion vector predictor from the motion vector of the current block. Here, the prediction unit of the decoding device can obtain the motion vector difference included in the above prediction-related information, and derive the motion vector of the current block by adding the above motion vector difference and the above motion vector predictor. The prediction unit of the decoding device can obtain or derive from the above prediction-related information a reference picture index indicating a reference picture, etc.

[0148] FIG. 7 is a flowchart (sequence diagram) showing a method of constructing a motion vector predictor candidate list.

[0149] As shown in FIG. 7, in one embodiment, first, a spatial candidate block for motion vector prediction is searched and inserted into a prediction candidate list (S700). Thereafter, in one embodiment, it is determined whether the number of spatial candidate blocks is less than 2 (S710). For example, in one embodiment, if the number of spatial candidate blocks is less than 2, a temporal candidate block is searched and additionally inserted into the prediction candidate list (S720). If the temporal candidate block is unavailable, a zero motion vector is used. That is, a zero motion vector can be additionally inserted into the prediction candidate list (S730). Thereafter, in one embodiment, the configuration of the preliminary candidate list is terminated (S740). Alternatively, in one embodiment, if the number of spatial candidate blocks is not less than 2, the configuration of the preliminary candidate list is terminated (S740). Here, the preliminary candidate list indicates an MVP candidate list.

[0150] On the other hand, when the MVP mode is applied, a reference picture index is explicitly signaled. In this case, it can be signaled separately as a reference picture index (refidxL0) for L0 prediction and a reference picture index (refidxL1) for L1 prediction. For example, when the MVP mode is applied and bi-prediction is applied, information regarding the above refidxL0 and information regarding refidxL1 can both be signaled.

[0151] When the MVP mode is applied, as described above, information regarding the MVD derived from the encoding device is signaled to the decoding device. The information regarding the MVD can include, for example, information indicating the x and y components of the absolute value and sign of the MVD. In this case, information indicating whether the absolute value of the MVD is greater than 0 and whether it is greater than 1, and information indicating the remainder of the MVD can be signaled step by step. For example, information indicating whether the absolute value of the MVD is greater than 1 can be signaled only when the value of the flag information indicating whether the absolute value of the MVD is greater than 0 is 1.

[0152] For example, information regarding MVD is encoded in an encoding device with a syntax as shown in the following table and signaled to a decoding device.

[0153]

Table 2

[0154] For example, in Table 2, the abs_mvd_greater0_flag syntax element indicates information regarding whether the difference (MVD) is greater than 0, and the abs_mvd_greater1_flag syntax element indicates information regarding whether the difference (MVD) is greater than 1. Also, the abs_mvd_minus2 syntax element indicates information regarding the value obtained by subtracting 2 from the difference (MVD), and the mvd_sign_flag syntax element indicates information regarding the sign of the difference (MVD). Also, in Table 2, [0] of each syntax element indicates that it is information regarding L0, and [1] indicates that it is information regarding L1.

[0155] For example, MVD[compIdx] is derived based on abs_mvd_greater0_flag[compIdx] * (abs_mvd_minus2[compIdx] + 2) * (1 - 2 * mvd_sign_flag[compIdx]). Here, compIdx (or cpIdx) indicates the index of each component and can have a value of 0 or 1. compIdx0 indicates the x component, and compIdx1 indicates the y component. However, this is an example, and values can be represented for each component using other coordinate systems instead of the x, y coordinate system.

[0156] On one hand, the MVD for L0 prediction (MVD L0) and the MVD for L1 prediction (MVD L1) may be separately signaled, and the information regarding the MVD may include information regarding MVD L0 and / or information regarding MVD L1. For example, when the MVP mode is applied to the current block and the BI prediction is applied, both the information regarding MVD L0 and the information regarding MVD L1 are signaled.

[0157] FIG. 8 is a diagram for explaining SMVD (Symmetric Motion Vector Differences).

[0158] When the BI prediction is applied, SMVD (Symmetric MVD) may be used in consideration of coding efficiency. In this case, signaling of a part of the motion information may be omitted. For example, when SMVD is applied to the current block, the information regarding refidxL0, the information regarding refidxL1, and the information regarding MVD L1 can be internally derived without being signaled from the encoding device to the decoding device. For example, when the MVP mode and the BI prediction are applied to the current block, flag information (e.g., SMVD flag information or sym_mvd_flag syntax element) for indicating whether SMVD can be applied is signaled, and when the value of the flag information is 1, the decoding device determines that SMVD is applied to the current block.

[0159] When the SMVD mode is applied (i.e., when the value of the SMVD flag information is 1), information regarding mvp_l0_flag, mvp_l1_flag, and MVD L0 (Motion Vector Difference L0) is explicitly signaled, and as described above, signaling of information regarding refidxL0, information regarding refidx1, and information regarding MVD L1 (Motion Vector Difference L1) is omitted and can be derived internally. For example, refidxL0 can be derived as an index that points to the previous reference picture closest to the current picture in terms of the POC procedure within the reference picture list 0 (which may also be called List0 or L0). refidxL1 can be derived as an index that points to the subsequent reference picture closest to the current picture in terms of the POC procedure within the reference picture list 1 (which may also be called List1 or L1). Alternatively, for example, both refidxL0 and refidxL1 can be derived as 0 respectively. Alternatively, for example, the above refidxL0 and refidxL1 can be derived as the minimum indices having the same POC difference in relation to the current picture respectively. Specifically, for example, when "[POC of the current picture] - [POC of the first reference picture indicated by refidxL0]" is called the first POC difference and "[POC of the current picture] - [POC of the second reference picture indicated by refidxL1]" is called the second POC difference, the value of refidxL0 that points to the first reference picture is derived as the refidxL0 of the current block and the value of refidxL1 that points to the second reference picture is derived as the refidxL1 of the current block only when the first POC difference and the second POC difference are the same. Also, for example, when there are multiple sets where the first POC difference and the second POC difference are the same, refidxL0 and refidxL1 of the set with the smallest difference among them can be derived as the refidxL0 and refidxL1 of the current block.

[0160] As shown in FIG. 8, reference picture list 0, reference picture list 1, and MVD L0, MVD L1 are shown. Here, MVD L1 is symmetric to MVD L0.

[0161] MVD L1 can be derived as minus (−) MVD L0. For example, the final (improved or corrected) motion information (motion vector: MV) for the current block is derived based on the following formula.

[0162] <Equation 1>

Number

[0163] In Equation 1, mvx0 and mvy0 represent the x and y components of the motion vector for L0 motion information or L0 prediction, and mvx1 and mvy1 represent the x and y components of the motion vector for L1 motion information or L1 prediction. Also, mvpx0 and mvpy0 represent the x and y components of the motion vector predictor for L0 prediction, and mvpx1 and mvpy1 represent the x and y components of the motion vector predictor for L1 prediction. Also, mvdx0 and mvdy0 represent the x and y components of the motion vector difference for L0 prediction.

[0164] On the other hand, in the MMVD mode, as a method of applying MVD (Motion Vector Difference) to the merge mode, the motion information directly used for generating the prediction sample of the current block (i.e., the current CU) can be implicitly derived. For example, an MMVD flag (e.g., mmvd_flag) indicating whether to use MMVD for the current block (i.e., the current CU) is signaled, and MMVD can be performed based on this MMVD flag. When MMVD is applied to the current block (e.g., when mmvd_flag is 1), additional information regarding MMVD can be signaled.

[0165] Here, the additional information regarding MMVD includes a merge candidate flag (e.g., mmvd_cand_flag) indicating whether the first candidate or the second candidate in the merge candidate list is used together with the MVD, a distance index (e.g., mmvd_distance_idx) for indicating the magnitude of motion, and a direction index (mmvd_direction_idx) for indicating the motion direction.

[0166] In the MMVD mode, two candidates (i.e., the first candidate or the second candidate) located at the first and second entries among the candidates in the merge candidate list can be used, and either one of the two candidates (i.e., the first candidate or the second candidate) can be used as the base MV. For example, a merge candidate flag (e.g., mmvd_cand_flag) can be signaled to indicate either one of the two candidates (i.e., the first candidate or the second candidate) in the merge candidate list.

[0167] Also, the distance index (e.g., mmvd_distance_idx) indicates information on the magnitude of motion and can indicate a predetermined offset from the start point. The offset may be added to the horizontal or vertical component of the start motion vector. The relationship between the distance index and the predetermined offset can be shown as in the following table.

[0168]

Table 3

[0169] Referring to Table 3 above, the distance of the MVD (e.g., MmvdDistance) is determined by the value of the distance index (e.g., mmvd_distance_idx), and the distance of the MVD (e.g., MmvdDistance) can be derived using integer sample precision or fractional sample precision based on the value of tile_group_fpel_mmvd_enabled_flag. For example, when tile_group_fpel_mmvd_enabled_flag is 1, it indicates that the distance of the MVD is derived using integer sample precision in the current tile group (or picture header), and when tile_group_fpel_mmvd_enabled_flag is 0, it indicates that the distance of the MVD is derived using fractional sample precision in the tile group (or picture header). In Table 1, the information (flag) for the tile group can be replaced with the information for the picture header. For example, tile_group_fpel_mmvd_enabled_flag can be replaced with ph_fpel_mmvd_enabled_flag (or ph_mmvd_fullpel_only_flag).

[0170] Also, the direction index (e.g., mmvd_direction_idx) indicates the direction of the MVD with respect to the starting point and indicates four directions as shown in Table 4 below. Here, the direction of the MVD can indicate the sign of the MVD. The relationship between the direction index and the MVD sign is shown as follows in the table.

[0171]

Table 4

[0172] Referring to Table 4 above, the sign of the MVD (e.g., MmvdSign) is determined by the value of the direction index (e.g., mmvd_direction_idx), and the sign of the MVD (e.g., MmvdSign) is derived for the L0 reference picture and the L1 reference picture.

[0173] Based on the distance index (e.g., mmvd_distance_idx) and the direction index (e.g., mmvd_direction_idx) as described above, the offset of the MVD can be calculated as in the following formula.

[0174] <Equation 2>

Number

[0175] <Equation 3>

Number

[0176] In Equation 2 and Equation 3, the MMVD distance (MmvdDistance[x0][y0]) and the MMVD signs (MmvdSign[x0][y0][0], MmvdSign[x0][y0][1]) are derived based on Table 3 and / or Table 4. In summary, in the MMVD mode, among the merge candidates of the merge candidate list derived based on the neighboring blocks, the merge candidate indicated by the merge candidate flag (e.g., mmvd_cand_flag) is selected, and the selected merge candidate can be used as the base candidate (e.g., MVP). Then, the motion information (i.e., the motion vector) of the current block can be derived by adding the MVD derived using the distance index (e.g., mmvd_distance_idx) and the direction index (e.g., mmvd_direction_idx) based on the base candidate.

[0177] Based on the motion information derived by the prediction mode, a predicted block for the current block can be derived. The predicted block includes a predicted sample (predicted sample array) of the current block. When the motion vector of the current block points to a fractional sample unit, an interpolation procedure can be performed, whereby the predicted sample of the current block can be derived based on the reference samples of the fractional sample unit within the reference picture. When dual prediction is applied, the predicted sample derived by the weighted sum or weighted average (according to the phase) of the predicted sample derived based on L0 prediction (i.e., prediction using the reference picture and MVL0 within reference picture list L0) and the predicted sample derived based on L1 prediction (i.e., prediction using the reference picture and MVL1 within reference picture list L1) can be used as the predicted sample of the current block. When dual prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions with respect to the current picture (i.e., when it corresponds to bidirectional prediction while being dual prediction), this may be called true dual prediction.

[0178] As described above, it is possible to generate a restored sample and a restored picture based on the derived predicted sample, and then procedures such as in-loop filtering can be executed.

[0179] As described above, according to this document, when dual prediction is applied to the current block, a prediction sample can be derived based on a weighted average. Conventionally, the dual prediction signal (i.e., the dual prediction sample) has been derived by a simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, the dual prediction sample has been derived as an average of the L0 prediction sample based on the L0 reference picture and MVL0 and the L1 prediction sample based on the L1 reference picture and MVL1. However, according to this document, when dual prediction is applied, the dual prediction signal (dual prediction sample) can be derived by a weighted average of the L0 prediction signal and the L1 prediction signal as follows.

[0180] In the embodiments related to the aforementioned MMVD, a method considering a long-term reference picture in the MVD derivation process of MMVD can be proposed, thereby enabling the compression efficiency to be maintained and increased in various applications. Also, the method proposed in the embodiments of this document can be similarly applied not only to the MMVD technology used in MERGE but also to the symmetric MVD technology SMVD used in the inter mode (MVP mode).

[0181] FIG. 9 is a diagram for explaining a method of deriving a motion vector in inter prediction.

[0182] In an embodiment of this document, an MV derivation method considering a long-term reference picture is used in the process of motion vector scaling of a temporal motion candidate (temporal merge candidate, or temporal mvp candidate). The temporal motion candidate can correspond to mvCol (mvLXCol). The temporal motion candidate may be referred to as "TMVP".

[0183] The following table explains the definition of the long-term reference picture.

[0184]

Table 5

[0185] Referring to Table 5 above, when LongTermRefPic (aPic, aPb, refIdx, LX) is 1 (true), the corresponding reference picture is marked as being used for long - term reference. For example, a reference picture that is not marked as being used for long - term reference can be a reference picture that is marked as being used for short - term reference. In another example, a reference picture that is not marked as being used for long - term reference and is not marked as being unused can be a reference picture that is marked as being used for short - term reference. Hereinafter, a reference picture that is marked as being used for long - term reference may be referred to as a long - term reference picture, and a reference picture that is marked as being used for short - term reference may be referred to as a short - term reference picture.

[0186] The following table explains the derivation of TMVP (mvLXCol).

[0187]

Table 6

[0188] Referring to FIG. 9 and Table 6, when the reference picture type pointed to by the current picture (e.g., indicating whether it is a long - term reference picture (LTRP) or a short - term reference picture (STRP)) is not the same as the type of the collocated reference picture pointed to by the collocated picture, the temporal motion vector (mvLXCol) is not used. That is, when all are long - term reference pictures or all are short - term reference pictures, colMV is derived, and when there are other types, colMV is not derived. Also, when all are long - term reference pictures and when the POC difference between the current picture and the reference picture of the current picture is the same as the POC difference between the collocated picture and the reference picture of the collocated picture, the collocated motion vector without scaling can be used as it is. When it is a short - term reference picture and the POC differences are different, the motion vector of the scaled collocated block is used.

[0189] In the embodiments of this document, the MMVD used in the MERGE / SKIP mode signals, for one coding block, the base motion vector index, the distance index, and the direction index as information for deriving the MVD information. When performing uni - directional prediction, the MVD is derived from the motion information, and when performing bi - directional prediction, symmetric MVD information is generated using mirroring and scaling methods.

[0190] When performing bi - directional prediction, the MVD information for L0 or L1 is scaled to generate the MVD for L1 or L0, but when referring to a long - term reference picture, changes in the MVD derivation process are required.

[0191] Figure 10 shows the MVD derivation process of MMVD according to an embodiment of this document. The method shown in Figure 10 can be for blocks to which bidirectional prediction is applied.

[0192] Referring to Figure 10, when the distance to the L0 reference picture is the same as the distance to the L1 reference picture, the derived MmvdOffset can be used directly as the MVD. When the POC differences (the POC difference between the L0 reference picture and the current picture and the POC difference between the L1 reference picture and the current picture) are different, according to the POC difference and whether it is a long-term or short-term reference picture, it is possible to derive the MVD by scaling or simple mirroring (i.e., -1*MmvdOffset).

[0193] As an example, the method of deriving a symmetric MVD using MMVD for blocks to which bidirectional prediction is applied does not conform to blocks that use long-term reference pictures. In particular, when the reference picture types in each direction are different, it is difficult to expect performance improvement when using MMVD. Therefore, in the following figures and embodiments, examples are introduced in which it is realized that MMVD is not applied when the reference picture types of L0 and L1 are different.

[0194] Figure 11 shows the MVD derivation process of MMVD according to another embodiment of this document. The method shown in Figure 11 can be for blocks to which bidirectional prediction is applied.

[0195] Referring to FIG. 11, different MVD derivation methods are applied depending on whether the reference picture referred to by the current picture (or current slice, current block) is a LTRP (Long-Term Reference Picture) or a STRP (Short-Term Reference Picture). In one example, when the method of the embodiment according to FIG. 11 is applied, a part of the standard document according to this embodiment is described as shown in the following table.

[0196]

Table 7-1

[0197]

Table 7-2

[0198] FIG. 12 shows the MVD derivation process of MMVD according to another embodiment of this document. The method shown in FIG. 12 can be for blocks to which bidirectional prediction is applied.

[0199] Referring to FIG. 12, different MVD derivation methods are applied depending on whether the reference picture referred to by the current picture (or current slice, current block) is a LTRP (Long-Term Reference Picture) or a STRP (Short-Term Reference Picture). In one example, when the method of the embodiment according to FIG. 12 is applied, a part of the standard document according to this embodiment is described as shown in the following table.

[0200]

Table 8-1

[0201]

Table 8-2

[0202] In summary, the MVD derivation process of MMVD that does not derive MVD when the reference picture types in each direction are different is described.

[0203] In one embodiment according to this document, MVD is not derived in all cases of referring to long-term reference pictures. That is, when even one of the L0 and L1 reference pictures is a long-term reference picture, MVD is set to 0, and MVD can be derived only when there is a short-term reference picture. This will be specifically described in the following drawings and tables.

[0204] FIG. 13 shows the MVD derivation process of MMVD according to one embodiment of this document. The method shown in FIG. 13 can be for blocks to which bidirectional prediction is applied.

[0205] Referring to FIG. 13 above, based on the highest priority condition (RefPicL0!=LTRP && RefPicL1!=STRP), when the current picture (or current slice, current block) refers only to short-term reference pictures, MVD for MMVD can be derived. In one example, when the method of the embodiment according to FIG. 13 is applied, a part of the standard document according to this embodiment is described as follows in the following table.

[0206]

Table 9-1

[0207]

Table 9-2

[0208] In one embodiment according to this document, when the reference picture types in each direction are different and there is a short-term reference picture, MVD is derived, and when there is a long-term reference picture, MVD is derived to 0. This will be specifically described in the following drawings and tables.

[0209] FIG. 14 shows the MVD derivation process of MMVD according to an embodiment of this document. The method shown in FIG. 14 can be for blocks to which bidirectional prediction is applied.

[0210] Referring to FIG. 14 above, when the reference picture types in each direction are different, when referring to a reference picture close to the current picture (short-term reference picture), MmvdOffset is applied, and when referring to a reference picture far from the current picture (long-term reference picture), the MVD has a value of 0. Here, a picture close to the current picture can be regarded as having a short-term reference picture, but if the close picture is a long-term reference picture, mmvdOffset can be applied to the motion vector of the list pointing to the short-term reference picture.

[0211]

Table 10

[0212] For example, the four paragraphs included in Table 10 above can sequentially replace the bottom block (content) of the flowchart included in FIG. 14.

[0213] In one example, when the method of the embodiment according to FIG. 14 is applied, a part of the standard document according to this embodiment is described as follows in the following table.

[0214]

Table 11-1

[0215]

Table 11-2

[0216] The following table shows a comparison table between the embodiments included in this document.

[0217]

Table 12

[0218] Referring to Table 12, a comparison is shown between methods of applying an offset considering a reference picture type for MVD derivation of MMVD described in the embodiments according to FIGS. 10 to 14. In Table 12, Embodiment A relates to an existing MMVD, Embodiment B shows the embodiments according to FIGS. 10 to 12, Embodiment C shows the embodiment according to FIG. 13, and Embodiment D shows the embodiment according to FIG. 14.

[0219] That is, in the embodiments according to FIGS. 10, 11, and 12, a method of deriving an MVD only when the reference picture types in both directions are the same is described, and in the embodiment according to FIG. 13, a method of deriving an MVD only when both directions are short-term reference pictures is described. In the case of the embodiment according to FIG. 13, if it is a long-term reference picture for unidirectional prediction, the MVD is set to 0. Also, in the embodiment according to FIG. 14, a method of deriving an MVD in only one direction when the reference picture types in both directions are different is described. Such differences between the embodiments indicate various features of the technology described in this document, and it can be understood by those of ordinary skill in the technical field to which this document pertains that the effects to be achieved by the embodiments according to this document based on the above features can be realized.

[0220] In the embodiments according to this document, when the reference picture type is a long-term reference picture, it has a separate process. When including a long-term reference picture, since scaling or mirroring based on POC difference (POCDiff) has no impact on performance improvement, the MVD in the direction with a short-term reference picture is assigned an MmvdOffset value, and the MVD in the direction with a long-term reference picture is assigned a value of 0. In one example, when this embodiment is applied, a part of the standard document according to this embodiment is described as follows in the following table.

[0221]

Table 13-1

[0222]

Table 13-2

[0223] In another example, a part of Table 13 above can be replaced with the following table. Referring to Table 14, Offset is applied based on the reference picture type instead of POCDiff.

[0224]

Table 14

[0225] In still another example, a part of Table 13 above can be replaced with the following table. Referring to Table 15, regardless of the reference picture type, MmvdOffset can always be set to L0 and -MmvdOffset can be set to L1.

[0226]

Table 15

[0227] According to one embodiment of this document, similar to the MMVD used in the aforementioned MERGE mode, the SMVD in the inter mode can be performed. When performing bidirectional prediction, whether symmetric MVD derivation is possible is signaled from the encoding device to the decoding device. When the related flag (e.g., sym_mvd_flag) is true (or its value is 1), the second-direction MVD (e.g., MVD L1) is derived by mirroring the first-direction MVD (e.g., MVD L0). In this case, scaling for the first-direction MVD may not be performed.

[0228] The following table shows the syntax related to the decoding of the decoding unit.

[0229] [Table 16]

[0230] [Table 17]

[0231] Referring to Table 16 and Table 17 above, when inter_pred_idc == PRED_BI and the reference pictures of L0 and L1 are available (e.g., RefIdxSymL0 > -1 && RefIdxSymL1 > -1), sym_mvd_flag is signaled.

[0232] The following table shows the decoding procedure for the MMVD reference index by way of an example.

[0233] [Table 18]

[0234] Referring to Table 18, the procedure for deriving the availability of the reference pictures of L0 and L1 is described. That is, if there is a reference picture in the forward direction among the L0 reference pictures, the reference picture index closest to the current picture is set to RefIdxSymL0, and the corresponding value is set to the reference index of L0. Also, if there is a reference picture in the backward direction among the L1 reference pictures, the reference picture index closest to the current picture is set to RefIdxSymL1, and the corresponding value is set to the reference index of L1.

[0235] The following Table 19 shows the decoding procedure for the MMVD reference index according to another example.

[0236]

Table 19

[0237] Referring to Table 19, when the L0 or L1 reference picture types are different as in the embodiments described with FIGS. 10, 11, and 12, that is, when long-term reference pictures and short-term reference pictures are used, after deriving the reference index for SMVD to prevent SMVD, if the reference picture types of L0 and L1 are different, do not use SMVD (see the bottom paragraph of Table 19).

[0238] In one embodiment of this document, similar to the MMVD used in the merge mode, in the inter mode, SMVD can be applied. As in the embodiment described with FIG. 13, when long-term reference pictures are used, in order to prevent SMVD, long-term reference pictures can be excluded in the process of deriving the reference index for SMVD as shown in the following table.

[0239]

Table 20

[0240] The following table according to another example of this embodiment shows an example of processing such that the SMVD is not applied when using a long-term reference picture after the derivation of the reference picture index for the SMVD.

[0241]

Table 21

[0242] In one embodiment of this document, when the reference picture type of the current picture and the reference picture type of the collocated picture are different in the colMV derivation process of the TMVP, the motion vector MV is set to 0. However, since the derivation methods in the cases of MMVD and SMVD are different, this is made uniform.

[0243] Even when the reference picture type of the current picture is a long-term reference picture and the reference picture type of the collocated picture is a long-term reference picture, the motion vector directly uses the collocated motion vector value. However, in MMVD and SMVD, in this case, the MV is set to 0. Here, the TMVP also sets the MV to 0 without additional derivation.

[0244] Also, even if the reference picture types are different, there may be a long-term reference picture close to the current picture. Therefore, instead of setting the MV to 0 considering this, the colMV can be used as the MV without scaling.

[0245] FIG. 15 is a diagram for explaining the SMVD according to one embodiment of this document.

[0246] For the derivation of SMVD, a method as shown in FIG. 15 can be used. That is, SMVD can be derived based on STRP (Short-Term Reference Picture) and / or LTRP (Long-Term Reference Picture). When using the mirrored L0 MVD for L1 MVD, if the types of the reference pictures are different, inaccurate MVD may be derived. This is because the ratio of distances (the distance between reference picture 0 and the current picture and the distance between reference picture 1 and the current picture) becomes large, and the correlation degree of the motion vectors in each direction decreases.

[0247] According to one embodiment of this document, the usability of the reference picture is checked, and if the conditions are met, sym_mvd_flag can be parsed. If sym_mvd_flag is true, the MVD of L1 (MVDL1) can be derived as the mirrored MVDL0 (the MVD of L0).

[0248] The following table shows a part of the coding unit syntax according to this embodiment.

[0249]

Table 22

[0250] Based on Table 22, the derivation procedure of sym_mvd_flag according to this embodiment can be described.

[0251] In this embodiment, the reference picture index for SMVD (RefIdxSymLX with X = 0, 1) can be derived. RefIdxSymL0 can indicate the index of the nearest reference picture having a POC smaller than the POC of the current picture. RefIdxSymL1 can indicate the index of the nearest reference picture having a POC larger than the POC of the current picture.

[0252] The following table describes, in the form of a standard document, a method for deriving a reference picture index for SMVD according to this embodiment.

[0253] [Table 23]

[0254] The following table shows the comparison results between embodiments. By considering the reference picture type according to the embodiments included in Table 24, the accuracy of MVD in SMVD can be improved. In Table 24, MVD can indicate MVD 0 (MVD of L0).

[0255] [Table 24]

[0256] Referring to Table 24, Example P shows an existing method for deriving SMVD. In Example Q, SMVD may be limited when a mixed reference picture type (e.g., STRP / LTRP or LTRP / STRP) is used at L0 and L1. In Example R, SMVD may be limited when referring to a long-term reference picture (LTRP).

[0257] The following table describes, in the form of a standard document, a method for deriving a reference picture index for SMVD according to Example Q of Table 24.

[0258] [Table 25]

[0259] The following table describes, in the form of a standard document, a method for deriving a reference picture index for SMVD according to Example Q of Table 24.

[0260]

Table 26

[0261]

Table 27

[0262] Referring to Table 26 and / or Table 27, the SMVD may be restricted when referring to the Long-Term Reference Picture (LTRP). For example, referring to Table 26, the long-term reference picture can be excluded in the reference picture checking process. Thereby, other reference pictures (e.g., not the long-term reference picture) can be considered for SMVD. Referring to Table 27, the SMVD may not be executed when the closest reference picture from the current picture is the long-term reference picture. For example, even if the reference picture list includes short-term reference pictures, the SMVD may not be executed when the closest reference picture from the current picture is the long-term reference picture.

[0263] In an example according to an embodiment of the present document, when the POC distance of L0 is greater than or equal to the POC of L1 in the MMVD procedure, the L1 MVD can be derived as a scaled or mirrored L0 MVD. When the POC distance of L0 is smaller than the POC of L1 in the MMVD procedure, the L0 MVD can be derived as a scaled or mirrored L1 MVD in the MMVD procedure.

[0264] FIG. 16 is a flowchart showing a method for deriving MMVD according to an embodiment of the present document.

[0265] In one embodiment of this document, considering the POC difference and / or reference picture type, the MVD can be derived by MMVD. Referring to FIG. 16, currPocDiffLX can mean the difference between the POC of the current picture and the POC of the reference picture LX. CurrPocDiffL0 and currPocDiffL1 can be compared with each other, and the type of the reference picture can be checked ("refPicList0!= LTRP" or "refPicList1!= LTRP"). Considering the conditions, MmvdOffset (derived using mmvd_cand_flag, mmvd_distance_idx, and / or mmvd_direction_idx) can be assigned as the same value, mirrored value, or scaled value as mMvdLX.

[0266] The following table shows a part of the standard document according to this embodiment.

[0267]

Table 28-1

[0268]

Table 28-2

[0269] When the current picture refers to one or more long-term reference pictures (LTRP), a mirroring procedure considering the POC distance may not be necessary. This is because the mirrored MVD obtained from a reference picture at a very far distance compared to other MVDs is not effective in terms of accuracy. A solution to this is described below.

[0270] The following table shows the comparison results between examples.

[0271]

Table 29

[0272] Referring to Table 19, Example X shows the existing method for deriving MMVD. In Example Y, the MMVD procedure can be restricted when one or more long-term reference pictures are referred to the current block. That is, in Example Y, the procedure for comparing the POC distance for the long-term reference picture can be omitted. In Example Z, for all cases, the derivation procedure of MMVD can be restricted. That is, in Example Z, for all cases, the procedure for comparing the POC distance can be omitted. In Table 19, offset can be referred to as MmvdOffset.

[0273] FIG. 17 is a flowchart showing a method for deriving MMVD according to an embodiment of this document. The flowchart of FIG. 17 can show the method for deriving MMVD according to the aforementioned Example Y.

[0274] Referring to FIG. 17, when the reference picture type is a long-term reference picture, the condition for comparing the POC difference can be removed, and the anchor MVD used for the mirroring procedure can be fixed to the L0 MVD.

[0275] The following table describes in the form of a standard document the method for deriving MMVD according to Example Y of Table 19.

[0276]

Table 30-1

[0277]

Table 30-2

[0278] FIG. 18 is a flowchart showing a method for deriving MMVD according to an embodiment of this document. The flowchart of FIG. 18 can show the method for deriving MMVD according to the aforementioned Example Z.

[0279] Referring to FIG. 18, in Example Z, for all cases, the derivation procedure of MMVD can be restricted. For all cases, the condition for comparing POC differences can be removed, and the anchor MVD used for the mirroring or scaling procedure can be fixed to the L0 MVD.

[0280] The following table describes in the form of a standard document the method for deriving MMVD according to Example Z of Table 19.

[0281]

Table 31-1

[0282]

Table 31-2

[0283] Also, in an example of this embodiment, for all cases, the condition for comparing POC differences can be removed, and only the case of mirroring may be used. The following table describes in the form of a standard document the method for deriving MMVD in this example.

[0284]

Table 32

[0285] The following drawings are created to illustrate a specific example of this specification. Since the names of the specific devices and the names of the specific signals / messages / fields described in the drawings are presented by way of example, the technical features of this specification are not limited to the specific names used in the following drawings.

[0286] Figures 19 and 20 schematically show an example of a video / image encoding method and related components according to an embodiment of this document. The method disclosed in FIG. 19 can be executed by the encoding device disclosed in FIG. 2. Specifically, for example, S1900 to S1970 in FIG. 19 can be executed by the prediction unit 220 of the encoding device, and S1980 can be executed by the residual processing unit 230 of the encoding device. S1990 can be executed by the entropy encoding unit 240 of the encoding device. The method disclosed in FIG. 19 can include the embodiments described above in this document.

[0287] Referring to FIG. 19, the encoding device constructs a motion vector predictor candidate list for the current block (S1900). For example, the encoding device can perform inter prediction on the current block considering the RD (Rate Distortion) cost to generate a prediction sample for the current block. Alternatively, for example, the encoding device can determine the inter prediction mode used to generate the prediction sample for the current block and derive motion information. Here, the inter prediction mode can be the MVP (Motion Vector Prediction) mode, but is not limited thereto. Here, the MVP mode can also be called the AMVP (Advanced Motion Vector Prediction) mode.

[0288] The encoding device can derive optimal motion information for the current block through motion estimation. For example, the encoding device can search for a highly correlated similar reference block within a determined search range in the reference picture in pixel units of fractions using the original block in the original picture for the current block, and derive motion information through this.

[0289] The encoding device can construct a motion vector predictor candidate list in order to indicate the derived motion information using a motion vector predictor and / or a motion vector difference. For example, the encoding device can construct a motion vector predictor candidate list based on a spatial neighboring candidate block and / or a temporal neighboring candidate block. Alternatively, the encoding device can also further use a zero motion vector when constructing the motion vector predictor candidate list. For example, when dual prediction is applied to the current block, an L0 motion vector predictor candidate list for L0 prediction and an L1 motion vector predictor candidate list for L1 prediction can each be constructed.

[0290] The encoding device determines the motion vector predictor of the current block based on the motion vector predictor candidate list (S1910). For example, the encoding device can determine the motion vector predictor for the current block based on the above-derived motion information (or motion vector) among the motion vector predictor candidates in the motion vector predictor candidate list. Alternatively, the encoding device can determine, within the motion vector predictor candidate list, the motion vector predictor with the smallest difference from the above-derived motion information (or motion vector). For example, when dual prediction is applied to the current block, the L0 motion vector predictor for L0 prediction and the L1 motion vector predictor for L1 prediction can each be determined from the L0 motion vector predictor candidate list and the L1 motion vector predictor candidate list.

[0291] The encoding device generates selection information indicating the motion vector predictor among the motion vector predictor candidate lists (S1920). For example, the selection information may also be referred to as index information, and may also be referred to as an MVP flag or an MVP index. That is, the encoding device can generate information indicating the motion vector predictor used to indicate the motion vector of the current block within the motion vector predictor candidate list. For example, when dual prediction is applied to the current block, selection information regarding the L0 motion vector predictor and selection information regarding the L1 motion vector predictor can each be generated.

[0292] The encoding device determines the motion vector difference for the current block based on the motion vector predictor (S1930). For example, the encoding device can determine the motion vector difference based on the derived motion information (or motion vector) for the current block and the motion vector predictor. Alternatively, the encoding device can determine the motion vector difference based on the difference between the derived motion information (or motion vector) for the current block and the motion vector predictor. For example, when dual prediction is applied to the current block, an L0 motion vector difference and an L1 motion vector difference can each be determined. Here, the L0 motion vector difference can be denoted as MvdL0, and the L1 motion vector difference can also be denoted as MvdL1.

[0293] The encoding device derives the reference picture list for the current block (S1940). In one example, the reference picture list can include reference picture list 0 (or L0, reference picture list L0), and reference picture list 1 (or L1, reference picture list L1). For example, the encoding device can configure the reference picture list for each slice included in the current picture.

[0294] The encoding device derives the POC difference between each of the reference pictures included in the above reference picture list and the current picture (S1950). In one example, the POC difference between the current picture and the previous reference picture from the current picture may be greater than 0. In another example, the POC difference between the current picture and the next reference picture from the current picture may be less than 0. However, this is only an illustration.

[0295] The encoding device derives a symmetric motion vector difference reference index based on the above POC difference (S1960). The symmetric motion vector difference reference index can point to a reference picture for the application of SMVD. The symmetric motion vector difference reference index can include a reference index L0 (RefIdxSumL0) and a reference index L1 (RefIdxSumL1).

[0296] The encoding device can derive a motion vector based on the symmetric MVD and the above motion vector predictor. The motion information can include the above motion vector.

[0297] The encoding device generates a predicted sample based on the above motion vector difference and the above symmetric motion vector difference reference index (S1970). For example, the above predicted sample can be generated based on the block (or sample) indicated by the above motion vector among the blocks (or samples) in the above reference picture pointed to by the above reference picture index.

[0298] The encoding device generates prediction-related information including the above inter prediction mode (S1950). The above prediction-related information can include information regarding MMVD, information regarding SMVD, and the like.

[0299] The encoding device generates residual information based on the above prediction sample (S1980). Specifically, the encoding device can derive a residual sample based on the above prediction sample and the original sample. The encoding device can derive residual information based on the above residual sample.

[0300] The encoding device encodes prediction-related information including the above selection information and information related to the motion vector difference (S1990). The encoding device encodes image / video information including the above prediction-related information and the above residual information. The encoded image / video information can be output in the form of a bitstream. The above bitstream can be transmitted to the decoding device via a network or a (digital) storage medium. The above prediction-related information can include information related to MMVD, information related to SMVD, etc.

[0301] The above image / video information can include various information according to the embodiments of this document. For example, the above image / video information can include information disclosed in at least one of Tables 1 to 32 described above.

[0302] In one embodiment, bi-prediction can be applied to the above current block. The above motion vector difference can include an L0 motion vector difference for L0 prediction and an L1 motion vector difference for L1 prediction. The above symmetric motion vector difference reference index can include an L0 symmetric motion vector difference reference index for L0 prediction and an L1 symmetric motion vector difference reference index for L1 prediction. The above L0 symmetric motion vector difference reference index and the above L1 symmetric motion vector difference reference index can be derived based on the short-term reference pictures included in the above reference picture list.

[0303] In one embodiment, the absolute value of the L1 motion vector difference may be the same as the absolute value of the L0 motion vector difference. The sign of the L1 motion vector difference may be different from the sign of the L0 motion vector difference.

[0304] In one embodiment, the image information may include a symmetric motion vector difference flag having a value of 1.

[0305] In one embodiment, the reference picture list may include a reference picture list L0 for L0 prediction and a reference picture list L1 for L1 prediction. The POC difference may include a first POC difference between the L0 short-term reference picture included in the reference picture list L0 and the current picture, and a second POC difference between the L1 short-term reference picture included in the reference picture list L1 and the current picture. The L0 symmetric motion vector difference reference index may point to the L0 short-term reference picture. The L1 symmetric motion vector difference reference index may point to the L1 short-term reference picture. The L0 symmetric motion vector difference reference index may be derived based on the first POC difference. The L1 symmetric motion vector difference reference index may be derived based on the second POC difference.

[0306] In one embodiment, the first POC difference may be the same as the second POC difference.

[0307] In one embodiment, the reference picture list can include a reference picture list L0 for L0 prediction. The reference picture list L0 can include a first L0 short-term reference picture and a second L0 short-term reference picture. The POC difference can include a third POC difference between the first L0 short-term reference picture and the current picture, and a fourth POC difference between the second L0 short-term reference picture and the current picture. Based on a comparison between the third POC difference and the fourth POC difference, a reference index pointing to the first L0 short-term reference picture can be derived as the L0 symmetric motion vector difference reference index.

[0308] In one embodiment, when the third POC difference is smaller than the fourth POC difference, a reference index pointing to the first L0 short-term reference picture can be derived as the L0 symmetric motion vector difference reference index.

[0309] FIGS. 21 and 22 schematically show an example of an image / video decoding method and related components according to an embodiment of this document. The method disclosed in FIG. 21 can be executed by the decoding device disclosed in FIG. 3. Specifically, for example, S2100 in FIG. 21 can be executed by the entropy decoding unit 310 of the decoding device, and S2110 to S2170 can be executed by the prediction unit 330 of the decoding device. The method disclosed in FIG. 21 can include the embodiments described above in this document.

[0310] Referring to FIG. 21, the decoding device receives / acquires image / video information (S2100). The decoding device can receive / acquire the above image / video information via a bitstream. The above image / video information can include prediction-related information (including prediction mode information) and residual information. The above prediction-related information can include information related to MMVD, information related to SMVD, etc. Also, the above image / video information can include various information according to the embodiments of this document. For example, the above image / video information can include information disclosed in at least one of Tables 1 to 32 described above.

[0311] The decoding device derives an inter prediction mode for the current block based on the above prediction-related information (S2110). Here, the inter prediction mode can include the merge mode, the AMVP mode (mode using motion vector predictor candidates), MMVD, and SMVD described above.

[0312] The decoding device constructs a motion vector predictor candidate list for the current block based on the inter prediction mode information (S2120). For example, the decoding device can construct a motion vector predictor candidate list based on spatial neighboring candidate blocks and / or temporal neighboring candidate blocks. Alternatively, the decoding device can also further use a zero motion vector when constructing the motion vector predictor candidate list. The motion vector predictor candidate list constructed here can be the same as the motion vector predictor candidate list constructed by the encoding device. For example, when dual prediction is applied to the current block, an L0 motion vector predictor candidate list for L0 prediction and an L1 motion vector predictor candidate list for L1 prediction can be respectively constructed.

[0313] The decoding device derives the motion vector of the current block based on the motion vector predictor candidate list (S2130). For example, the decoding device can derive a motion vector predictor candidate for the current block within the motion vector predictor candidate list based on the above-described selection information, and can derive the motion information (or motion vector) of the current block based on the derived motion vector predictor candidate. Alternatively, the motion information (or motion vector) of the current block can be derived based on the above-derived motion vector predictor candidate and the motion vector difference derived based on the information regarding the above-described motion vector difference. For example, when dual prediction is applied to the current block, the L0 motion vector predictor for L0 prediction and the L1 motion vector predictor for L1 prediction can be derived from the L0 motion vector predictor candidate list and the L1 motion vector predictor candidate list respectively based on the selection information for L0 prediction and the selection information for L1 prediction. For example, when dual prediction is applied to the current block, the L0 motion vector difference and the L1 motion vector difference can be derived based on the information regarding the motion vector difference respectively. Here, the L0 motion vector difference can be represented by MvdL0, and the L1 motion vector difference can also be represented by MvdL1. Also, the motion vector of the current block can be derived by the L0 motion vector and the L1 motion vector respectively.

[0314] The decoding device derives the reference picture list for the above current block (S2140). In one example, the reference picture list can include reference picture list 0 (or L0, reference picture list L0) and reference picture list 1 (or L1, reference picture list L1). For example, the decoding device can configure the reference picture list for each slice included in the current picture.

[0315] The decoding device derives the POC difference between each of the reference pictures included in the above reference picture list and the current picture (S2150). In one example, the POC difference between the current picture and the previous reference picture from the current picture may be greater than 0. In another example, the POC difference between the current picture and the next reference picture from the current picture may be less than 0. However, this is only an illustration.

[0316] The decoding device derives motion information including a reference picture index for SMVD based on the above POC difference (S2160). The decoding device can derive a reference index for SMVD. The reference index for SMVD can point to a reference picture for the application of SMVD. The reference index for SMVD can include a reference index L0 (RefIdxSumL0) and a reference index L1 (RefIdxSumL1).

[0317] The decoding device generates a predicted sample based on the above motion information (S2170). The decoding device can generate the predicted sample based on the motion vector and the reference picture index included in the above motion information. For example, the predicted sample can be generated based on the block (or sample) indicated by the motion vector among the blocks (or samples) in the reference picture pointed to by the above reference picture index.

[0318] In one embodiment, the motion vector can be derived based on a motion vector predictor derived based on the motion vector predictor candidate list and a motion vector difference. Bi-prediction can be applied to the current block. The motion vector difference can include an L0 motion vector difference for L0 prediction and an L1 motion vector difference for L1 prediction. The symmetric motion vector difference reference index can include an L0 symmetric motion vector difference reference index for L0 prediction and an L1 symmetric motion vector difference reference index for L1 prediction. The L1 motion vector difference can be derived based on the L0 motion vector difference. The L0 symmetric motion vector difference reference index and the L1 symmetric motion vector difference reference index can be derived based on short-term reference pictures included in the reference picture list.

[0319] In one embodiment, the image information can include a symmetric motion vector difference flag. Based on the symmetric motion vector difference flag having a value of 1, the L0 symmetric motion vector difference reference index and the L1 symmetric motion vector difference reference index can be derived.

[0320] In one embodiment, the absolute value of the L1 motion vector difference can be the same as the absolute value of the L0 motion vector difference. The sign of the L1 motion vector difference can be different from the sign of the L0 motion vector difference.

[0321] In one embodiment, the reference picture list can include a reference picture list L0 for L0 prediction and a reference picture list L1 for L1 prediction. The POC difference can include a first POC difference between the L0 short-term reference picture included in the reference picture list L0 and the current picture, and a second POC difference between the L1 short-term reference picture included in the reference picture list L1 and the current picture. The L0 symmetric motion vector difference reference index can point to the L0 short-term reference picture. The L1 symmetric motion vector difference reference index can point to the L1 short-term reference picture. The L0 symmetric motion vector difference reference index can be derived based on the first POC difference. The L1 symmetric motion vector difference reference index can be derived based on the second POC difference.

[0322] In one embodiment, the first POC difference can be the same as the second POC difference.

[0323] In one embodiment, the reference picture list can include a reference picture list L0 for L0 prediction. The reference picture list L0 can include a first L0 short-term reference picture and a second L0 short-term reference picture. The POC difference can include a third POC difference between the first L0 short-term reference picture and the current picture, and a fourth POC difference between the second L0 short-term reference picture and the current picture. Based on a comparison between the third and fourth POC differences, a reference index pointing to the first L0 short-term reference picture can be derived as the L0 symmetric motion vector difference reference index.

[0324] In one embodiment, when the third POC difference is smaller than the fourth POC difference, a reference index pointing to the first L0 short-term reference picture can be derived as the L0 symmetric motion vector difference reference index.

[0325] In the foregoing embodiments, the method has been described based on a flowchart as a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps, and a certain step may occur in a different step and a different order from those described above, or simultaneously. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, and different steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of the embodiments of this document.

[0326] The method according to the embodiments of the foregoing document can be embodied in the form of software, and the encoding device and / or decoding device according to this document can be included in, for example, devices that perform image processing such as TVs, computers, smartphones, set-top boxes, and display devices.

[0327] In this document, when an embodiment is implemented in software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory may be inside or outside the processor and may be connected to the processor by various well-known means. The processor can include an ASIC (Application-Specific Integrated Circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (Read-Only Memory), a RAM (Random Access Memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored in a digital storage medium.

[0328] In addition, the decoding device and encoding device to which the embodiments of this document are applied may include a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT video (Over The Top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (Virtual Reality) device, an AR (Augmented (Argument) Reality) device, an image phone video device, a transportation means terminal (e.g., a vehicle (including an autonomous driving vehicle) terminal, an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and can be used to process video signals or data signals. For example, the OTT video (Over The Top video) device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.

[0329] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a recording medium readable by the computer. Multimedia data having a data structure according to the embodiments of this document can also be stored in a recording medium readable by the computer. The above-mentioned recording medium readable by the computer includes all types of storage devices and distributed storage devices in which data readable by the computer is stored. The above-mentioned recording medium readable by the computer can include, for example, Blu-ray Disc (BD), Universal Serial (Universal Serial) Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Further, the above-mentioned recording medium readable by the computer includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a recording medium readable by the computer or transmitted via a wired or wireless communication network.

[0330] In addition, the embodiments of this document can be embodied in a computer program product by program code, and the above program code can be executed by a computer according to the embodiments of this document. The above program code can be stored on a carrier readable by a computer.

[0331] FIG. 23 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0332] Referring to FIG. 23, the content streaming system to which the embodiments of this document are applied can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0333] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and serves to transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server may be omitted.

[0334] The bitstream can be generated by an encoding method or a method for generating a bitstream to which the embodiments of this document are applicable. The streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0335] The streaming server transmits multimedia data to a user device based on a user request via a web server. The web server serves as a medium to inform the user of what services are available. If the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include another control server. In this case, the control server serves to control commands / responses between each device within the content streaming system.

[0336] The streaming server can receive content from a media storage device (repository) and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0337] In the example of the above user device, there may be a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a navigation device, a slate PC, a tablet PC, an ULTRABOOK (registered trademark), a wearable device (for example, a smartwatch (watch-type terminal), a smart glass (glass-type terminal), an HMD (Head Mounted Display)), a digital TV, a desktop computer, a digital signature (signi), etc.

[0338] Each server in the above content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.

[0339] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and embodied as a device, and the technical features of the device claims in this specification can be combined and embodied as a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and embodied as a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and embodied as a method.

Claims

1. An image decoding method performed by a decoding device, comprising: obtaining image information including prediction related information from a bitstream; deriving an inter prediction mode based on the prediction related information; constructing a motion vector predictor candidate list for a current block based on the inter prediction mode; deriving a motion vector for the current block based on the motion vector predictor candidate list; deriving a reference picture list for the current block, the reference picture list including an L0 reference picture list and an L1 reference picture list; determining whether an L0 reference picture in the L0 reference picture list is a short-term reference picture; deriving a first picture order count (POC) difference between the L0 reference picture in the L0 reference picture list and a current picture; deriving an L0 symmetric motion vector differential reference index; determining whether an L1 reference picture in the L1 reference picture list is a short-term reference picture; deriving a second POC difference between the L1 reference picture in the L1 reference picture list and the current picture; deriving an L1 symmetric motion vector differential reference index; generating a prediction sample based on the motion vector, the L0 symmetric motion vector differential reference index, and the L1 symmetric motion vector differential reference index; the motion vector is derived based on a motion vector predictor derived based on the motion vector predictor candidate list and a motion vector differential; Bi-prediction is applied to the current block; the motion vector differentials include an L0 motion vector differential for an L0 prediction and an L1 motion vector differential for an L1 prediction; the L1 motion vector differential is derived based on the L0 motion vector differential; the L0 symmetric motion vector differential reference index is derived based on the first POC differential and a determination that the L0 reference picture in the L0 reference picture list is the short-term reference picture; the L0 symmetric motion vector differential reference index is derived by using a reference index for the L0 reference picture determined as the short-term reference picture rather than a long-term reference picture; the L1 symmetric motion vector differential reference index is derived based on the second POC differential and a determination that the L1 reference picture in the L1 reference picture list is the short-term reference picture; The method, wherein the L1 symmetric motion vector differential reference index is derived by using a reference index for the L1 reference picture determined as the short-term reference picture but not a long-term reference picture.

2. 1. An image encoding method performed by an encoding device, comprising: constructing a motion vector predictor candidate list for the current block; determining a motion vector predictor for the current block based on the motion vector predictor candidate list; generating selection information indicative of the motion vector predictor in the motion vector predictor candidate list; determining a motion vector differential for the current block based on the motion vector predictor; deriving a reference picture list including reference pictures, the reference picture list including an L0 reference picture list and an L1 reference picture list; determining whether an L0 reference picture in the L0 reference picture list is a short-term reference picture; deriving a first picture order count (POC) difference between the L0 reference picture in the L0 reference picture list and a current picture; deriving an L0 symmetric motion vector differential reference index; determining whether an L1 reference picture in the L1 reference picture list is a short-term reference picture; deriving a second POC difference between the L1 reference picture in the L1 reference picture list and the current picture; deriving an L1 symmetric motion vector differential reference index; generating a prediction sample based on the motion vector differential, the L0 symmetric motion vector differential reference index, and the L1 symmetric motion vector differential reference index; generating residual information based on the prediction samples; encoding image information including said residual information and prediction related information including said selection information and information regarding said motion vector differentials; Bi-prediction is applied to the current block; the motion vector differentials include an L0 motion vector differential for an L0 prediction and an L1 motion vector differential for an L1 prediction; the L0 symmetric motion vector differential reference index is derived based on the first POC differential and a determination that the L0 reference picture in the L0 reference picture list is the short-term reference picture; the L0 symmetric motion vector differential reference index is derived by using a reference index for the L0 reference picture determined as the short-term reference picture rather than a long-term reference picture; the L1 symmetric motion vector differential reference index is derived based on the second POC differential and a determination that the L1 reference picture in the L1 reference picture list is the short-term reference picture; The method, wherein the L1 symmetric motion vector differential reference index is derived by using a reference index for the L1 reference picture determined as the short-term reference picture but not a long-term reference picture.

3. A method for transmitting data relating to an image, comprising the steps of: obtaining a bitstream relating to the image, the bitstream comprising: constructing a motion vector predictor candidate list for the current block; determining a motion vector predictor for the current block based on the motion vector predictor candidate list; generating selection information indicative of the motion vector predictor in the motion vector predictor candidate list; determining a motion vector differential for the current block based on the motion vector predictor; deriving a reference picture list including reference pictures, the reference picture list including an L0 reference picture list and an L1 reference picture list; determining whether an L0 reference picture in the L0 reference picture list is a short-term reference picture; deriving a first picture order count (POC) difference between the L0 reference picture in the L0 reference picture list and a current picture; deriving an L0 symmetric motion vector differential reference index; determining whether an L1 reference picture in the L1 reference picture list is a short-term reference picture; deriving a second POC difference between the L1 reference picture in the L1 reference picture list and the current picture; deriving an L1 symmetric motion vector differential reference index; generating a prediction sample based on the motion vector differential, the L0 symmetric motion vector differential reference index, and the L1 symmetric motion vector differential reference index; generating residual information based on the prediction samples; encoding image information including the residual information and prediction related information including the selection information and information regarding the motion vector differential; transmitting the data including the bitstream; Bi-prediction is applied to the current block; the motion vector differentials include an L0 motion vector differential for an L0 prediction and an L1 motion vector differential for an L1 prediction; the L0 symmetric motion vector differential reference index is derived based on the first POC differential and a determination that the L0 reference picture in the L0 reference picture list is the short-term reference picture; the L0 symmetric motion vector differential reference index is derived by using a reference index for the L0 reference picture determined as the short-term reference picture rather than a long-term reference picture; the L1 symmetric motion vector differential reference index is derived based on the second POC differential and a determination that the L1 reference picture in the L1 reference picture list is the short-term reference picture; The method, wherein the L1 symmetric motion vector differential reference index is derived by using a reference index for the L1 reference picture determined as the short-term reference picture but not a long-term reference picture.

Citation Information

Patent Citations

  • Symmetric motion vector difference coding

    WO2020132272A1

  • Symmetric motion vector difference coding

    WO2020221256A1