History-based image coding method and apparatus

The method enhances image/video compression efficiency by using a history-based motion vector prediction buffer to improve inter prediction and motion vector prediction, addressing the challenges of high-resolution content compression and transmission.

JP2025089405AActive Publication Date: 2025-06-12LG ELECTRONICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025047080
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-10-04
Filing Date
2025-03-21
Publication Date
2025-06-12
Estimated Expiration
2039-10-04

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos requires a highly efficient image/video compression technology to effectively compress, transmit, store, and reproduce such content, while existing methods struggle with efficient inter prediction and motion vector prediction.

Method used

The proposed method involves deriving a history-based motion vector prediction (HMVP) buffer for current blocks, constructing a motion information candidate list based on HMVP candidates, and using this information to generate prediction samples and restored samples, with efficient initialization and management of the HMVP buffer.

Benefits of technology

This approach improves overall image/video compression efficiency, reduces data transmission requirements for residual processing, and supports parallel processing through effective HMVP buffer management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089405000001_ABST
    Figure 2025089405000001_ABST
Patent Text Reader

Abstract

To provide an image decoding method.SOLUTION: A method includes the steps of: deriving a HMVP buffer for a current block; configuring a motion information candidate list on the basis of a HMVP candidate included in the HMVP buffer; deriving motion information of the current block on the basis of the motion information candidate list; deriving the reference picture index of the current block on the basis of the motion information; deriving the motion vector of the current block on the basis of the motion information; generating prediction samples for the current block on the basis of the reference picture index and the motion vector; and generating reconstructed samples on the basis of the prediction samples. One or more tiles are included in a current picture. The current picture includes a plurality of tile rows and tile columns. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in the current picture.SELECTED DRAWING: Figure 27
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to image coding, and more particularly, to a history-based image coding method and apparatus therefor.

Background Art

[0002] Recently, the demand for high-resolution and high-quality images / videos such as 4K or 8K and above UHD (Ultra High Definition) images / videos has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using existing media such as wired and wireless broadband lines, or storing image / video data using existing storage media, the transmission cost and storage cost increase.

[0003] Also, recently, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) contents, and holograms have been increasing, and the broadcasting of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] Therefore, in order to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above, a highly efficient image / video compression technology is required.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technical problem of this document is to provide a method and apparatus for increasing image coding efficiency.

[0006] Another technical problem of this document is to provide an efficient inter prediction method and apparatus.

[0007] Another technical problem of this document is to provide a method and an apparatus for deriving a history-based motion vector.

[0008] Another technical problem of this document is to provide a method and an apparatus for efficiently deriving HMVP (History-based Motion Vector Prediction) candidates.

[0009] Another technical problem of this document is to provide a method and an apparatus for efficiently updating an HMVP buffer.

[0010] Another technical problem of this document is to provide a method and an apparatus for efficiently initializing an HMVP buffer.

Means for Solving the Problem

[0011] According to one embodiment of this document, an image decoding method performed by a decoding device is provided. The above method includes steps of deriving an HMVP (History-based Motion Vector Prediction) buffer for a current block, constructing a motion information candidate list based on HMVP candidates in the HMVP buffer, deriving motion information of the current block based on the motion information candidate list, generating a prediction sample for the current block based on the motion information, and generating a restored sample based on the prediction sample. In the current picture, one or more tiles exist, and the HMVP buffer is initialized with the first CTU of the CTU row having the current block in the current tile.

[0012] According to another embodiment of this document, a decoding device for performing image decoding is provided. The decoding device derives an HMVP (History-based Motion Vector Prediction) buffer for a current block, constructs a motion information candidate list based on the HMVP candidates in the HMVP buffer, derives the motion information of the current block based on the motion information candidate list, and generates a prediction sample for the current block based on the motion information, and includes a prediction unit and an addition unit that generates a restored sample based on the prediction sample. In the current picture, one or more tiles exist, and the HMVP buffer is initialized with the first CTU in the CTU row having the current block in the current tile.

[0013] According to still another embodiment of this document, an image encoding method performed by an encoding device is provided. The above method includes steps of deriving an HMVP (History-based Motion Vector Prediction) buffer for a current block, constructing a motion information candidate list based on the HMVP candidates in the HMVP buffer, deriving the motion information of the current block based on the motion information candidate list, generating a prediction sample for the current block based on the motion information, deriving a residual sample based on the prediction sample, and encoding image information having information about the residual sample. In the current picture, one or more tiles exist, and the HMVP buffer is initialized with the first CTU in the CTU row having the current block in the current tile.

[0014] According to still another embodiment of the present document, an encoding apparatus that performs image encoding is provided. The encoding apparatus derives an HMVP (History-based Motion Vector Prediction) buffer for a current block, constructs a motion information candidate list based on HMVP candidates in the HMVP buffer, derives the motion information of the current block based on the motion information candidate list, and generates a prediction sample for the current block based on the motion information. A prediction unit, a residual processing unit that derives a residual sample based on the prediction sample, and an entropy encoding unit that encodes image information having information about the residual sample. In the current picture, one or more tiles exist, and the HMVP buffer is initialized with the first CTU in the CTU row having the current block in the current tile.

[0015] According to still another embodiment of the present document, a digital storage medium storing image data having encoded image information generated by an image encoding method performed by an encoding apparatus is provided.

[0016] According to still another embodiment of the present document, a digital storage medium storing image data having encoded image information that causes a decoding apparatus to perform an image decoding method is provided.

Advantages of the Invention

[0017] According to one embodiment of the present document, the overall image / video compression efficiency can be improved.

[0018] According to one embodiment of the present document, the amount of data to be transmitted required for residual processing can be reduced through efficient inter prediction.

[0019] According to one embodiment of the present document, the HMVP buffer can be efficiently managed.

[0020] According to one embodiment of the present document, parallel processing can be supported through efficient HMVP buffer management.

[0021] According to one embodiment of this document, a motion vector for inter prediction can be efficiently derived.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

DETAILED DESCRIPTION OF THE INVENTION

[0023] The methods presented in this document can be subject to various modifications and can have various embodiments. Specific embodiments will be illustrated in the drawings and will be described in detail. The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the technical concept of the methods presented in this document. Singular expressions include the expression "at least one" unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is not precluded in advance.

[0024] On the one hand, each component in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions and the like. It does not mean that each component is realized by separate hardware or separate software. For example, among the components, two or more components can be combined to form one component, and one component can also be divided into multiple components. As long as the implementation forms in which the components are integrated and / or separated do not deviate from the essence of the method disclosed in this document, they are included in the disclosure scope of this document.

[0025] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (Versatile Video Coding) standard. Also, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the EVC (Essential Video Coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd Generation Of Audio Video Coding Standard), or the next-generation video / image coding standard (such as H.267 or H.268, etc.).

[0026] This document presents various embodiments for video / image coding. Unless otherwise mentioned, the above embodiments can also be executed in combination with each other.

[0027] In this document, a video can mean a collection of a series of images over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can contain one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can contain one or more tiles. A brick can represent a rectangular region of CTU rows within a tile in a picture. A tile can be partitioned into multiple bricks, each of which can be composed of one or more CTU rows within the tile. Also, a tile that is not partitioned into multiple bricks may be also referred to as a brick.A brick scan can show a specific sequential ordering of CTUs partitioning a picture, where the CTUs can be ordered consecutively in a CTU raster scan within a brick, bricks within a tile can be ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture can be ordered consecutively in a raster scan of the tiles of the picture (A brick scan is a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture). A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture (A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture). The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width that can be specified by syntax elements in the picture parameter set (The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set).The tile row is a rectangular region of CTUs having a width specified by syntax elements in the picture parameter set and a height equal to the height of the picture. A tile scan is a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that may be exclusively contained in a single NAL unit. A slice may consists of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile.In this document, tile groups and slices can be used interchangeably. For example, in this document, a tile group / tile group header is also referred to as a slice / slice header.

[0028] A pixel or pel can mean the smallest unit that makes up one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample generally indicates a pixel or the value of a pixel, and may indicate only the pixel / pixel value of the luma component, or may indicate only the pixel / pixel value of the chroma component.

[0029] A unit can indicate the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the corresponding region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0030] In this document, the terms “ / ” and “,” are interpreted as “and / or.” For example, “A / B” is interpreted as “A and / or B,” and “A, B” is interpreted as “A and / or B.” Additionally, “A / B / C” means “at least one of A, B, and / or C.” Also, “A, B, C” also means “at least one of A, B, and / or C.”

[0031] Further, in the document, the term “or” is interpreted as “and / or.” For example, the expression “A or B” can mean 1) only A, 2) only B, or 3) both A and B. In other words, the term “or” in this document can be interpreted to mean “additionally or alternatively.”

[0032] Hereinafter, embodiments of this document will be described in more detail with reference to the attached drawings. Hereinafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions for the same components can be omitted.

[0033] FIG. 1 schematically shows an example of a video / image coding system to which an embodiment of this document can be applied.

[0034] As shown in FIG. 1, the video / image coding system can include a first device (source device) and a second device (receiver device). The source device can transmit encoded video / image information or data to the receiver device in a file or streaming form via a digital storage medium or a network.

[0035] The source device can include a video source, an encoding device, and a transmitting unit. The receiver device can include a receiving unit, a decoding device, and a renderer. The encoding device is also called a video / image encoding device, and the decoding device is also called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.

[0036] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., and in this case, it can replace the video / image capture process as the process of generating related data.

[0037] The encoding device can encode the input video / image. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0038] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the above bitstream and transmit it to the decoding device.

[0039] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0040] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit (display).

[0041] FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding apparatus to which an embodiment of this document can be applied. Hereinafter, the video encoding apparatus can include an image encoding apparatus.

[0042] As shown in FIG. 2, the encoding apparatus 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filtering unit 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 is also called a reconstructor or a reconstructed block generator. The aforementioned image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to the embodiment. Also, the memory 270 can include a DPB (Decoded Picture Buffer) and can also be configured by a digital storage medium. The above hardware components can further include the memory 270 as an internal / external component.

[0043] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit is also called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary tree. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary tree can be applied. Alternatively, the binary-tree structure may be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to this document can be executed. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with an optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the aforementioned final coding unit.The prediction unit is a unit of sample prediction, or the conversion unit is a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0044] The term "unit" can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to one picture (or image) for a pixel or a pel.

[0045] The encoding device 200 can subtract the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 200 is also called the subtraction unit 231. The prediction unit can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including the predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0046] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The non - directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is just an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0047] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The above motion information can include a motion vector and a reference picture index. The above motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the above reference block and the reference picture including the above temporal neighboring block may be the same or different. The above temporal neighboring block may also be called by names such as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the above temporal neighboring block is also called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be executed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal is not transmitted.In the case of the motion information prediction (Motion Vector Prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and by signaling the motion vector difference, the motion vector of the current block can be indicated.

[0048] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This is also called Combined Inter and Intra Prediction (CIIP). Also, the prediction unit may be based on the Intra Block Copy (IBC) prediction mode for the prediction of a block, or may be based on the palette mode. The above IBC prediction mode or palette mode can be used for content image / video coding such as games, for example, like SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be executed in a way similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can also be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index.

[0049] The prediction signal generated through the above prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when the relationship information between pixels is represented by a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square or can be applied to a block of a variable size that is not square.

[0050] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients is also referred to as residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (Network Abstraction Layer) units. The video / image information can further include information regarding various parameter sets such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / image information. The video / image information can be encoded through the above-described encoding procedure and included in the above bitstream.The above bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal can be configured as internal / external elements of the encoding device 200, or the transmission unit can also be included in the entropy encoding unit 240.

[0051] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222 to the restored residual signal. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 is also called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can also be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0052] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied during the picture encoding and / or restoration process.

[0053] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information regarding the filtering and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The information regarding the filtering can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0054] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches in the encoding device 200 and the decoding device, and can also improve the coding efficiency.

[0055] The memory 270 DPB can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the block from which the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0056] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which the embodiments of this document can be applied.

[0057] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-predictor 331 and an intra-predictor 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above can be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (Decoded Picture Buffer) and can also be configured by a digital storage medium. The above hardware component can further include the memory 360 as an internal / external component.

[0058] When a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information is processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the above bitstream. The decoding device 300 can execute decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a reproducing device.

[0059] The decoding device 300 can receive the signal output from the encoding device in FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the above bitstream to derive information (such as video / image information) necessary for image restoration (or picture restoration). The above video / image information can further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the above video / image information can further include general constraint information. Also, the decoding device can decode a picture based on the information regarding the above parameter set and / or the above general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the above decoding procedure and obtained from the above bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient for the residual. More specifically, the CABAC entropy decoding method receives the BIN corresponding to each syntax element in the bitstream, determines a context model using the decoding target syntax element information adjacent to the decoding target block and the decoding information of the decoding target block or the information of the symbol / BIN decoded in the previous step, predicts the occurrence probability of the BIN based on the determined context model, and executes arithmetic decoding of the BIN to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / BIN for the next symbol / BIN context model after determining the context model. Among the information related to prediction in the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value for which entropy decoding is performed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, the information related to filtering among the information decoded by the entropy decoding unit 310 can be provided to the filtering unit 350. On the other hand, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document is also called a video / image / picture decoding device, and the above decoding device can also be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The above information decoder can include the above entropy decoding unit 310, and the above sample decoder can include at least one of the above inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.

[0060] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the above reordering can be performed based on the coefficient scan order executed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (for example, quantization step size information) to obtain transform coefficients.

[0061] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).

[0062] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0063] The prediction unit 330 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This is also called Combined Inter and intra Prediction (CIIP). In addition, the prediction unit may be based on the Intra Block Copy (IBC) prediction mode for the prediction of a block, or may be based on the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be executed in a way similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can also be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / image information.

[0064] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0065] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The above motion information can include a motion vector and a reference picture index. The above motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0066] The addition unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331) to the obtained residual signal. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0067] The addition unit 340 is also called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block within the current picture, can be output after filtering as described later, or can be used for inter prediction of the next picture.

[0068] On the other hand, in the picture decoding process, LMCS (Luma Mapping with Chroma Scaling) can also be applied.

[0069] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be sent to the memory 360, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0070] The (modified) restored picture stored in the DPB of the memory 360 can be used as a reference picture by the inter prediction unit 332. The memory 360 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 331.

[0071] In this specification, the embodiments described in the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can be applied to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300 in the same or corresponding manner, respectively.

[0072] As described above, prediction is performed to increase the compression efficiency when performing video coding. Thereby, a predicted block including prediction samples for the current block which is a block to be coded can be generated. Here, the predicted block includes prediction samples in a spatial domain (domain) (or pixel domain). The predicted block is derived identically in an encoding device and a decoding device, and the encoding device can increase the image coding efficiency by signaling information (residual information) regarding a residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and can generate a restored block including restored samples by combining the residual block and the predicted block, and can generate a restored picture including the restored block.

[0073] The residual information can be generated through conversion and quantization procedures. For example, the encoding device can derive a residual block between the original block and the predicted block, perform a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, and perform a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, thereby signaling the relevant residual information (through a bitstream) to the decoding device. Here, the residual information can include information such as value information, position information, conversion technique, conversion kernel, and quantization parameter of the quantized conversion coefficients. The decoding device can perform an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse-quantize / inverse-convert the quantized conversion coefficients for reference for inter prediction of subsequent pictures to derive a residual block, and can generate a restored picture based on this.

[0074] When inter prediction is applied, the prediction unit of the encoding device / decoding device can perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values or motion information) of picture(s) other than the current picture (Inter prediction can be a prediction derived in a manner that is dependent on data elements(e.g., sample values or motion information) of picture(s) other than the current picture). When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The above motion information can include a motion vector and a reference picture index. The above motion information can further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the above reference block and the reference picture including the above temporal neighboring block may be the same or different. The above temporal neighboring block may also be called by names such as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the above temporal neighboring block is also called a collocated picture (colPic).For example, a motion information candidate list can be configured based on adjacent blocks of a current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block can be signaled. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and (normal) merge mode, the motion information of the current block is the same as the motion information of the selected adjacent block. In the case of skip mode, different from the merge mode, the residual signal is not transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected adjacent block can be used as a motion vector predictor, and a motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.

[0075] Video / image encoding procedures based on inter prediction can generally include, for example, the following.

[0076] FIG. 4 shows an example of an inter prediction-based video / image encoding method.

[0077] The encoding device performs inter prediction on the current block (S400). The encoding device can derive the inter prediction mode and motion information of the current block and generate a prediction sample of the current block. Here, the inter prediction mode determination, motion information derivation, and prediction sample generation procedures may be executed simultaneously, or one procedure may be executed before the other procedures. For example, the inter prediction unit of the encoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block, the motion information derivation unit derives the motion information of the current block, and the prediction sample derivation unit can derive the prediction sample of the current block. For example, the inter prediction unit of the encoding device can search for a block similar to the current block within a certain area (search area) of the reference picture through motion estimation, and derive a reference block whose difference from the current block is the smallest or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode applied to the current block among various prediction modes. The encoding device can compare the RD cost for the various prediction modes and determine the optimal prediction mode for the current block.

[0078] For example, when the skip mode or the merge mode is applied to the current block, the encoding device constructs a merge candidate list to be described later, and among the reference blocks pointed to by the merge candidates included in the merge candidate list, a reference block whose difference between the current block and the current block is the smallest or below a certain criterion can be derived. In this case, the merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0079] As another example, when the (A)MVP mode is applied to the current block, the encoding device configures an (A)MVP candidate list described later, and can use the motion vector of the mvp (motion vector predictor) candidate selected from among the mvp candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the above-described motion estimation can be used as the motion vector of the current block, and among the mvp candidates, the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block becomes the selected mvp candidate. An MVD (Motion Vector Difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information regarding the MVD can be signaled to the decoding device. Further, when the (A)MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and signaled to the decoding device separately.

[0080] The encoding device can derive a residual sample based on the prediction sample (S410). The encoding device can derive the residual sample by comparing the original sample of the current block with the prediction sample.

[0081] The encoding device encodes image information including prediction information and residual information (S420). The encoding device can output the encoded image information in the form of a bitstream. The above prediction information is information related to the above prediction procedure, and can include prediction mode information (for example, skip flag, merge flag, or mode index, etc.) and information related to motion information. The information related to the above motion information can include candidate selection information (for example, merge ind, for example, mvp flag or mvp index) which is information for deriving a motion vector. Also, the information related to the above motion information can include information related to the above-mentioned MVD and / or reference picture index information. Also, the information related to the above motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The above residual information is information related to the above residual samples. The above residual information can include information related to the quantized transform coefficients for the above residual samples.

[0082] The output bitstream can be stored in a (digital) storage medium and transmitted to the decoding device, or can be transmitted to the decoding device via a network.

[0083] On the other hand, as described above, the encoding device can generate a restored picture (including restored samples and restored blocks) based on the above reference samples and the above residual samples. This is to derive the same prediction result in the encoding device as that executed in the decoding device, thereby improving the coding efficiency. Therefore, the encoding device can store the restored picture (or restored samples, restored blocks) in the memory and utilize it as a reference picture for inter prediction. As described above, an in-loop filtering procedure and the like can be further applied to the above restored picture.

[0084] The video / image decoding procedure based on inter prediction can generally include, for example, the following.

[0085] FIG. 5 shows an example of an inter-prediction-based video / image decoding method.

[0086] As shown in FIG. 5, the decoding device can execute operations corresponding to the operations executed by the encoding device. The decoding device can perform prediction on the current block based on the received prediction information to derive a prediction sample.

[0087] Specifically, the decoding device can determine a prediction mode for the current block based on the received prediction information (S500). The decoding device can determine which inter-prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0088] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block or whether the (A)MVP mode is determined. Alternatively, one can be selected from various inter-prediction mode candidates based on the mode index. The inter-prediction mode candidates can include a skip mode, a merge mode, and / or the (A)MVP mode, or can include various inter-prediction modes described later.

[0089] The decoding device derives motion information of the current block based on the determined inter-prediction mode (S510). For example, when the skip mode or the merge mode is applied to the current block, the decoding device can construct a merge candidate list described later and select one merge candidate from the merge candidates included in the merge candidate list. The selection can be executed based on the above-described selection information (merge index). The motion information of the current block can be derived using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0090] As another example, when the (A)MVP mode is applied to the current block, the decoding device constructs an (A)MVP candidate list described later, and can use the motion vector of the mvp (motion vector predictor) candidate selected from among the mvp candidates included in the (A)MVP candidate list as the mvp of the current block. The above selection can be executed based on the above-described selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the mvp and the MVD of the current block. Also, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index in the reference picture list for the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.

[0091] On the other hand, as will be described later, the motion information of the current block can be derived without constructing a candidate list. In this case, the motion information of the current block can be derived by the procedure disclosed in the prediction mode described later. In this case, the candidate list construction as described above can be omitted.

[0092] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S520). In this case, the reference picture can be derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, in some cases, a prediction sample filtering procedure for all or part of the prediction samples of the current block can be further executed.

[0093] For example, the inter prediction unit of the decoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. Based on the prediction mode information received by the prediction mode determination unit, it determines the prediction mode for the current block, derives the motion information (such as motion vectors and / or reference picture indexes, etc.) of the current block based on the information related to the motion information received by the motion information derivation unit, and the prediction sample derivation unit can derive the prediction sample of the current block.

[0094] The decoding device generates a residual sample for the current block based on the received residual information (S530). The decoding device can generate a restored sample for the current block based on the prediction sample and the residual sample, and generate a restored picture based on this (S540). As described above, subsequently, in-loop filtering procedures and the like can be further applied to the restored picture.

[0095] FIG. 6 exemplarily shows the inter prediction procedure.

[0096] As shown in FIG. 6 and as described above, the inter prediction procedure can include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. As described above, the inter prediction procedure can be executed by the encoding device and the decoding device. In this document, the coding device can include the encoding device and / or the decoding device.

[0097] As shown in FIG. 6, the coding device determines an inter prediction mode for the current block (S600). For the prediction of the current block within a picture, various inter prediction modes can be used. For example, various modes such as a merge mode, a skip mode, an MVP (Motion Vector Prediction) mode, an affine mode, a sub-block merge mode, an MMVD (Merge with MVD) mode, etc. can be used. A DMVR (Decoder side Motion Vector Refinement) mode, an AMVR (Adaptive Motion Vector Resolution) mode, a Bi-prediction with CU-level weight (BCW), a Bi-Directional Optical Flow (BDOF), etc. can be used as or instead of an accompanying mode. The affine mode is also called an affine motion prediction mode. The MVP mode is also called an AMVP (Advanced Motion Vector Prediction) mode. In this document, some modes and / or motion information candidates derived by some modes can also be included as one of the motion information-related candidates of other modes. For example, an HMVP candidate can be added as a merge candidate of the above merge / skip mode, or can be added as an mvp candidate of the above MVP mode. When the above HMVP candidate is used as a motion information candidate of the above merge mode or skip mode, the above HMVP candidate is also called an HMVP merge candidate.

[0098] Prediction mode information indicating an inter prediction mode of a current block can be signaled from an encoding device to a decoding device. The prediction mode information can be included in a bitstream and received by the decoding device. The prediction mode information can include index information indicating one of a plurality of candidate modes. Alternatively, the inter prediction mode can also be indicated via hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags. For example, a skip flag is signaled to indicate whether the skip mode can be applied, a merge flag is signaled when the skip mode is not applied to indicate whether the merge mode can be applied, it is indicated that the MVP mode is applied when the merge mode is not applied, or a flag for additional classification can also be signaled. The affine mode can be signaled as an independent mode, or can also be signaled as a mode subordinate to the merge mode or the MVP mode, etc. For example, the affine mode can include an affine merge mode and an affine MVP mode.

[0099] The coding device derives motion information for the current block (S610). The motion information derivation can be derived based on the inter prediction mode.

[0100] The coding device can perform inter prediction using the motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block in the original picture for the current block to search for similar reference blocks with high correlation within a determined search range in the reference picture in units of fractional pixels, thereby deriving motion information. The similarity of the blocks can be derived based on the difference in phase-based sample values. For example, the similarity of the blocks can be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, the motion information can be derived based on the reference block with the smallest SAD in the search area. The derived motion information can be signaled to the decoding device in various ways based on the inter prediction mode.

[0101] The coding device performs inter prediction based on the motion information for the above current block (S620). The coding device can derive prediction sample(s) for the current block based on the motion information. The current block including the prediction sample is also called a predicted block.

[0102] On the other hand, in inter prediction, according to the conventional merge or AMVP mode, a method of reducing the amount of motion information by using the motion vectors of spatially / temporally adjacent blocks of the current block as motion information candidates has been used. For example, the adjacent blocks used to derive the motion information candidates of the current block could include the lower left corner adjacent block, left adjacent block, upper right corner adjacent block, upper adjacent block, and upper left corner adjacent block of the current block.

[0103] FIG. 7 exemplarily shows the spatial adjacent blocks used for deriving motion information candidates in the conventional merge or AMVP mode.

[0104] Basically, the above spatial adjacent blocks are limited to the blocks in contact with the current block. This is to enhance hardware realizability. To derive information of blocks far from the current block, problems such as an increase in line buffer occur. However, using motion information of non-adjacent blocks to derive candidate motion information of the current block can constitute various candidates, thus bringing about performance improvement. The HMVP (History based Motion Vector Prediction) method can be used to use motion information of non-adjacent blocks without increasing the line buffer. In this document, HMVP can indicate History based Motion Vector Prediction or History based Motion Vector Predictor. According to this document, inter prediction can be efficiently executed using HMVP, and parallel processing can be supported. For example, in the embodiments of this document, various methods for managing the history buffer for parallelization processing are proposed, and parallel processing can be supported based on this. However, supporting parallel processing does not mean that parallel processing must be executed. Considering hardware performance and service form, the coding device may or may not execute parallel processing. For example, when the coding device is equipped with a multi-core processor, the coding device can parallel-process a part of slices, blocks, and / or tiles. On the other hand, when the coding device is equipped with a single-core processor or a multi-core processor, the coding device can also execute sequential processing while reducing the calculation and memory burden.

[0105] The HMVP candidates by the above-described HMVP method can include motion information of previously coded blocks. For example, the motion information of previously coded blocks according to the block coding order within the current picture is not considered as the motion information of the current block when the previously coded block does not adjoin the current block. However, the HMVP candidates can be considered as candidates for the motion information of the current block (e.g., merge candidates or MVP candidates) without considering whether the previously coded block adjoins the current block. In this case, a plurality of HMVP candidates can be stored in the buffer. For example, when the merge mode is applied to the current block, the HMVP candidates (HMVP merge candidates) can be added to the merge candidate list. In this case, the HMVP candidates can be added next to the spatial merge candidates and the temporal merge candidates included in the merge candidate list.

[0106] According to the HMVP method, the motion information of previously coded blocks can be stored in tabular form and used as motion information candidates (e.g., merge candidates) for the current block. A table (or buffer, list) containing multiple HMVP candidates can be maintained during the encoding / decoding procedure. The above table (or buffer, list) is also referred to as the HMVP table (or buffer, list). According to one embodiment of this document, the above table (or buffer, list) can be initialized when encountering a new slice. Alternatively, according to one embodiment of this document, the above table (or buffer, list) can be initialized when contacting a new CTU row. When the above table is initialized, the number of HMVP candidates contained in the table can be set to 0. The size of the above table (or buffer, list) can be fixed to a specific value (e.g., 5, etc.). For example, if there is an inter-coded block, the relevant motion information can be added as a new HMVP candidate at the last entry of the above table. The above (HMVP) table is also referred to as the (HMVP) buffer or (HMVP) list.

[0107] Figure 8 schematically shows an example of an HMVP candidate-based decoding procedure. Here, the HMVP candidate-based decoding procedure can include an HMVP candidate-based inter-prediction procedure.

[0108] As shown in FIG. 8, the decoding device loads an HMVP table including one or more HMVP candidates, and decodes a block based on at least one of the one or more HMVP candidates. Specifically, for example, the decoding device can derive motion information of a current block based on at least one of the one or more HMVP candidates, perform an inter prediction on the current block based on the motion information, and derive a predicted block (including predicted samples). As described above, a restored block can be generated based on the predicted block. The derived motion information of the current block can be updated in the table. In this case, the motion information can be added as a new HMVP candidate as the last entry of the table. When the number of HMVP candidates already included in the table is the same as the size of the table, the candidate that first entered the table is deleted, and the derived motion information can be added as a new HMVP candidate to the last entry of the table.

[0109] FIG. 9 exemplarily shows HMVP table update according to the FIFO rule, and FIG. 10 exemplarily shows HMVP table update according to the restricted FIFO rule.

[0110] The FIFO (First-In-First-Out) rule can be applied to the table. For example, when the table size S is 16, this indicates that 16 HMVP candidates can be included in the table. When more than 16 HMVP candidates are generated from previously coded blocks, the FIFO rule can be applied, whereby the table can include the latest coded up to 16 motion information candidates. In this case, as shown in FIG. 9, the FIFO rule is applied to remove the oldest HMVP candidate, and a new HMVP candidate can be added.

[0111] On the one hand, in order to further improve the coding efficiency, a restricted FIFO rule can also be applied as shown in FIG. 10. As shown in FIG. 10, when inserting an HMVP candidate into the table, first, a redundancy check can be applied. Thereby, it can be determined whether an HMVP candidate having the same motion information already exists in the table. If an HMVP candidate having the same motion information exists in the table, the HMVP candidate having the same motion information is removed from the table, and the HMVP candidates after the removed HMVP candidate move one by one (i.e., each index - 1), and then new HMVP candidates can be inserted.

[0112] As described above, the HMVP candidate can be used in the merge candidate list construction procedure. In this case, for example, all HMVP candidates that can be inserted from the last entry to the first entry in the table can be inserted after the spatial merge candidates and the temporal merge candidates. In this case, a pruning check can be applied to the HMVP candidate. The maximum number of allowed merge candidates can be signaled, and when the total number of available merge candidates reaches the maximum number of merge candidates, the merge candidate list construction procedure can be terminated.

[0113] Similarly, the HMVP candidate can also be used in (A) the MVP candidate list construction procedure. In this case, the motion vectors of the last k HMVP candidates in the HMVP table can be added after the TMVP candidates that make up the MVP candidate list. In this case, for example, an HMVP candidate having the same reference picture as the MVP target reference picture can be used for constructing the MVP candidate list. Here, the MVP target reference picture can indicate the reference picture for the inter - prediction of the current block to which the MVP mode is applied. In this case, a pruning check can be applied to the HMVP candidate. The above k is, for example, 4. However, this is only an example, and the above k can have various values such as 1, 2, 3, 4, etc.

[0114] On the other hand, when the total number of merge candidates is the same as or greater than 15, a truncated unary plus fixed length (with 3 bits) binarization method as shown in Table 1 below can be applied for merge index coding.

[0115] [Table 1]

[0116] The above table assumes the case where Nmrg = 15, and Nmrg indicates the total number of merge candidates.

[0117] On the other hand, when developing a solution applying a video codec, parallel processing can also be supported in image / video coding for implementation optimization.

[0118] FIG. 11 exemplarily shows WPP (Wavefront Parallel Processing), which is one of the techniques for parallel processing.

[0119] As shown in FIG. 11, when WPP is applied, parallel processing can be performed in CTU row units. In this case, dependencies exist with respect to the positions pointed to by the arrows when coding (encoding / decoding) the blocks indicated by X. Therefore, it is necessary to wait for the coding of the upper right CTU of the block currently being coded to be completed. Also, when WPP is applied, the initialization of the CABAC probability table (or context information) can be done in slice units, and for parallel processing including entropy encoding / decoding, the CABAC probability table (or context information) must be initialized in CTU row units. WPP can be regarded as a technique proposed to determine an efficient initialization position. When WPP is applied, each LCT row can be called a sub-stream, and when the coding device has multiple processing cores, parallel processing can be supported. For example, when WPP is applied, if three processing cores decode in parallel, the first processing core can decode sub-stream 0, the second processing core can decode sub-stream 1, and the third processing core can decode sub-stream 2. When WPP is applied, after coding is performed on the nth (n is an integer) sub-stream, after the coding of the second CTU or LCU of the nth sub-stream is completed, coding on the (n + 1)th sub-stream can be performed. For example, in the case of entropy coding, if the entropy coding of the second LCU of the nth sub-stream is completed, the first LCU of the (n + 1)th sub-stream can be entropy-coded based on the context information of the second LCU of the nth sub-stream. At this time, the number of sub-streams within a slice can be the same as the number of LCU rows. Also, the number of sub-streams within a slice can be the same as the number of entry points. At this time, the number of entry points can be specified by the number of entry point offsets.For example, the number of entry points can have a value that is 1 greater than the number of entry point offsets. Information regarding the number of entry point offsets and / or information regarding the values of the offsets can be included in and encoded in the video / image information described above, and can be signaled to the decoding device via a bitstream. On the other hand, when the coding device includes one processing core, coding processing can be performed in units of one substream, through which the memory load and coding dependency can be reduced.

[0120] The above-described HMVP method stores, as candidates, motion information derived in the coding procedure of each block by the size of a predetermined buffer (HMVP table). In this case, as many candidates as the number of buffers can be satisfied as disclosed in FIG. 9 without additional conditions, or candidates can be satisfied so as not to overlap through duplicate checking between newly added candidates and candidates existing in the buffer (HMVP table). Thereby, various candidates can be configured. However, when developing a solution applying a video codec, since the point in time when the HMVP candidates are filled in the buffer cannot generally be known, it is impossible to achieve parallel processing with or without applying WPP.

[0121] FIG. 12 exemplarily shows problems when applying a general HMVP method in consideration of parallel processing.

[0122] As shown in FIG. 12, when parallelizing on a per CTU row basis like WPP, dependency issues of the HMVP buffer may occur. For example, the HMVP buffer for the first CTU in the N (N>=1)-th CTU row is not filled until the coding (encoding / decoding) of the block existing in the (N-1)-th CTU row, for example, the block within the last CTU of the (N-1)-th CTU row, is completed. That is, when parallel processing is applied under the current structure, the decoding device cannot know whether the HMVP candidates stored in the current HMVP buffer match the HMVP buffer used for decoding the current (target) block. This is because there may be a difference between the HMVP buffer derived at the time of coding the current block when applying sequential processing and the HMVP buffer derived at the time of coding the current block when applying parallel processing.

[0123] In one embodiment of this document, in order to solve the above problems, when applying HMVP, parallel processing is supported by initializing the history management buffer (HMVP buffer).

[0124] FIG. 13 exemplarily shows a method for initializing the history management buffer (HMVP buffer) according to one embodiment of this document.

[0125] As shown in FIG. 13, the HMVP buffer can be initialized for each first CTU in a CTU row. That is, when coding the first CTU in a CTU row, by initializing the HMVP buffer, the number of HMVP candidates included in the HMVP buffer can be set to 0. As described above, by initializing the HMVP buffer for each CTU row, even when parallel processing is supported, the HMVP candidates derived in the coding process of the CTU located in the left direction of the current block can be used without restrictions. In this case, for example, if the current CU, which is the current block, is located at the first CTU in the CTU row and the current CU corresponds to the first CU of the first CTU, the number of HMVP candidates included in the HMVP buffer is 0. Also, for example, if a CU coded before the current CU in the CTU row is coded in inter mode, HMVP candidates can be derived based on the motion information of the previously coded CU and included in the HMVP buffer.

[0126] FIG. 14 exemplarily shows an HMVP buffer management method according to an embodiment.

[0127] As shown in FIG. 14, the HMVP buffer can be initialized in slice units, and for each CTU in a slice, it is possible to determine whether the coding target CTU (current CTU) is the first CTU in each CTU row. In FIG. 14, it is described that, as an example, when (ctu_idx % Num) is 0, it is determined to be the first CTU. At this time, Num means the number of CTUs in each CTU row. As another example, when using the aforementioned brick concept, when (ctu_idx_in_brick % BrickWidth) is 0, it can be determined that it is the first CTU in the CTU row (within the corresponding brick). Here, ctu_idx_in_brick indicates the index of the corresponding CTU within the above brick, and BrickWidth represents the width of the corresponding brick in CTU units. That is, BrickWidth can indicate the number of CTU columns within the corresponding brick. When the current CTU is the first CTU in the CTU row, the HMVP buffer is initialized (that is, the number of candidates in the HMVP buffer is set to 0), and if not, the HMVP buffer is maintained. Thereafter, through the prediction process for each CU within the corresponding CTU (for example, merge or MVP mode-based), at this time, the candidates stored in the HMVP buffer can be included as motion information candidates (for example, merge candidates or MVP candidates) for the merge mode or MVP mode. The motion information of the target block (current block) derived in the inter prediction process based on the merge mode or MVP mode or the like is stored (updated) in the HMVP buffer as a new HMVP candidate. In this case, the aforementioned duplicate check process can also be further executed. Thereafter, the aforementioned procedure can be repeated for the CU and CTU.

[0128] As another example, when applying HMVP, it is also possible to remove the CTU unit dependency by initializing the HMVP buffer for each CTU.

[0129] FIG. 15 exemplarily shows an HMVP buffer management method according to another embodiment.

[0130] As shown in Fig. 15, the current CTU can perform HMVP buffer initialization for each CTU without determining whether the current CTU is the first CTU in each CTU row. In this case, since the HMVP buffer is initialized in CTU units, the motion information of the blocks existing within the CTU is stored in the HMVP table. In this case, the HMVP candidates can be derived based on the motion information of the blocks (e.g., CUs) within the same CTU, and as described below, HMVP buffer initialization becomes possible without determining whether the current CTU is the first CTU in each CTU row.

[0131] As described above, the HMVP buffer can be initialized in slice units, and through this, it is possible to use the motion vectors of blocks that are spatially separated from the current block. However, in this case, since parallel processing support is not possible within the slice, in the above-described embodiments, etc., a method of initializing the buffer in CTU rows or CTU units has been proposed. That is, according to the embodiments of this document, etc., the HMVP buffer can be initialized in slice units, and within the slice, it can be initialized in CTU row units.

[0132] On the other hand, when coding (encoding / decoding) one picture, the picture can be divided in slice units and / or the picture can also be divided in tile units. For example, considering error resilience, the picture can be divided in slice units, or in order to encode / decode a partial region within the picture, the picture can also be divided in tile units. When one picture is divided into a plurality of tiles, when applying the HMVP management buffer, performing initialization in CTU row units within the picture, that is, initializing the HMVP buffer with the first CTU in each CTU row within the picture, is not suitable for the tile structure for encoding / decoding a part of the picture.

[0133] Fig. 16 exemplarily shows the HMVP buffer initialization method in the tile structure.

[0134] As shown in FIG. 16, in the case of tiles 1 and 3, since the HMVP management buffer is not initialized for each tile unit, dependencies (HMVP) occur with tiles 0 and 2 respectively. Therefore, when tiles exist, it is possible to initialize the HMVP buffer in the following manner.

[0135] As an example, the HMVP buffer can be initialized in CTU units. It is natural that this can be applied without distinguishing tiles, slices, etc.

[0136] As another example, the HMVP buffer can be initialized for the first CTU of each tile.

[0137] FIG. 17 shows an example of a method for initializing the HMVP buffer for the first CTU of a tile according to another embodiment.

[0138] As shown in FIG. 17, when coding the first CTU of each tile, the HMVP buffer is initialized. That is, when coding tile 0, HMVP buffer 0 is initialized and used, and when coding tile 1, HMVP buffer 1 can be initialized and used.

[0139] As yet another example, the HMVP buffer can be initialized for the first CTU of the CTU rows within each tile.

[0140] FIG. 18 shows an example of a method for initializing the HMVP management buffer for the first CTU of the CTU rows within each tile according to yet another embodiment.

[0141] As shown in FIG. 18, the HMVP buffer can be initialized for each CTU row of each tile. For example, the HMVP buffer is initialized at the first CTU of the first CTU row of tile n, the HMVP buffer is initialized at the first CTU of the second CTU row of tile n, and the HVMP buffer can be initialized at the first CTU of the third CTU row of tile n. In this case, if the coding device has a multi-core processor, the coding device can initialize and use HMVP buffer 0 for the first CTU row of tile n, initialize and use HVMP buffer 1 for the second CTU row of tile n, and initialize and use HMVP buffer 2 for the third CTU row of tile n, and support parallel processing through this. On the other hand, when the coding device is equipped with a single-core processor, the coding device can initialize and reuse the HVMP buffer at the first CTU of each CTU row within each tile according to the coding order.

[0142] On the other hand, due to the tile division structure and the slice division structure, tiles and slices can exist simultaneously within one picture.

[0143] FIG. 19 shows an example of a structure in which tiles and slices exist simultaneously.

[0144] FIG. 19 exemplarily shows a case where one picture is divided into 4 tiles and there are 2 slices within each tile. As shown in FIG. 19, there may be a case where both slices and tiles exist within one picture, and it is possible to initialize the HMVP buffer as follows.

[0145] As an example, the HMVP buffer can be initialized in units of CTUs. Such a method can be applied without distinguishing whether the CTU is located in a tile or in a slice.

[0146] As another example, the HMVP buffer can be initialized for the first CTU within each tile.

[0147] FIG. 20 shows an example of a method for initializing the HMVP buffer for the first CTU in each tile.

[0148] As shown in FIG. 20, the HMVP buffer can be initialized with the first CTU of each tile. Even when there are multiple slices within one tile, the HMVP buffer initialization can be performed with the first CTU in the tile.

[0149] As yet another example, the HMVP buffer initialization can also be performed for each slice existing within a tile.

[0150] FIG. 21 shows an example of a method for initializing the HMVP buffer for each slice in a tile.

[0151] As shown in FIG. 21, the HMVP buffer can be initialized with the first CTU of each slice in the tile. Therefore, when there are multiple slices within one tile, the HMVP buffer initialization can be performed for each of the above-mentioned multiple slices. In this case, when processing the first CTU of each slice, the above HMVP buffer initialization can be performed.

[0152] On the other hand, there may be multiple tiles within one picture without slices. Alternatively, there can also be multiple tiles within one slice. In such cases, the HMVP buffer initialization can be performed as follows.

[0153] As an example, the HMVP buffer can be initialized for each tile group unit.

[0154] FIG. 22 shows an example of initializing the HMVP buffer for the first CTU of the first tile in a tile group.

[0155] As shown in FIG. 22, one picture can be divided into two tile groups, and each tile group (TileGroup0, TileGroup1) can be further divided into a plurality of tiles. In this case, the HMVP buffer can be initialized for the first CTU of the first tile within one tile group.

[0156] As another example, the HMVP buffer can be initialized on a per-tile basis within a tile group.

[0157] FIG. 23 shows an example of initializing the HMVP buffer for the first CTU of each tile within a tile group.

[0158] As shown in FIG. 23, one picture can be divided into two tile groups, and each tile group (TileGroup0, TileGroup1) can be further divided into a plurality of tiles. In this case, the HMVP buffer can be initialized for the first CTU of each tile within one tile group.

[0159] As yet another example, the HMVP buffer can be initialized for each CTU row of each tile within a tile group.

[0160] FIG. 24 shows an example of initializing the HMVP buffer for each CTU row of each tile within a tile group.

[0161] As shown in FIG. 24, one picture can be divided into two tile groups, and each tile group (TileGroup0, TileGroup1) can be further divided into a plurality of tiles. In this case, the HMVP buffer can be initialized with the first CTU of each CTU row of each tile within one tile group.

[0162] Alternatively, in this case as well, the HMVP management buffer can be initialized on a per-CTU basis. It goes without saying that this can be applied without distinguishing tiles, slices, tile groups, etc.

[0163] FIG. 25 and FIG. 26 schematically show an example of a video / image encoding method and related components including an inter prediction method according to one or more embodiments of the present document. The method disclosed in FIG. 25 can be performed by the encoding apparatus disclosed in FIG. 2. Specifically, for example, S2500 to S2530 in FIG. 25 can be performed by the prediction unit 220 of the encoding apparatus, S2540 in FIG. 25 can be performed by the residual processing unit 230 of the encoding apparatus, and S2550 in FIG. 25 can be performed by the entropy encoding unit 240 of the encoding apparatus. The method disclosed in FIG. 25 can include the embodiments described above in the present document and the like.

[0164] As shown in FIG. 25, the encoding apparatus derives an HMVP buffer for the current block (S2500). The encoding apparatus can perform the HMVP buffer management method described above in embodiments of the present document and the like. For example, the HMVP buffer can be initialized in units of slice, tile, or tile group. And / or the HMVP buffer can be initialized in units of CTU rows. In this case, the HMVP buffer can be initialized in units of CTU rows including the current block within the current tile. Here, a tile can indicate a rectangular region such as a CTU in a picture. A tile can be defined based on a specific tile row and a specific tile column within the picture. For example, one or more tiles may exist in the current picture. In this case, the HMVP buffer can be initialized with the first CTU of the CTU row including the current block within the current tile. Alternatively, one or more slices may exist in the current picture. In this case, the HMVP buffer can be initialized with the first CTU of the CTU row including the current block within the current slice. Alternatively, one or more tile groups may exist in the current picture. In this case, the HMVP buffer can be initialized with the first CTU of the CTU row including the current block within the current tile group.

[0165] The encoding device can determine whether the current CTU is the first CTU in the CTU row. In this case, the HMVP buffer can be initialized with the first CTU in the CTU row where the current CTU containing the current block is located. In other words, the HMVP buffer can be initialized when processing the first CTU in the CTU row where the current CTU containing the current block is located. When it is determined that the current CTU containing the current block is the first CTU in the CTU row within the current tile, the HMVP buffer includes HMVP candidates derived based on the motion information of the blocks processed before the current block within the current CTU. When it is determined that the current CTU is not the first CTU in the CTU row within the current tile, the HMVP buffer can include HMVP candidates derived based on the motion information of the blocks processed before the current block within the CTU row within the current tile. Also, for example, when the current CU, which is the current block, is located at the first CTU in the CTU row within the current tile and the current CU corresponds to the first CU of the first CTU, the number of HMVP candidates included in the HMVP buffer is 0. Also, for example, if the CU (e.g., the CU coded before the current CU within the current CTU and / or the CU within the CTU coded before the current CTU within the CTU row) coded before the current CU within the CTU row within the current tile is coded in inter mode, HMVP candidates can be derived based on the motion information of the previously coded CU and included in the HMVP buffer.

[0166] When the current picture is divided into a plurality of tiles, the HMVP buffer can be initialized in units of CTU rows within each tile.

[0167] The above-mentioned HMVP buffer can be initialized in units of CTU rows within a tile or a slice. For example, when a specific CTU in the above-mentioned CTU row is not the first CTU in the above-mentioned CTU row in the current picture, but the specific CTU is the first CTU in the above-mentioned CTU row in the current tile or the current slice, the above-mentioned HMVP buffer can be initialized with the above-mentioned specific CTU.

[0168] When the above-mentioned HMVP buffer is initialized, the number of HMVP candidates included in the above-mentioned HMVP buffer can be set to 0.

[0169] The encoding device constructs a motion information candidate list based on the above-mentioned HMVP buffer (S2510). The above-mentioned HMVP buffer can include HMVP candidates, and the motion information candidate list including the above-mentioned HMVP candidates can be constructed.

[0170] As an example, when the merge mode is applied to the current block, the motion information candidate list can be a merge candidate list. As another example, when the (A)MVP mode is applied to the current block, the motion information candidate list can be an MVP candidate list. When the merge mode is applied to the current block, the above-mentioned HMVP candidate can be added to the merge candidate list when the number of available merge candidates (including, for example, spatial merge candidates and temporal merge candidates) in the merge candidate list for the current block is less than a predetermined maximum number of merge candidates. In this case, the above-mentioned HMVP candidate can be inserted after the spatial candidates and temporal candidates in the merge candidate list. In other words, a larger index value than the index values assigned to the spatial candidates and temporal candidates in the merge candidate list can be assigned to the above-mentioned HMVP candidate. When the (A)MVP mode is applied to the current block, the above-mentioned HMVP candidate can be added to the MVP candidate list when the number of available MVP candidates (derived based on spatially adjacent blocks and temporally adjacent blocks) in the MVP candidate list for the current block is less than 2.

[0171] The encoding device can derive the motion information of the current block based on the above motion information candidate list (S2520).

[0172] The encoding device can derive the motion information of the current block based on the above motion information candidate list. For example, when the merge mode or MVP mode is applied to the current block, the above HMVP candidates included in the HMVP buffer can be used as merge candidates or MVP candidates. For example, when the merge mode is applied to the current block, the above HMVP candidates included in the HMVP buffer are included as candidates in the merge candidate list, and the above HMVP candidates can be indicated among the candidates included in the merge candidate list based on the merge index. The above merge index is prediction-related information and can be included in the image / video information described later. In this case, the above HMVP candidates can be indexed within the merge candidate list with a lower priority than the spatial merge candidates and temporal merge candidates included in the above merge candidate list. That is, an index value assigned to the above HMVP candidates can be assigned a value higher than the index values of the above spatial merge candidates and temporal merge candidates. As another example, when the MVP mode is applied to the current block, the above HMVP candidates included in the HMVP buffer are included as candidates in the merge candidate list, and the above HMVP candidates can be indicated among the candidates included in the MVP candidate list based on the MVP flag (or MVP index). The above MVP flag (or MVP index) is prediction-related information and can be included in the image / video information described later.

[0173] The encoding device generates a prediction sample for the current block based on the above-derived motion information (S2530). The encoding device performs inter prediction (motion compensation) based on the above motion information, and can derive a prediction sample using the reference sample pointed to by the above motion information on the reference picture.

[0174] The encoding device generates a residual sample based on the above prediction sample (S2540). The encoding device can generate a residual sample based on the original sample for the current block and the prediction sample for the current block.

[0175] The encoding device derives information regarding the residual sample based on the above residual sample, and encodes the image / video information including the information regarding the residual sample (S2550). The information regarding the residual sample can be called residual information, and can include information regarding quantized transform coefficients. The encoding device can perform a transform / quantization procedure on the above residual sample to derive quantized transform coefficients.

[0176] The encoded image / video information can be output in the form of a bitstream. The above bitstream can be transmitted to a decoding device via a network or a storage medium. The image / video information can further include prediction-related information, and the above prediction-related information can further include information regarding various prediction modes (e.g., merge mode, MVP mode, etc.), MVD information, etc.

[0177] FIG. 27 and FIG. 28 schematically show an example of an image decoding method including an inter prediction method according to an embodiment of this document and related components. The method disclosed in FIG. 27 can be performed by the decoding device disclosed in FIG. 3. Specifically, for example, S2700 to S2730 in FIG. 27 can be performed by the prediction unit 330 of the above decoding device, and S2740 can be performed by the addition unit 340 of the above decoding device. The method disclosed in FIG. 27 can include the embodiments described above in this document.

[0178] As shown in FIG. 27, the decoding device derives an HMVP buffer for the current block (S2700). The decoding device can perform the HMVP buffer management method described above in the embodiments of this document. For example, the HMVP buffer can be initialized on a per-slice, per-tile, or per-tile-group basis. And / or, the HMVP buffer can be initialized on a per-CTU row basis. In this case, the HMVP buffer can be initialized on a per-CTU row basis within the slice, tile, or tile group. Here, a tile can indicate a rectangular region such as a CTU within a picture. A tile can be specified based on a specific tile row and a specific tile column within the picture. For example, there may be one or more tiles within the current picture. In this case, the HMVP buffer can be initialized with the first CTU of the CTU row containing the current block within the current tile. Alternatively, there may be one or more slices within the current picture. In this case, the HMVP buffer can be initialized with the first CTU of the CTU row containing the current block within the current slice. Alternatively, there may be one or more tile groups within the current picture. In this case, the HMVP buffer can be initialized with the first CTU of the CTU row containing the current block within the current tile group.

[0179] The decoding device can determine whether the current CTU is the first CTU in the CTU row. In this case, the HMVP buffer can be initialized with the first CTU in the CTU row where the current CTU containing the current block is located. In other words, the HMVP buffer can be initialized when processing the first CTU in the CTU row where the current CTU containing the current block is located. When it is determined that the current CTU containing the current block is the first CTU in the CTU row within the current tile, the HMVP buffer includes HMVP candidates derived based on the motion information of the blocks processed earlier than the current block within the current CTU. When it is determined that the current CTU is not the first CTU in the CTU row within the current tile, the HMVP buffer can include HMVP candidates derived based on the motion information of the blocks processed earlier than the current block within the CTU row within the current tile. Also, for example, when the current CU, which is the current block, is located at the first CTU in the CTU row within the current tile and the current CU corresponds to the first CU of the first CTU, the number of HMVP candidates included in the HMVP buffer is 0. Also, for example, if the CU (e.g., the CU coded earlier than the current CU in the current CTU and / or the CU within the CTU coded earlier than the current CTU in the CTU row) coded earlier than the current CU in the CTU row within the current tile is coded in inter mode, HMVP candidates are derived based on the motion information of the previously coded CU and can be included in the HMVP buffer.

[0180] When the current picture is divided into a plurality of tiles, the HMVP buffer can be initialized in units of CTU rows within each tile.

[0181] The above-mentioned HMVP buffer can be initialized in units of CTU rows within a tile or a slice. For example, if a specific CTU in the above CTU row is not the first CTU in the above CTU row in the current picture, but the first CTU in the above CTU row in the current tile or the current slice, the above HMVP buffer can be initialized with the above specific CTU.

[0182] When the above HMVP buffer is initialized, the number of HMVP candidates included in the above HMVP buffer can be set to 0.

[0183] The decoding device constructs a motion information candidate list based on the above HMVP buffer (S2710). The above HMVP buffer can include HMVP candidates, and the above motion information candidate list including the above HMVP candidates can be constructed.

[0184] As an example, when the merge mode is applied to the current block, the above motion information candidate list can be a merge candidate list. As another example, when the (A)MVP mode is applied to the current block, the above motion information candidate list can be an MVP candidate list. When the merge mode is applied to the current block, the above HMVP candidate can be added to the above merge candidate list when the number of available merge candidates (including, for example, spatial merge candidates and temporal merge candidates) in the merge candidate list for the current block is less than a predetermined maximum number of merge candidates. In this case, the above HMVP candidate can be inserted after the above spatial and temporal candidates in the above merge candidate list. In other words, the above HMVP candidate can be assigned an index value larger than the index assigned to the above spatial and temporal candidates in the above merge candidate list. When the (A)MVP mode is applied to the current block, the above HMVP candidate can be added to the above MVP candidate list when the number of available MVP candidates (derived based on spatially adjacent blocks and temporally adjacent blocks) in the MVP candidate list for the current block is less than 2.

[0185] The decoding device can derive the motion information of the current block based on the motion information candidate list (S2720).

[0186] The encoding device can derive the motion information of the current block based on the motion information candidate list. For example, when the merge mode or the MVP mode is applied to the current block, the HMVP candidates included in the HMVP buffer can be used as merge candidates or MVP candidates. For example, when the merge mode is applied to the current block, the HMVP candidates included in the HMVP buffer are included as candidates in the merge candidate list, and the HMVP candidate among the candidates included in the merge candidate list can be indicated based on the merge index obtained from the bitstream. In this case, the HMVP candidate can be assigned an index within the merge candidate list with a lower priority than the spatial merge candidates and the temporal merge candidates included in the merge candidate list. That is, an index value assigned to the HMVP candidate can be a value higher than the index values of the spatial merge candidate and the temporal merge candidate. As another example, when the MVP mode is applied to the current block, the HMVP candidates included in the HMVP buffer are included as candidates in the merge candidate list, and the HMVP candidate among the candidates included in the MVP candidate list can be indicated based on the MVP flag (or, MVP index) obtained from the bitstream.

[0187] The decoding device generates a prediction sample for the current block based on the derived motion information (S2730). The decoding device can derive a prediction sample by performing inter prediction (motion compensation) based on the motion information and using the reference sample pointed to by the motion information on the reference picture. The current block including the prediction sample is also called a predicted block.

[0188] The decoding device generates a restored sample based on the above prediction sample (S2740). As described above, a restored block / picture can be generated based on the above restored sample. The decoding device can obtain residual information (including information on quantized transform coefficients) from the above bitstream, can derive a residual sample based on the above residual information, and can generate the above restored sample based on the above prediction sample and the above residual sample, as described above. Thereafter, if necessary, in-loop filtering procedures such as deblock filtering, SAO, and / or ALF procedures can be applied to the above restored picture to improve subjective / objective image quality, as described above.

[0189] In the above-described embodiments, the method is described based on a flowchart in a series of steps or blocks, but the embodiments are not limited to the order of the steps, and a certain step may occur in a different order or simultaneously with steps different from the above. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps may be included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.

[0190] The method according to the above-described embodiments of this document can be realized in software form, and the encoding device and / or decoding device according to this document can be included in devices that execute image processing such as TVs, computers, smartphones, set-top boxes, display devices, etc.

[0191] In this document, when an embodiment is realized by software, the above-described method can be realized by a module (process, function, etc.) that performs the above-described functions. The module can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (Application-Specific Integrated Circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (Read-Only Memory), a RAM (Random Access Memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be realized and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be realized and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for realization (e.g., information on instructions) or an algorithm can be stored in a digital storage medium.

[0192] In addition, the decoding device and encoding device to which the embodiment(s) of this document is / are applied can be included in a multimedia broadcast transceiver device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT video (Over The Top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (Virtual Reality) device, an AR (Augmented Reality) device, a picture phone video device, a transportation means terminal (e.g., a vehicle (including an autonomous driving vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process video signals or data signals. For example, an OTT video (Over The Top video) device can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.

[0193] Also, the processing method to which the embodiment(s) of this document is / are applicable can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document can also be stored in a computer-readable recording medium. The above computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which data that can be read by a computer is stored. The above computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial (Universal Serial) Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Also, the above computer-readable recording medium includes a medium realized in the form of a carrier wave (for example, transmission via the Internet). Also, a bit stream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0194] Also, the embodiment(s) of this document can be realized by a computer program product with program code, and the above program code can be performed by a computer according to the embodiment(s) of this document. The above program code can be stored on a carrier readable by a computer.

[0195] FIG. 29 shows an example of a content streaming system to which the embodiment disclosed in this document can be applied.

[0196] As shown in FIG. 29, the content streaming system to which the embodiment of this document is applied generally can include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0197] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and serves to transmit this to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate a bitstream, the encoding server can be omitted.

[0198] The bitstream can be generated by an encoding method or a bitstream generation method to which the embodiments of this document are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0199] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium for notifying the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server serves to control commands / responses between each device within the content streaming system.

[0200] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0201] Examples of the user device include a mobile phone, smartphone, laptop computer, digital broadcast terminal, PDA (Personal Digital Assistants), PMP (Portable Multimedia Player), navigation device, slate PC, tablet PC, ULTRABOOK (registered trademark), wearable device (e.g., smartwatch, smart glass (glass-type terminal), HMD (Head Mounted Display)), digital TV, desktop computer, digital signature, etc.

[0202] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.

Claims

1. An image decoding method performed by a decoding device, comprising: deriving a History-based Motion Vector Prediction (HMVP) candidate list for the current block; constructing a motion information candidate list based on the HMVP candidates included in the HMVP candidate list, where whether to add the HMVP candidate to the motion information candidate list is checked after a time candidate in the motion information candidate list; deriving motion information of the current block based on the motion information candidate list; deriving a reference picture index for the current block based on the motion information; deriving a motion vector for the current block based on the motion information; generating a prediction sample for the current block based on the reference picture index and the motion vector; generating reconstructed samples based on the predicted samples; the current picture includes one or more tiles; the current picture includes a plurality of tile columns and tile rows; A tile is a rectangular region of a coding tree unit (CTU) within a specific tile column and a specific tile row in the current picture, The HMVP candidate list is updated based on the motion information of the previous block; The HMVP candidate list is initialized with a first CTU for each CTU row of each tile; A method, wherein upon initialization of the HMVP candidate list, a number of HMVP candidates included in the HMVP candidate list is set to zero.

2. An image encoding method performed by an encoding device, comprising: deriving a History-based Motion Vector Prediction (HMVP) candidate list for the current block; constructing a motion information candidate list based on the HMVP candidates included in the HMVP candidate list, where whether to add the HMVP candidate to the motion information candidate list is checked after a time candidate in the motion information candidate list; deriving motion information of the current block based on the motion information candidate list; deriving a reference picture index for the current block based on the motion information; deriving a motion vector for the current block based on the motion information; generating a prediction sample for the current block based on the reference picture index and the motion vector; deriving residual samples based on the prediction samples; and encoding image information including information about the residual samples; One or more tiles are in the current picture, the current picture includes a plurality of tile columns and tile rows; A tile is a rectangular region of a coding tree unit (CTU) within a specific tile column and a specific tile row in the current picture, The HMVP candidate list is updated based on the motion information of the previous block; The HMVP candidate list is initialized with a first CTU for each CTU row of each tile; A method, wherein upon initialization of the HMVP candidate list, a number of HMVP candidates included in the HMVP candidate list is set to zero.

3. A method for transmitting data relating to an image, comprising the steps of: Obtaining a bitstream generated by a method, the method comprising: deriving a History-based Motion Vector Prediction (HMVP) candidate list for the current block; constructing a motion information candidate list based on the HMVP candidates included in the HMVP candidate list, where whether to add the HMVP candidate to the motion information candidate list is checked after a time candidate in the motion information candidate list; deriving motion information of the current block based on the motion information candidate list; deriving a reference picture index for the current block based on the motion information; deriving a motion vector for the current block based on the motion information; generating a prediction sample for the current block based on the reference picture index and the motion vector; deriving residual samples based on the prediction samples; generating the bitstream by encoding image information including information about the residual samples; transmitting the data including the bitstream; One or more tiles are in the current picture, the current picture includes a plurality of tile columns and tile rows; A tile is a rectangular region of a coding tree unit (CTU) within a specific tile column and a specific tile row in the current picture, The HMVP candidate list is updated based on the motion information of the previous block; The HMVP candidate list is initialized with a first CTU for each CTU row of each tile; A transmission method, wherein the number of HMVP candidates included in the HMVP candidate list is set to zero based on the HMVP candidate list being initialized.

Citation Information

Patent Citations

  • Resetting of look up table per slice / tile / LCU row

    WO2020003266A1

  • Partial / full pruning when adding a HMVP candidate to merge / amvp

    WO2020003275A1

  • Method and apparatus for history-based motion vector prediction with parallel processing

    WO2020018241A1

  • Method and apparatus for history-based motion vector prediction

    WO2020018297A1