Image decoding method and device for generating prediction sample by applying determined prediction mode

The video decoding method enhances coding efficiency by applying a regular merge mode when default modes are unavailable, effectively compressing high-resolution video data for reduced transmission and storage costs.

JP2025107337APending Publication Date: 2025-07-17LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025076172
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-19
Filing Date
2025-05-01
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, including immersive media like VR and AR, necessitates a highly efficient video compression technology to reduce transmission and storage costs while addressing challenges in deriving prediction samples when default merge modes are unavailable.

Method used

A video decoding method that determines a regular merge mode when MMVD, CIIP, and partitioning modes are not available, using merge index information to generate prediction samples based on a merge candidate list.

Benefits of technology

Improves video coding efficiency by enabling efficient inter prediction even when default merge modes are not selectable, thereby reducing transmission and storage costs for high-resolution video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025107337000001_ABST
    Figure 2025107337000001_ABST
Patent Text Reader

Abstract

To provide an image decoding method and device for generating a prediction sample by applying a determined prediction mode.SOLUTION: When an inter-prediction type for a current block represents bi-prediction, weight index information for a candidate in a merge candidate list or a sub-block merge candidate list can be derived, and thus the coding efficiency can be improved.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to a video decoding method and apparatus for generating a prediction sample by applying a determined prediction mode.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) contents, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] Accordingly, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce information of high-resolution and high-quality images / videos having various characteristics as described above.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technical problem of this document is to provide a method and apparatus for increasing video coding efficiency.

[0006] Another technical problem of this document is to provide a method and apparatus for deriving a prediction sample based on a default merge mode when the final merge mode cannot be selected.

[0007] Still another technical problem of this document is to provide a method and an apparatus for deriving a prediction sample by applying a regular merge mode as a default merge mode.

Means for Solving the Problem

[0008] According to an embodiment of this document, a video decoding method performed by a decoding apparatus is provided. The method includes steps of obtaining video information including inter-prediction mode information and residual information via a bitstream, generating residual samples based on the residual information, applying a prediction mode determined based on the inter-prediction mode information to generate prediction samples of the current block, and generating restored samples based on the prediction samples and the residual samples. The inter-prediction mode information includes a general merge flag indicating whether a merge mode is available for the current block. Based on the general merge flag, if the merge mode is available for the current block and when the MMVD mode (merge mode with motion vector difference), the merge subblock mode, the CIIP mode (combined inter-picture merge and intra-picture prediction mode), and the partitioning mode for predicting by dividing the current block into two partitions are not available, a regular merge mode is determined. The inter-prediction mode information includes merge index information indicating one candidate among merge candidates included in a merge candidate list generated by applying the regular merge mode, and the prediction samples are generated using the merge index information.

[0009] According to another embodiment of the present document, a video encoding method performed by an encoding device is provided. The method includes determining an inter prediction mode of a current block and generating inter prediction mode information representing the inter prediction mode; generating prediction samples of the current block based on the determined prediction mode; generating residual information based on residual samples for the current block; and encoding video information including the inter prediction mode information and the residual information. The inter prediction mode information includes a general merge flag indicating whether a merge mode is available for the current block. When a MMVD mode (merge mode with motion vector difference), a merge subblock mode, a CIIP mode (combined inter-picture merge and intra-picture prediction mode), and a partitioning mode in which the current block is divided into two partitions for prediction are not available, a regular merge mode is determined.

[0010] The inter prediction mode information includes merge index information indicating one candidate among merge candidates included in a merge candidate list generated by applying the regular merge mode.

[0011] According to yet another embodiment of the present document, there is provided a computer-readable digital storage medium storing a bitstream including video information that causes a decoding apparatus to perform a video decoding method. The video decoding method includes steps of: obtaining video information including inter prediction mode information and residual information via a bitstream; generating residual samples based on the residual information; applying a prediction mode determined based on the inter prediction mode information to generate prediction samples of the current block; and generating restored samples based on the prediction samples and the residual samples. The inter prediction mode information includes a general merge flag indicating whether a merge mode is available for the current block. Based on the general merge flag, when the merge mode is available for the current block and the MMVD mode (merge mode with motion vector difference), the merge subblock mode, the CIIP mode (combined inter-picture merge and intra-picture prediction mode), and the partitioning mode for dividing the current block into two partitions for prediction are not available, a regular merge mode is determined.

[0012] The inter prediction mode information includes merge index information indicating one candidate among merge candidates included in a merge candidate list generated by applying the regular merge mode, and the prediction samples are generated using the merge index information.

Advantages of the Invention

[0013] According to the present document, general video / video compression efficiency can be improved.

[0014] According to this document, when the merge mode cannot be finally selected, efficient inter prediction can be performed by applying the default merge mode.

[0015] According to this document, when the merge mode cannot be finally selected, regular merge mode is applied, and motion information is derived based on candidates pointed to by merge index information, so that efficient inter prediction can be performed.

Brief Description of the Drawings

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10a

Figure 10b

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Mode for Carrying Out the Invention

[0017] The disclosure of this document can be modified in various ways and can have various embodiments. Specific embodiments are illustrated in the drawings and will be described in detail. However, this is not intended to limit the present disclosure to specific embodiments. The terms used in this document are merely used to describe specific embodiments and are not intended to limit the technical concept of the embodiments of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this document are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the document, and it should be understood that the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is not precluded in advance.

[0018] On the other hand, each configuration in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is implemented by separate hardware or separate software. For example, among the configurations, two or more configurations can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are included in the scope of rights of this document as long as they do not deviate from the essence of this document.

[0019] This document relates to video / video coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (versatile video coding) standard. Also, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268, etc.).

[0020] In this document, various embodiments related to video / video coding are presented, and unless otherwise stated, the embodiments can be combined with each other.

[0021] Hereinafter, preferred embodiments of this document will be described with reference to the accompanying drawings. Hereinafter, the same reference numerals are used for the same components in the drawings, and duplicate descriptions of the same components can be omitted.

[0022] FIG. 1 schematically shows an example of a video / image coding system to which the present disclosure can be applied.

[0023] As shown in FIG. 1, the video / image coding system can include a first device (source device) and a second device (receiver device). The source device can transmit encoded video / image information or data in a file or streaming form to the receiver device via a digital storage medium or a network.

[0024] The source device can include a video source, an encoding device, and a transmitting unit. The receiver device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. The transmitter can be included in the encoding device. The receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.

[0025] The video source can obtain video / images through processes such as capture, synthesis, or generation of video / images. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can generate (electronically) video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.

[0026] The encoding device can encode the input video / image. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0027] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0028] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0029] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0030] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).

[0031] This document presents various embodiments related to video / image coding. Unless otherwise stated, the embodiments can also be executed in combination with each other.

[0032] In this document, "video" can mean a collection of a series of images over time. "Picture" generally means a unit representing one image at a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (coding tree units). One picture can be composed of one or more slices / titles.

[0033] A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan can represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice can contain a plurality of consecutive CTU rows within one tile of a picture which can be contained in a plurality of complete tiles or one NAL unit. In this document, tile group and slice can be used interchangeably. For example, in this document, tile group / tile group header can be called slice / slice header.

[0034] On the other hand, one picture can be partitioned into two or more sub-pictures. A sub-picture can be an rectangular region of one or more slices within a picture.

[0035] A pixel or pel can mean the smallest unit that makes up one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, can indicate only the pixel / pixel value of the luma component, or can indicate only the pixel / pixel value of the chroma component. Or, a sample can mean a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can mean a conversion coefficient in the frequency domain.

[0036] A unit can indicate the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to that region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can include a set (or array) of samples (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.

[0037] In this document, "A or B" can mean "only A", "only B", or "both A and B". In other words, in this document, "A or B" can be interpreted as "A and / or B". For example, in this document, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".

[0038] The slashes ( / ) or commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".

[0039] In this document, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this document, expressions such as "at least one of A or B" or "at least one of A and / or B" can be analyzed in the same way as "at least one of A and B".

[0040] Also, in this document, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0041] Also, the parentheses used in this document can mean "for example". Specifically, when it is shown as "prediction (intra prediction)", it may be that "intra prediction" is proposed as an example of "prediction". In other words, "prediction" in this document is not limited to "intra prediction", and it may be that "intra prediction" is proposed as an example of "prediction". Also, when it is shown as "prediction (i.e., intra prediction)", it may be that "intra prediction" is proposed as an example of "prediction".

[0042] The technical features individually described within one drawing in this document can be implemented individually or simultaneously.

[0043] FIG. 2 is a diagram schematically explaining the configuration of a video / image encoding apparatus to which the present disclosure can be applied. Hereinafter, the video encoding apparatus can include an image encoding apparatus.

[0044] As shown in FIG. 2, the encoding apparatus 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor (231). The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0045] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that a coding unit of an optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the above-described final coding unit.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0046] The unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can represent a set of samples or transform coefficients, etc., consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).

[0047] The encoding device 200 can subtract the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoding device 200 can be called the subtraction unit 231. The prediction unit can perform a prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0048] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the vicinity (neighbor) of the current block depending on the prediction mode, or can also be located far away. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is an example, and a greater or lesser number of directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the surrounding blocks.

[0049] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (col CU), etc., and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on the peripheral block, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, different from the merge mode, a residual signal may not be transmitted.In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0050] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index.

[0051] The prediction signal generated via the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when representing the relationship information between pixels as a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to pixel blocks having the same size of a square and can also be applied to blocks of variable size that are not square.

[0052] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can encode, together with or separately from the quantized transform coefficients, information necessary for video / image restoration (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / image information. The video / image information can be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be provided in the entropy encoding unit 240.

[0053] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.

[0054] On the other hand, LMCS (luma mapping with chrom ascaling) can also be applied during the picture encoding and / or restoration process.

[0055] The filtering unit 260 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0056] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. Through this, when inter prediction is applied, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.

[0057] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks in which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatial neighboring blocks or temporal neighboring blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0058] On the one hand, in this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation can be omitted. When the quantization / inverse quantization is omitted, the quantized transformation coefficient can be called a transformation coefficient. When the transformation / inverse transformation is omitted, the transformation coefficient can be called a coefficient or a residual coefficient, or, for the sake of uniformity of expression, can still be called a transformation coefficient.

[0059] Also, in this document, the quantized transformation coefficient and the transformation coefficient can be respectively referred to as a transformation coefficient and a scaled transformation coefficient. In this case, the residual information can include information about the transformation coefficient (etc.), and the information about the transformation coefficient (etc.) can be signaled via a residual coding syntax. The transformation coefficient can be derived based on the residual information (or the information about the transformation coefficient (etc.)), and the scaled transformation coefficient can be derived via an inverse transformation (scaling) for the transformation coefficient. The residual sample can be derived based on the inverse transformation (transformation) for the scaled transformation coefficient. This can be similarly applied / expressed in other parts of this document.

[0060] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which the present disclosure can be applied.

[0061] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filtering unit 350, and a memory 360. The predictor 330 can include an intra predictor 331 and an inter predictor 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 322. The above-described entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 can be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0062] If a bitstream including video / image information is input, the decoding apparatus 300 can restore an image corresponding to the process in which the video / image information is processed by the encoding apparatus of FIG. 2. For example, the decoding apparatus 300 can derive units / blocks based on block division related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using the processing unit applied in the encoding apparatus. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided according to a quad tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding apparatus 300 can be reproduced via a reproducing apparatus.

[0063] The decoding device 300 can receive the signal output from the encoding device of FIG. 3 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the information of the surrounding and decoded blocks to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin according to the determined context model, performs arithmetic decoding of the bin, and can generate the symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction units (inter prediction unit 332 and intra prediction unit 331), and the residual value obtained by performing entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.

[0064] In the inverse quantization unit 321, the quantized transform coefficient can be inverse quantized to output a transform coefficient. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficient using a quantization parameter (for example, quantization step size information) to obtain a transform coefficient.

[0065] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0066] The prediction unit can perform prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0067] The prediction unit 330 can generate a prediction signal based on various prediction methods to be described later. For example, for prediction of one block, the prediction unit can apply not only intra prediction or inter prediction but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, for prediction of a block, the prediction unit can be based on the intra block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included and signaled in the video / image information.

[0068] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred can be located at the periphery (neighbor) of the current block depending on the prediction mode, or can be located remotely. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The intra prediction unit 331 can also utilize the prediction mode applied to neighboring blocks to determine the prediction mode to be applied to the current block.

[0069] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub - blocks or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be executed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0070] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0071] The adder 340 can be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed within the current picture, and as will be described later, it can also be output after filtering, or it can be used for inter prediction of the next picture.

[0072] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0073] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be sent to the memory 360, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0074] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of the spatially neighboring blocks or the temporally neighboring blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.

[0075] In this specification, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 200 can be applied to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300 in the same or corresponding manner, respectively.

[0076] On the other hand, as described above, prediction is performed to improve the compression efficiency when performing video coding. Thereby, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same manner in the encoding device and the decoding device, and the encoding device can improve the image coding efficiency by signaling information (residual information) regarding the residual between the original block, which is not the original sample value of the original block itself, and the predicted block to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a restored block including restored samples, and generate a restored picture including the restored block.

[0077] The residual information can be generated through conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and thus signal the relevant residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter prediction of subsequent pictures to derive a residual block, and generate a restored picture based on this.

[0078] On the one hand, for the prediction of the current block in a picture, various inter-prediction modes can be used. For example, various prediction modes such as merge mode, skip mode, MVP (motion vector prediction) mode, Affine mode, sub-block merge mode, MMVD (merge with MVD) mode, etc. can be used. DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, Bi-prediction with CU-level weight (BCW), Bi-directional optical flow (BDOF), etc. can be used as accompanying modes or can be used instead. The Affine mode can also be called the affine motion prediction mode. The MVP mode can also be called the AMVP (advanced motion vector prediction) mode. In this document, some motion information candidates derived by some modes and / or some modes can also be included as one of the motion information related candidates of other modes. For example, the HMVP candidate can be added as a merge candidate in the merge / skip mode or can be added as an mvp candidate in the MVP mode.

[0079] Inter prediction mode information indicating the current block's inter prediction mode can be signaled from an encoding device to a decoding device. The inter prediction mode information can be included in a bitstream and received by the decoding device. The inter prediction mode information can include index information indicating one of a number of candidate modes. Alternatively, the inter prediction mode can also be indicated via hierarchical signaling of flag information. In this case, the inter prediction mode information can include one or more flags. For example, a skip flag can be signaled to indicate whether to apply the skip mode, and when the skip mode is not applied, a merge flag can be signaled to indicate whether to apply the merge mode. When the merge mode is not applied, it can be indicated that the MVP mode is applied, or flags for additional classification can be further signaled. The affine mode can also be signaled as an independent mode, or can be signaled as a mode subordinate to the merge mode or the MVP mode, etc. For example, the affine mode can include an affine merge mode and an affine MVP mode.

[0080] On the one hand, information indicating whether list 0 (L0) prediction, list 1 (L1) prediction, or bi - prediction is used for the current block (current coding unit) can be signaled. The said information can be referred to as motion prediction direction information, inter - prediction direction information, or inter - prediction indication information, and can be configured / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether the above - mentioned list 0 (L0) prediction, list 1 (L1) prediction, or bi - prediction is used for the current block (current coding unit). In this document, for the convenience of explanation, the inter - prediction type (L0 prediction, L1 prediction, or BI prediction) indicated by the inter_pred_idc syntax element can be expressed as the motion prediction direction. L0 prediction can also be represented by pred_L0, L1 prediction by pred_L1, and bi - prediction by pred_BI. For example, the following prediction types can be represented by the value of the inter_pred_idc syntax element.

[0081] As described above, one picture can include one or more slices. A slice can have one type among the slice types including intra (I) slice, predictive (P) slice, and bi-predictive (B) slice. The slice type can be indicated based on slice type information. For blocks within an I slice, inter prediction is not used for prediction, and only intra prediction can be used. Of course, in this case, it is also possible to code and signal the original sample value without prediction. For blocks within a P slice, intra prediction or inter prediction can be used, and when inter prediction is used, only uni prediction can be used. On the other hand, for blocks within a B slice, intra prediction or inter prediction can be used, and when inter prediction is used, up to maximum bi prediction can be used.

[0082] L0 and L1 can include reference pictures that have been encoded / decoded before the current picture. For example, L0 can include reference pictures before and / or after the current picture in POC order, and L1 can include reference pictures after and / or before the current picture in POC order. In this case, a relatively lower reference picture index can be assigned to L0 for reference pictures before the current picture in POC order, and a relatively lower reference picture index can be assigned to L1 for reference pictures after the current picture in POC order. In the case of a B slice, bi prediction can be applied, and in this case as well, uni-directional bi prediction can be applied, or both-directional bi prediction can be applied. Both-directional bi prediction can be called true bi prediction.

[0083] For example, information regarding the inter prediction mode of the current block can be coded and signaled at a level such as CU (CU syntax), or can be implicitly determined depending on conditions. In this case, for some modes, they are explicitly signaled, and for some of the remaining modes, they can be implicitly derived.

[0084] For example, the CU syntax can carry information regarding the (inter) prediction mode as follows. The CU syntax can be as shown in Table 1 below.

[0085] [Table 1-1]

[0086] [Table 1-2]

[0087] [Table 1-3]

[0088] [Table 1-4]

[0089] [Table 1-5]

[0090] [Table 1-6]

[0091] [Table 1-7]

[0092] [Table 1-8]

[0093]

Table 1-9

[0094]

Table 1-10

[0095] In Table 1 above, cu_skip_flag can indicate whether the skip mode is applied to the current block (CU).

[0096] When pred_mode_flag is 0, it can be specified that the current block is coded in the inter prediction mode, and when pred_mode_flag is 1, it can be specified that the current coding unit is coded in the intra prediction mode.

[0097] When pred_mode_ibc_flag is 1, it can be specified that the current block is coded in the IBC prediction mode, and when pred_mode_ibc_flag is 0, it can be specified that the current block (CU) is not coded in the IBC prediction mode.

[0098] Also, if pcm_flag[x0][y0] is 1, it can be specified that the pcm_sample() syntax structure exists and the transform_tree() syntax structure does not exist in the current block including the luma coding block at position (x0, y0). If pcm_flag[x0][y0] is the same as 0, it can be specified that the pcm_sample() syntax structure does not exist. That is, pcm_flag can represent whether the puls coding modulation (PCM) mode is applied to the current block. When the PCM mode is applied to the current block, prediction / transformation / quantization, etc. are not applied, and the values of the original samples in the current block can be coded and signaled.

[0099] Also, if intra_mip_flag[x0][y0] is 1, it can be specified that the intra prediction type for the luma sample is the matrix-based intra prediction (MIP), and if intra_mip_flag[x0][y0] is 0, it can be specified that the intra prediction type for the luma sample is not the matrix-based intra prediction. That is, intra_mip_flag can represent whether the MIP prediction mode (type) is applied to the luma samples of the current block.

[0100] intra_chroma_pred_mode[x0][y0] can specify the intra prediction mode for the chroma samples from the current block.

[0101] general_merge_flag[x0][y0] can specify whether the inter-prediction parameter for the current block is inferred from an adjacent inter-predicted partition. That is, general_merge_flag can indicate that the general merge mode is available. For example, when the value of general_merge_flag is 1, regular merge mode, MMVD (merge mode with motion vector difference) mode, and merge subblock mode (subblock merge mode) can be used. For example, when the value of general_merge_flag is 1, the merge data syntax can be parsed from the encoded video / image information (or bitstream), and the merge data syntax can be configured / coded as shown in Table 2 below.

[0102]

Table 2-1

[0103]

Table 2-2

[0104] In Table 2 above, if regular_merge_flag[x0][y0] is 1, it can be specified that the regular merge mode is used to generate the inter-prediction parameter of the current block. That is, regular_merge_flag can indicate whether the merge mode (regular merge mode) is applied to the current block.

[0105] If mmvd_merge_flag[x0][y0] is 1, it can be specified that the MMVD mode (merge mode with motion vector difference mode) is used to generate the inter prediction parameters of the current block. That is, mmvd_merge_flag can indicate whether MMVD is applied to the current block.

[0106] mmvd_cand_flag[x0][y0] can specify whether to use the first (0) or second (1) candidate in the merge candidate list together with the motion vector difference value derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0].

[0107] mmvd_distance_idx[x0][y0] can specify the index used to derive MmvdDistance[x0][y0].

[0108] mmvd_direction_idx[x0][y0] can specify the index used to derive MmvdSign[x0][y0].

[0109] merge_subblock_flag[x0][y0] can specify the sub-block based inter prediction parameters for the current block. That is, merge_subblock_flag can indicate whether the sub-block merge mode (or affine merge mode) is applied to the current block.

[0110] merge_subblock_idx[x0][y0] can specify the merge candidate index from the sub-block based merge candidate list.

[0111] ciip_flag[x0][y0] can specify whether CIIP (combined inter-picture merge and intra-picture prediction) prediction is applied to the current block.

[0112] merge_triangle_idx0[x0][y0] can specify the first merge candidate index of the triangular shape based motion compensation candidate list.

[0113] merge_triangle_idx1[x0][y0] can specify the second merge candidate index of the triangular shape based motion compensation candidate list.

[0114] merge_idx[x0][y0] can specify the merge candidate index in the merge candidate list.

[0115] On the other hand, referring to the CU syntax again, mvp_l0_flag[x0][y0] can specify the motion vector predictor index in list 0. That is, mvp_l0_flag can indicate the candidate selected for the MVP derivation of the current block in the MVP candidate list 0 when the MVP mode is applied.

[0116] mvp_l1_flag[x0][y0] has the same meaning as mvp_l0_flag, and l0 and list 0 can be replaced by l1 and list 1 respectively.

[0117] inter_pred_idc[x0][y0] can specify whether list 0, list 1 or bi-prediction is used in the current coding unit.

[0118] If sym_mvd_flag[x0][y0] is 1, it can be specified that there is no syntax structure of mvd_coding(x0,y0,refList,cpIdx) for the syntax elements ref_idx_l0[x0][y0], ref_idx_l1[x0][y0], and refList 1. That is, sym_mvd_flag can indicate whether symmetric MVD is used in mvd coding.

[0119] "ref_idx_l0[x0][y0]" can specify the list 0 reference picture index for the current block.

[0120] "ref_idx_l1[x0][y0]" has the same meaning as ref_idx_l0, and l0, L0, and list 0 can be replaced by l1, L1, and list 1 respectively.

[0121] If inter_affine_flag[x0][y0] is 1, when decoding a P or B slice, it can be specified that affine model based motion compensation is used to generate the predicted samples of the current block.

[0122] When cu_affine_type_flag[x0][y0] is 1, when decoding a P or B slice, 6-parameter affine model based motion compensation can be specified to be used to generate the prediction samples of the current block. When cu_affine_type_flag[x0][y0] is 0, when decoding a P or B slice, 4-parameter affine model based motion compensation can be specified to be used to generate the prediction samples of the current block.

[0123] amvr_flag[x0][y0] can specify the resolution of the motion vector difference value. The array indices x0, y0 can specify the position (x0, y0) of the top-left luma sample of the coding block considered with respect to the top-left luma sample of the picture. When amvr_flag[x0][y0] is 0, it can be specified that the resolution of the motion vector difference value is 1 / 4 of the luma sample. When amvr_flag[x0][y0] is 1, the resolution of the motion vector difference value can be further specified by amvr_precision_flag[x0][y0].

[0124] When amvr_precision_flag[x0][y0] is 0, when inter_affine_flag[x0][y0] is 0, the resolution of the motion vector difference value becomes 1 integer luma sample, and otherwise, it can be specified to be 1 / 16 of the luma sample. When amvr_precision_flag[x0][y0] is 1, when inter_affine_flag[x0][y0] is 0, the resolution of the motion vector difference value becomes 4 luma samples, and otherwise, it can be specified to be 1 integer luma sample.

[0125] bcw_idx[x0][y0] can specify a weight index for bi-prediction using CU weight values.

[0126] FIG. 4 shows an example of a video / video encoding method based on inter prediction, and FIG. 5 shows an example schematically showing an inter prediction unit in an encoding device. The inter prediction unit in the encoding device of FIG. 5 can be applied to be identical or corresponding to the inter prediction unit 221 of the encoding device 200 of FIG. 2 described above.

[0127] As shown in FIGS. 4 and 5, the encoding device performs inter prediction on the current block (S400). The encoding device can derive an inter prediction mode and motion information of the current block and generate a predicted sample of the current block. Here, the inter prediction mode determination, motion information derivation, and predicted sample generation procedures can be performed simultaneously, or any one of the procedures can be performed first than the other procedures.

[0128] For example, the inter prediction unit 221 of the encoding device may include a prediction mode determination unit 221_1, a motion information derivation unit 221_2, and a prediction sample derivation unit 221_3. The prediction mode determination unit 221_1 determines the prediction mode for the current block, the motion information derivation unit 221_2 derives the motion information of the current block, and the prediction sample derivation unit 221_3 can derive the prediction sample of the current block. For example, the inter prediction unit 221 of the encoding device can search for a block similar to the current block within a certain region (search region) of the reference picture via motion estimation, and derive a reference block whose difference from the current block is the smallest or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode applied to the current block among various prediction modes. The encoding device can compare the RD costs for the various prediction modes and determine the optimal prediction mode for the current block.

[0129] For example, when the skip mode or the merge mode is applied to the current block, the encoding device configures a merge candidate list described later, and among the reference blocks pointed to by the merge candidates included in the merge candidate list, a reference block whose difference from the current block is the smallest or below a certain criterion can be derived. In this case, the merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0130] As another example, when the (A)MVP mode is applied to the current block, the encoding device may construct an (A)MVP candidate list, and among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list, the motion vector of the selected mvp candidate may be used as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the above-described motion estimation may be used as the motion vector of the current block, and among the mvp candidates, the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block may be the selected mvp candidate. An MVD (motion vector difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index may be composed of reference picture index information and may be signaled to the decoding device separately.

[0131] The encoding device may derive a residual sample based on the prediction sample (S410). The encoding device may derive the residual sample by comparing the original sample of the current block with the prediction sample.

[0132] The encoding device encodes video information including prediction information and residual information (S420). The encoding device can output the encoded video information in the form of a bitstream. The prediction information can include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information related to motion information as information related to the prediction procedure. The information related to motion information can include candidate selection information (e.g., merge index) which is information for deriving a motion vector. Also, the information related to motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information related to residual samples. The residual information can include information related to quantized transform coefficients for the residual samples.

[0133] The output bitstream can be stored in a (digital) storage medium and transmitted to the decoding device, or can also be transmitted to the decoding device via a network.

[0134] Also, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to derive the same prediction result in the encoding device as that performed in the decoding device, and through this, the coding efficiency can be improved. Therefore, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in the memory and utilize it as a reference picture for inter prediction. As described above, an in-loop filtering procedure or the like can be further applied to the reconstructed picture.

[0135] FIG. 6 shows an example of a video / video decoding method based on inter prediction, and FIG. 7 shows an example schematically showing an inter prediction unit in a decoding apparatus. The inter prediction unit in the decoding apparatus of FIG. 7 can be applied so as to be identical or corresponding to the inter prediction unit 332 of the decoding apparatus 300 of FIG. 3 described above.

[0136] As shown in FIGS. 6 and 7, the decoding apparatus can perform an operation corresponding to the operation performed by the encoding apparatus. The decoding apparatus can perform prediction on the current block based on the received prediction information and derive a prediction sample.

[0137] Specifically, the decoding apparatus can determine a prediction mode for the current block based on the received prediction information (S600). The decoding apparatus can determine which inter prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0138] The inter prediction mode candidates can include a skip mode, a merge mode, and / or an (A) MVP mode, or can include various inter prediction modes.

[0139] The decoding apparatus derives motion information of the current block based on the determined inter prediction mode (S610). For example, when the skip mode or the merge mode is applied to the current block, the decoding apparatus can configure a merge candidate list and select any one of the merge candidates included in the merge candidate list. Here, the selection can be performed based on the above-described selection information (merge index). The motion information of the current block can be derived using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0140] As another example, when the (A)MVP mode is applied to the current block, the decoding device constructs an (A)MVP candidate list, and among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list, the motion vector of the selected mvp candidate can be used as the mvp of the current block. Here, the selection can be performed based on the above-described selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the mvp and MVD of the current block. Also, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index within the reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.

[0141] On the other hand, the motion information of the current block can be derived without constructing a candidate list. In this case, the motion information of the current block can be derived according to the procedure disclosed in the prediction mode. In this case, the candidate list construction as described above can be omitted.

[0142] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S620). In this case, the reference picture is derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, the prediction sample filtering procedure can be further performed on all or part of the prediction samples of the current block depending on the case.

[0143] For example, the inter prediction unit 332 of the decoding apparatus may include a prediction mode determination unit 332_1, a motion information derivation unit 332_2, and a prediction sample derivation unit 332_3. The prediction mode determination unit 332_1 determines a prediction mode for the current block based on the received prediction mode information, the motion information derivation unit 332_1 derives motion information (such as a motion vector and / or a reference picture index) of the current block based on information related to the received motion information, and the prediction sample derivation unit 332_3 can derive a prediction sample of the current block.

[0144] The decoding apparatus generates a residual sample for the current block based on the received residual information (S630). The decoding apparatus generates a restored sample for the current block based on the prediction sample and the residual sample, and can generate a restored picture based on this (S640). As described above, an in-loop filtering procedure or the like can be further applied to the restored picture later.

[0145] As described above, the inter prediction procedure can include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The inter prediction procedure can be performed in the encoding apparatus and the decoding apparatus as described above.

[0146] On the other hand, in deriving the motion information of the current block, motion information candidate(s) can be derived based on spatial adjacent block(s) and temporal adjacent block(s), and a motion information candidate for the current block can be selected based on the derived motion information candidate(s). At this time, the selected motion information candidate can be used as the motion information of the current block.

[0147] FIG. 8 is a diagram for explaining the merge mode in inter prediction.

[0148] When the merge mode is applied, the motion information of the current prediction block is not directly transmitted, but the motion information of the current prediction block is induced using the motion information of the adjacent prediction block. Therefore, by transmitting flag information indicating that the merge mode has been used and a merge index indicating which adjacent prediction block has been used, the motion information of the current prediction block can be indicated. The merge mode can be called a regular merge mode.

[0149] The encoding device has to search for merge candidate blocks used to derive the motion information of the current prediction block in order to perform the merge mode. For example, up to five of the merge candidate blocks can be used, but the embodiments (etc.) of this document are not limited to this. And the maximum number of the merge candidate blocks can be transmitted in the slice header or the tile group header, but the embodiments (etc.) of this document are not limited to this. After searching for the merge candidate blocks, the encoding device can generate a merge candidate list, and among these, the merge candidate block having the smallest cost can be selected as the final merge candidate block.

[0150] This document can provide various embodiments for the merge candidate blocks constituting the merge candidate list.

[0151] For example, the merge candidate list can use five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate can be utilized. As a specific example, in the case of a spatial merge candidate, the block shown in FIG. 4 can be used as the spatial merge candidate. Hereinafter, the spatial merge candidate or the spatial MVP candidate described later can be called an SMVP, and the temporal merge candidate or the temporal MVP candidate described later can be called a TMVP.

[0152] The merge candidate list for the current block can be configured based on, for example, the following procedure.

[0153] The coding device (encoding device / decoding device) can insert the spatial merge candidates derived by searching for the spatial adjacent blocks of the current block into the merge candidate list. For example, the spatial adjacent blocks can include the lower left corner adjacent block, the left adjacent block, the upper right corner adjacent block, the upper adjacent block, and the upper left corner adjacent block of the current block. However, this is only an example, and in addition to the spatial adjacent blocks described above, additional adjacent blocks such as the right adjacent block, the lower adjacent block, and the lower right corner adjacent block can also be used as the spatial adjacent blocks. The coding device can search for the spatial adjacent blocks based on the priority, detect the available blocks, and derive the motion information of the detected blocks as the spatial merge candidates. For example, the encoding device or the decoding device can search for the five blocks shown in FIG. 8 in the order of A1->B1->B0->A0->B2, index the available candidates sequentially, and configure them in the merge candidate list.

[0154] The coding device can search for the temporal neighboring blocks of the current block and insert the derived temporal merge candidates into the merge candidate list. The temporal neighboring blocks can be located on a reference picture that is a picture different from the current picture in which the current block is located. The reference picture on which the temporal neighboring blocks are located can be called a collocated picture or a col picture. The temporal neighboring blocks can be searched in the order of the peripheral blocks of the lower right corner and the lower right center block of the co-located block with respect to the current block on the col picture. On the other hand, when motion data compression is applied, specific motion information can be stored as representative motion information for each fixed storage unit in the col picture. In this case, it is not necessary to store the motion information for all blocks within the fixed storage unit, and the motion data compression effect can be obtained through this. In this case, the fixed storage unit can be determined in advance, for example, in units of 16×16 samples or 8×8 samples, or the size information for the fixed storage unit can also be signaled from the encoding device to the decoding device. When motion data compression is applied, the motion information of the temporal neighboring blocks can be replaced with the representative motion information of the fixed storage unit in which the temporal neighboring blocks are located. That is, in this case, from the perspective of implementation, based on the motion information of the prediction block that covers the position arithmetically left-shifted after arithmetically right-shifting by a certain value based on the coordinates (upper left sample position) of the temporal neighboring blocks, which are not the prediction blocks located at the coordinates of the temporal neighboring blocks, the temporal merge candidates can be derived.For example, when the fixed storage unit is a 2n×2n sample unit, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>n)<<n), (yTnb>>n)<<n)) can be used for the temporal merge candidate. Specifically, for example, when the fixed storage unit is a 16×16 sample unit, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>4)<<4), (yTnb>>4)<<4)) can be used for the temporal merge candidate. Alternatively, for example, when the fixed storage unit is an 8×8 sample unit, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>3)<<3), (yTnb>>3)<<3)) can be used for the temporal merge candidate.

[0155] The coding device can check whether the number of current merge candidates is smaller than the number of maximum merge candidates. The number of the maximum merge candidates can be predefined or signaled from the encoding device to the decoding device. For example, the encoding device can generate information regarding the number of the maximum merge candidates, encode it, and transmit it to the decoder in the form of a bitstream. If all the numbers of the maximum merge candidates are filled, the subsequent candidate addition process cannot proceed.

[0156] If the number of the current merge candidates is less than the maximum number of merge candidates according to the confirmation result, the coding device can insert additional merge candidates into the merge candidate list. For example, the additional merge candidates can include at least one of history based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate (when the slice / tile group type of the current slice / tile group is type B), and / or zero vector merge candidate.

[0157] If the number of the current merge candidates is not less than the maximum number of merge candidates according to the confirmation result, the coding device can end the configuration of the merge candidate list. In this case, the encoding device can select an optimal merge candidate from the merge candidates that configure the merge candidate list based on rate-distortion (RD) cost, and can signal selection information (e.g., merge index) indicating the selected merge candidate to the decoding device. The decoding device can select the optimal merge candidate based on the merge candidate list and the selection information.

[0158] As described above, the motion information of the selected merge candidate can be used as the motion information of the current block, and the predicted sample of the current block can be derived based on the motion information of the current block. The encoding device can derive the residual sample of the current block based on the predicted sample, and can signal the residual information regarding the residual sample to the decoding device. As described above, the decoding device can generate a restored sample based on the residual sample derived based on the residual information and the predicted sample, and can generate a restored picture based on this.

[0159] When the skip mode is applied, the motion information of the current block can be derived in the same way as when the merge mode is applied above. However, when the skip mode is applied, the residual signal for the block is omitted, and thus the predicted sample can be immediately used as the restored sample. The skip mode can be applied, for example, when the value of the cu_skip_flag syntax element is 1.

[0160] FIG. 9 is a diagram for explaining the MMVD mode (merge mode with motion vector difference mode) in inter prediction.

[0161] The MMVD mode is a method of applying MVD (motion vector difference) to the merge mode in which the motion information derived to generate the predicted sample of the current block is directly used.

[0162] For example, an MMVD flag (e.g., mmvd_flag) indicating whether to use MMVD for the current block (i.e., the current CU) can be signaled, and MMVD can be performed based on this MMVD flag. When MMVD is applied to the current block (e.g., when mmvd_flag is 1), additional information for MMVD can be signaled.

[0163] Here, the additional information for MMVD can include a merge candidate flag (e.g., mmvd_cand_flag) indicating whether the first candidate or the second candidate in the merge candidate list is used with the MVD, a distance index (e.g., mmvd_distance_idx) for representing the motion magnitude, and a direction index (e.g., mmvd_direction_idx) for representing the motion direction.

[0164] In the MMVD mode, among the candidates in the merge candidate list, two candidates located at the first and second entries (i.e., the first candidate or the second candidate) can be used, and any one of the two candidates (i.e., the first candidate or the second candidate) can be used as the base MV. For example, a merge candidate flag (e.g., mmvd_cand_flag) can be signaled to represent any one of the two candidates (i.e., the first candidate or the second candidate) in the merge candidate list.

[0165] Also, a distance index (e.g., mmvd_distance_idx) represents motion magnitude information and can indicate a predetermined offset from the starting point. Referring to FIG. 9, the offset can be added to the horizontal or vertical component of the starting motion vector. The relationship between the distance index and the predetermined offset can be represented as shown in Table 3 below.

[0166]

Table 3

[0167] As shown in Table 3 above, the MVD distance (e.g., MmvdDistance) is determined according to the value of the distance index (e.g., mmvd_distance_idx), and the MVD distance (e.g., MmvdDistance) can be derived using integer sample precision or fractional sample precision based on the value of slice_fpel_mmvd_enabled_flag. For example, when slice_fpel_mmvd_enabled_flag is 1, it means that the MVD distance is derived using integer sample precision in the current slice, and when slice_fpel_mmvd_enabled_flag is 0, it can mean that the MVD distance is derived using fractional sample precision in the current slice.

[0168] In addition, the direction index (e.g., mmvd_direction_idx) represents the direction of the MVD based on the starting point and can represent four directions as shown in Table 4 below. At this time, the direction of the MVD can represent the sign of the MVD. The relationship between the direction index and the MVD sign can be represented as shown in Table 4 below.

[0169]

Table 4

[0170] As shown in Table 4 above, the sign of the MVD (e.g., MmvdSign) is determined according to the value of the direction index (e.g., mmvd_direction_idx), and the sign of the MVD (e.g., MmvdSign) can be derived for the L0 reference picture and the L1 reference picture.

[0171] Based on the distance index (e.g., mmvd_distance_idx) and the direction index (e.g., mmvd_direction_idx) as described above, the offset of the MVD can be calculated as shown in the following Equation 1.

[0172]

Equation

[0173] That is, in the MMVD mode, among the merge candidates in the merge candidate list derived based on adjacent blocks, the merge candidate indicated by the merge candidate flag (e.g., mmvd_cand_flag) is selected, and the selected merge candidate can be used as the base candidate (e.g., MVP). Then, the motion information (i.e., the motion vector) of the current block can be derived by adding the MVD derived using the distance index (e.g., mmvd_distance_idx) and the direction index (e.g., mmvd_direction_idx) based on the base candidate.

[0174] FIG. 10a and FIG. 10b exemplarily show CPMV for affine motion prediction.

[0175] Conventionally, only one motion vector could be used to represent the motion of a coding block. That is, a translation motion model could be used. However, although such a method might represent the optimal motion in block units, coding efficiency could be improved if the optimal motion vector could be determined in sample units rather than the optimal motion of each sample. For this purpose, an affine motion model can be used. The affine motion prediction method for coding using the affine motion model could be as follows.

[0176] The affine motion prediction method can represent the motion vector for each sample unit of a block by using two, three, or four motion vectors. For example, the affine motion model can represent four types of motion. Among the motions that the affine motion model can represent, the affine motion model that represents three types of motion (translation, scale, rotate) can be called a similarity (or simplified) affine motion model. However, the affine motion model is not limited to the motion models described above.

[0177] Affine motion prediction can determine the motion vector of the sample position included in a block by using two or more control point motion vectors (CPMV). At this time, the set of motion vectors can be represented as an affine motion vector field (MVF).

[0178] For example, FIG. 6a can represent the case where two CPMVs are used, which can be called a 4-parameter affine model. In this case, the motion vector at the (x, y) sample position can be determined as shown in, for example, Equation 2.

[0179] For example, FIG. 10a can show the case where two CPMVs are used, which can be called a 4-parameter affine model. In this case, the motion vector at the (x, y) sample position can be determined as shown in, for example, Equation 2.

[0180] [Number]

[0181] For example, FIG. 10b can show the case where three CPMVs are used, which can be called a 6-parameter affine model. In this case, the motion vector at the (x, y) sample position can be determined as shown in, for example, Equation 3.

[0182] [Number]

[0183] In Equations 2 and 3, {v x , v y} can represent the motion vector at the (x, y) position. Also, {v 0x , v 0y} can represent the CPMV of the control point (CP: Control Point) at the upper left corner position of the coding block, {v 1x , v 1y} can represent the CPMV of the CP at the upper right corner position, and {v 2x , v 2y} can represent the CPMV of the CP at the lower left corner position. Also, W can represent the width of the current block, and H can represent the height of the current block.

[0184] FIG. 11 exemplarily shows the case where the affine MVF is determined in sub-block units.

[0185] In the encoding / decoding process, the affine MVF can be determined in sample units or in predefined sub-block units. For example, when determining in sample units, a motion vector can be obtained based on each sample value. Or, for example, when determining in sub-block units, among the sub-blocks, a motion vector of the corresponding block can be obtained based on the central (center lower right side, that is, the lower right sample among the central 4 samples) sample value. That is, in affine motion prediction, the motion vector of the current block can be derived in sample units or sub-block units.

[0186] In the case of FIG. 11, the affine MVF is determined in 4x4 sub-block units, but the size of the sub-block can be variously deformed.

[0187] That is, when affine prediction is available, the motion models applicable to the current block can include three types (a translational motion model, a 4-parameter affine motion model, and a 6-parameter affine motion model). Here, the translational motion model can represent a model in which a conventional block unit motion vector is used, the 4-parameter affine motion model can represent a model in which 2 CPMVs are used, and the 6-parameter affine motion model can represent a model in which 3 CPMVs are used.

[0188] On the other hand, affine motion prediction can include an affine MVP (or affine inter) mode or an affine merge mode.

[0189] FIG. 12 is a diagram for explaining an affine merge mode or a sub-block merge mode in inter prediction.

[0190] For example, in the affine merge mode, the CPMV can be determined by an affine motion model of neighboring blocks coded with affine motion prediction. For example, neighboring blocks coded with affine motion prediction in the search order can be used for the affine merge mode. That is, when at least one of the neighboring blocks is coded with affine motion prediction, the current block can be coded in the affine merge mode. Here, the affine merge mode can be called AF_MERGE.

[0191] When the affine merge mode is applied, the CPMV of the current block can be derived using the CPMV of the neighboring blocks. In this case, the CPMV of the neighboring blocks can be used as the CPMV of the current block as it is, or the CPMV of the neighboring blocks can be modified based on the size of the neighboring blocks and the size of the current block, etc., and used as the CPMV of the current block.

[0192] On the one hand, in the case of the affine merge mode in which motion vectors (MVs) are derived in sub-block units, it can be called the sub-block merge mode, which can be indicated based on a sub-block merge flag (or the merge_subblock_flag syntax element). Alternatively, when the value of the merge_subblock_flag syntax element is 1, it can be indicated that the sub-block merge mode is applied. In this case, the affine merge candidate list described later can also be called the sub-block merge candidate list. In this case, the sub-block merge candidate list can further include candidates derived by SbTMVP described later. In this case, the candidate derived by the SbTMVP can be used as the candidate at the 0th index of the sub-block merge candidate list. In other words, the candidate derived by the SbTMVP can be positioned before the inherited affine candidate or the constructed affine candidate described later within the sub-block merge candidate list.

[0193] When the affine merge mode is applied, an affine merge candidate list can be constructed for CPMV derivation for the current block. For example, the affine merge candidate list can include at least one of the following candidates: 1) Inherited affine merge candidates. 2) Constructed affine merge candidates. 3) Zero motion vector candidates (or zero vectors). Here, the inherited affine merge candidate is a candidate derived based on the CPMVs of neighboring blocks when the neighboring blocks are coded in the affine mode, the constructed affine merge candidate is a candidate derived by constructing CPMVs based on the MVs of the neighboring blocks of the current block for each CPMV unit, and the zero motion vector candidate can represent a candidate composed of CPMVs whose value is 0.

[0194] The affine merge candidate list can be configured, for example, as follows.

[0195] There can be at most two inherited affine candidates, and the inherited affine candidates can be derived from the affine motion models of adjacent blocks. The adjacent blocks can include one left adjacent block and the upper adjacent block. The candidate block can be positioned as shown in FIG. 8. The scan order for the left predictor can be A1->A0, and the scan order for the above predictor can be B1->B0->B2. Only one inherited candidate can be selected from each of the left and upper sides. A pruning check cannot be performed between the two inherited candidates.

[0196] When an adjacent affine block is confirmed, the control point motion vector of the confirmed block can be used to derive CPMV candidates in the affine merge list of the current block. Here, the adjacent affine block can represent a block coded in the affine prediction mode among the adjacent blocks of the current block. For example, as shown in FIG. 12, when the bottom-left adjacent block A is coded in the affine prediction mode, the motion vectors v2, v3, and v4 of the top-left corner, top-right corner, and bottom-left corner of the adjacent block A can be obtained. When the adjacent block A is coded in the 4-parameter affine motion model, two CPMVs of the current block can be calculated by v2 and v3. When the adjacent block A is coded in the 6-parameter affine motion model, three CPMVs of the current block can be calculated by v2, v3, and v4.

[0197] FIG. 13 is a diagram for explaining the positions of candidates in the affine merge mode or the sub-block merge mode.

[0198] An affine candidate constructed in the affine merge mode or sub-block merge mode can mean a candidate constructed by combining the translational motion information adjacent to each control point. The motion information of the control point is derived from the specified spatial adjacency and temporal adjacency. CPMVk (k = 0, 1, 2, 3) can represent the k-th control point.

[0199] As shown in FIG. 13, for CPMV0, blocks can be checked in the order of B2 ≧ B3 ≧ A2, and the motion vector of the first available block can be used. For CPMV1, blocks can be checked in the order of B1 ≧ B0, and for CPMV2, blocks can be checked in the order of A1 ≧ A0. The TMVP (temporal motion vector predictor) can be used as CPMV3 if available.

[0200] After the motion vectors of the four control points are obtained, the affine merge candidate can be generated based on the obtained motion information. The combination of control point motion vectors can correspond to any one of {CPMV0, CPMV1, CPMV2}, {CPMV0, CPMV1, CPMV3}, {CPMV0, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV3}, {CPMV0, CPMV1}, and {CPMV0, CPMV2}.

[0201] The combination of three CPMVs can form a 6-parameter affine merge candidate, and the combination of two CPMVs can form a 4-parameter affine merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different from each other, the related combinations of control point motion vectors can be discarded.

[0202] FIG. 14 is a diagram for explaining SbTMVP in inter prediction.

[0203] SbTMVP (subblock-based temporal motion vector prediction) can also be called ATMVP (advanced temporal motion vector prediction). SbTMVP can utilize the motion field in the collocated picture (which can also be called the col picture) to improve motion vector prediction and the merge mode for CUs within the current picture.

[0204] For example, SbTMVP can predict motion at the subblock (or sub-CU) level. Also, SbTMVP can apply a motion shift before fetching temporal motion information from the col picture. Here, the motion shift can be obtained from one of the motion vectors of the spatial neighboring blocks of the current block.

[0205] SbTMVP can predict the motion vectors of subblocks (or sub-CUs) within the current block (or CU) in two steps.

[0206] In the first step, the spatial neighboring blocks can be tested in the order of A1, B1, B0, and A0 in FIG. 4. The first spatial neighboring block having a motion vector that uses the col picture as its reference picture can be identified, and the motion vector can be selected with the applied motion shift. If no such motion is identified from the spatial neighboring blocks, the motion shift can be set to (0, 0).

[0207] In the second step, the motion shift confirmed in the first step can be applied to obtain sub-block level motion information (motion vectors and reference indexes) from the col picture. For example, the motion shift can be added to the coordinates of the current block. For example, the motion shift can be set with the motion of A1 in FIG. 8. In this case, the motion information of the corresponding block in the col picture for each sub-block can be used to derive the motion information of the sub-block. Temporal motion scaling can be applied to align the reference picture of the temporal motion vector with the reference picture of the current block.

[0208] A combined sub-block based merge list including all SbTVMP candidates and affine merge candidates can be used for signaling the affine merge mode. Here, the affine merge mode can be called the sub-block based merge mode. The SbTVMP mode can be available or unavailable depending on the flag included in the SPS (sequence parameter set). When the SbTMVP mode is available, the SbTMVP predictor can be added as the first entry in the list of sub-block based merge candidates, and the affine merge candidates can come next. The maximum allowable size of the affine merge candidate list can be 5.

[0209] The size of the sub CU (or sub-block) used in SbTMVP can be fixed to 8×8, and similar to the affine merge mode, the SbTMVP mode can be applied only to blocks where both the width and height are 8 or more. The encoding logic for additional SbTMVP merge candidates can be the same as that of other merge candidates. That is, an RD check using additional rate-distortion (RD) cost for each CU in the P or B slice can be performed to determine whether to use the SbTMVP candidates.

[0210] FIG. 15 is a diagram for explaining the CIIP mode (combined inter-picture merge and intra-picture prediction mode) in inter prediction.

[0211] CIIP (Combined Inter and Intra Prediction) can be applied to the current CU. For example, when the CU is coded in the merge mode, the CU contains at least 64 luma samples (i.e., when the product of the CU width and the CU height is 64 or more), and when both the CU width and the CU height are smaller than 128 luma samples, an additional flag (e.g., ciip_flag) can be signaled to indicate whether the CIIP mode is applied to the current CU.

[0212] In CIIP prediction, an inter prediction signal and an intra prediction signal can be combined. In the CIIP mode, the inter prediction signal P_inter can be derived using the same inter prediction process applied to the regular merge mode. The intra prediction signal P_intra can be derived by an intra prediction process having a planar mode.

[0213] The intra prediction signal and the inter prediction signal can be combined using weighted averaging and can be as shown in Equation 4 below. The weighting value can be calculated according to the coding modes of the upper and left adjacent blocks shown in FIG. 15.

[0214]

Equation

[0215] In the formula 4, when the upper adjacent block is available and intra-coded, isIntraTop can be set to 1, and if not, isIntraTop can be set to 0. When the left adjacent block is available and intra-coded, isIntraLeft can be set to 1, and if not, isIntraLeft can be set to 0. When (isIntraLeft + isIntraLeft) is 2, wt can be set to 3, and when (isIntraLeft + isIntraLeft) is 1, wt can be set to 2. Otherwise, wt can be set to 1.

[0216] FIG. 16 is a diagram for explaining a partitioning mode in inter prediction.

[0217] As shown in FIG. 16, when the partitioning mode is applied, the CU can be evenly divided into two triangular partitions using a diagonal split or an anti-diagonal split. However, this is only an example of the partitioning mode, and the CU can be evenly or unevenly divided into partitions of various shapes.

[0218] For each partition of the CU, only unidirectional prediction can be allowed. That is, each partition can have one motion vector and one reference index. The unidirectional prediction constraint is to ensure that, similar to bi-prediction, only two motion compensation predictions are required for each CU.

[0219] When the partitioning mode is applied, a flag representing the split direction (diagonal direction or anti-diagonal direction) and two merge indices (for each partition) can be further signaled.

[0220] After predicting each partition, the boundary lines in the diagonal or anti-diagonal direction can be adjusted using a blending processing with adaptive weights based on the sample values.

[0221] On the other hand, when the merge mode or skip mode is applied, in order to generate prediction samples, as described above, motion information can be derived based on a regular merge mode, a merge mode with motion vector difference, a merge subblock mode, a combined inter-picture merge and intra-picture prediction mode, or a partitioning mode. Each mode can be enabled or disabled via an on / off flag in the SPS (Sequence parameter set). If the on / off flag for a specific mode in the SPS is disabled, the syntax that is explicitly transmitted for the prediction mode in CU or PU units cannot be signaled.

[0222] Table 5 below relates to the process of deriving the merge mode or skip mode from the conventional merge_data syntax. In Table 5 below, CUMergeTriangleFlag[x0][y0] can correspond to the on / off flag for the partitioning mode described above in FIG. 12, and merge_triangle_split_dir[x0][y0] can represent the split direction (diagonal direction or opposite diagonal direction) when the partitioning mode is applied. Also, merge_triangle_idx0[x0][y0] and merge_triangle_idx1[x0][y0] can represent the two merge indices for each partition when the partitioning mode is applied.

[0223]

Table 5-1

[0224]

Table 5-2

[0225] On the other hand, each prediction mode including the regular merge mode, MMVD mode, merge sub-block mode, CIIP mode, and partitioning mode can be enabled or disabled from the SPS (Sequence parameter set) as shown in Table 6 below. In Table 6 below, sps_triangle_enabled_flag can correspond to the flag for enabling or disabling the partitioning mode described above in FIG. 12 from the SPS.

[0226]

Table 6-1

[0227]

Table 6-2

[0228]

Table 6-3

[0229]

Table 6-4

[0230]

Table 6-5

[0231]

Table 6-6

[0232] The merge_data syntax in Table 5 above can be parsed or derived based on the flags of SPS in Table 6 above and the conditions under which each prediction mode can be used. Organizing all cases according to the conditions under which the SPS flags and each prediction mode can be used, it may be as shown in Tables 7 and 8 below. Table 7 shows the number of cases when the current block is in the merge mode, and Table 8 shows the number of cases when the current block is in the skip mode. In Tables 7 and 8 below, Triangle or TRI can correspond to the partitioning mode described in FIG. 12.

[0233]

Table 7

[0234]

Table 8

[0235] As an example when mentioned in Table 7 and Table 8 above, the case where the current block is 4x16 and in skip mode will be described. When the merge sub-block mode, MMVD mode, CIIP mode, and partitioning mode are all enabled in the SPS, if regular_merge_flag[x0][y0], mmvd_flag[x0][y0], and merge_subblock_flag[x0][y0] are all 0 in the merge_data syntax, the motion information for the current block must be derived in the partitioning mode. However, even if the partitioning mode is enabled from the on / off flag in the SPS, it cannot be used as a prediction mode unless it additionally satisfies the conditions in Table 9 below. In Table 9 below, MergeTriangleFlag[x0][y0] can correspond to the on / off flag for the partitioning mode, and sps_triangle_enabled_flag can correspond to the flag that enables or disables the partitioning mode from the SPS.

[0236]

Table 9

[0237] Referring to Table 9 above, if the current slice is a P slice, prediction samples cannot be generated via the partitioning mode, so there may be a case where the decoder cannot decode the bitstream any further. In order to solve the problems that occur in exceptional cases where decoding is not performed because the final prediction mode cannot be selected by each on / off flag of the SPS and the merge data syntax, this document proposes a default merge mode. The default merge mode can be pre-defined in various ways or induced via additional syntax signaling.

[0238] In one embodiment, when the MMVD mode, the merge sub-block mode, the CIIP mode, and the partitioning mode for dividing the current block into two partitions for prediction are not available, the regular merge mode can be applied to the current block. That is, when the merge mode cannot be finally selected for the current block, the regular merge mode can be applied as the default merge mode.

[0239] For example, based on a general merge flag indicating whether the merge mode is available for the current block, if the merge mode is available for the current block but the merge mode cannot be finally selected for the current block, the regular merge mode can be applied as the default merge mode.

[0240] At this time, among the merge candidates included in the merge candidate list generated by applying the regular merge mode, the prediction sample of the current block can be generated using the merge index information indicating one candidate. For example, the motion information of the current block can be derived based on the merge index information, and the prediction sample can be generated based on the derived motion information.

[0241] The merge data syntax resulting from this can be as shown in Table 10 below.

[0242]

Table 10-1

[0243]

Table 10-2

[0244] Referring to Table 10 and Table 6 above, based on the case where the MMVD mode is not available, the flag sps_mmvd_enabled_flag for enabling or disabling the MMVD mode from the SPS may be 0, or the first flag (mmvd_merge_flag[x0][y0]) indicating whether the MMVD mode is applied may be 0.

[0245] Also, based on the case where the merge sub-block mode is not available, the flag sps_affine_enabled_flag for enabling or disabling the merge sub-block mode from the SPS may be 0, or the second flag (merge_subblock_flag[x0][y0]) indicating whether the merge sub-block mode is applied may be 0.

[0246] Also, based on the case where the CIIP mode is not available, the flag sps_ciip_enabled_flag for enabling or disabling the CIIP mode from the SPS may be 0, or the third flag (ciip_flag[x0][y0]) indicating whether the CIIP mode is applied may be 0.

[0247] Also, based on the case where the partitioning mode is not available, the flag sps_triangle_enabled_flag for enabling or disabling the partitioning mode from the SPS may be 0, or the fourth flag (MergeTriangleFlag[x0][y0]) indicating whether the partitioning mode is applied may be 0.

[0248] On the one hand, based on the case where the value of a general merge flag indicating whether the merge mode is available for the current block is 1, the values of a first flag (mmvd_merge_flag[x0][y0]) indicating whether the MMVD mode is applied, a second flag (merge_subblock_flag[x0][y0]) indicating whether the merge sub-block mode is applied, and a third flag (ciip_flag[x0][y0]) indicating whether the CIIP mode is applied can be signaled.

[0249] Also, for example, based on the case where the partitioning mode is disabled based on the flag sps_triangle_enabled_flag, a fourth flag (MergeTriangleFlag[x0][y0]) indicating whether the partitioning mode is applied can be set to 0.

[0250] In another embodiment, based on the case where the regular merge mode, MMVD mode, merge sub-block mode, CIIP mode, and the partitioning mode for predicting by dividing the current block into two partitions are not available, the regular merge mode can be applied to the current block. That is, when the merge mode cannot be finally selected for the current block, the regular merge mode can be applied as the default merge mode.

[0251] For example, when the merge mode is available for the current block based on a general merge flag indicating whether the merge mode is available for the current block, but the merge mode cannot be finally selected for the current block, the regular merge mode can be applied as the default merge mode.

[0252] For example, based on the case where the MMVD mode is not available, the flag sps_mmvd_enabled_flag for enabling or disabling the MMVD mode from the SPS may be 0, or the first flag (mmvd_merge_flag[x0][y0]) indicating whether the MMVD mode is applied may be 0.

[0253] Also, based on the case where the merge sub-block mode is not available, the flag sps_affine_enabled_flag for enabling or disabling the merge sub-block mode from the SPS may be 0, or the second flag (merge_subblock_flag[x0][y0]) indicating whether the merge sub-block mode is applied may be 0.

[0254] Also, based on the case where the CIIP mode is not available, the flag sps_ciip_enabled_flag for enabling or disabling the CIIP mode from the SPS may be 0, or the third flag (ciip_flag[x0][y0]) indicating whether the CIIP mode is applied may be 0.

[0255] Also, based on the case where the partitioning mode is not available, the flag sps_triangle_enabled_flag for enabling or disabling the partitioning mode from the SPS may be 0, or the fourth flag (MergeTriangleFlag[x0][y0]) indicating whether the partitioning mode is applied may be 0.

[0256] Also, based on the case where the regular merge mode is not available, a fifth flag (regular_merge_flag[x0][y0]) indicating whether the regular merge mode is applied can be 0. That is, even when the value of the fifth flag is 0, the regular merge mode can be applied to the current block based on the case where the MMVD mode, merge sub-block mode, CIIP mode, and partitioning mode are not available.

[0257] At this time, among the merge candidates included in the merge candidate list of the current block, the motion information of the current block can be derived based on the first candidate, and a predicted sample can be generated based on the derived motion information.

[0258] In yet another embodiment, the regular merge mode can be applied to the current block based on the case where the regular merge mode, MMVD mode, merge sub-block mode, CIIP mode, and partitioning mode for dividing the current block into two partitions for prediction are not available. That is, when the merge mode cannot be finally selected for the current block, the regular merge mode can be applied as the default merge mode.

[0259] For example, when the merge mode is available for the current block based on a general merge flag indicating whether the merge mode is available for the current block, but the merge mode cannot be finally selected for the current block, the regular merge mode can be applied as the default merge mode.

[0260] For example, based on the case where the MMVD mode is not available, the flag sps_mmvd_enabled_flag that enables or disables the MMVD mode from the SPS may be 0, or the first flag (mmvd_merge_flag[x0][y0]) indicating whether the MMVD mode is applied may be 0.

[0261] Also, based on the case where the merge sub-block mode is not available, the flag sps_affine_enabled_flag that enables or disables the merge sub-block mode from the SPS may be 0, or the second flag (merge_subblock_flag[x0][y0]) indicating whether the merge sub-block mode is applied may be 0.

[0262] Also, based on the case where the CIIP mode is not available, the flag sps_ciip_enabled_flag that enables or disables the CIIP mode from the SPS may be 0, or the third flag (ciip_flag[x0][y0]) indicating whether the CIIP mode is applied may be 0.

[0263] Also, based on the case where the partitioning mode is not available, the flag sps_triangle_enabled_flag that enables or disables the partitioning mode from the SPS may be 0, or the fourth flag (MergeTriangleFlag[x0][y0]) indicating whether the partitioning mode is applied may be 0.

[0264] Also, based on the case where the regular merge mode is not available, a fifth flag (regular_merge_flag[x0][y0]) indicating whether the regular merge mode is applied can be 0. That is, even when the value of the fifth flag is 0, based on the case where the MMVD mode, merge sub-block mode, CIIP mode, and partitioning mode are not available, the regular merge mode can be applied to the current block.

[0265] At this time, the (0,0) motion vector can be derived as the motion information of the current block, and a predicted sample of the current block can be generated based on the (0,0) motion information. The (0,0) motion vector can perform prediction by referring to the 0th reference picture in the L0 reference list (reference list). However, if the 0th reference picture (RefPicList[0][0]) in the L0 reference list does not exist, prediction can be performed by referring to the 0th reference picture (RefPicList[1][0]) in the L1 reference list.

[0266] FIGS. 17 and 18 schematically show an example of a video / video encoding method and related components according to an embodiment(s) of this document.

[0267] The method disclosed in FIG. 17 can be performed by the encoding device disclosed in FIG. 2 or FIG. 18. Specifically, for example, S1700 to S1710 in FIG. 17 can be performed by the prediction unit 220 of the encoding device 200 in FIG. 18, S1720 in FIG. 17 can be performed by the residue processing unit 230 of the encoding device 200 in FIG. 18, and S1730 in FIG. 17 can be performed by the entropy encoding unit 240 of the encoding device 200 in FIG. 18. The method disclosed in FIG. 17 can include the embodiments described above in this document.

[0268] As shown in FIG. 17, the encoding device can determine the inter-prediction mode of the current block and generate inter-prediction mode information representing the inter-prediction mode (S1700). For example, the encoding device can determine at least one of various modes such as the regular merge mode, skip mode, MVP (motion vector prediction) mode, MMVD mode (merge mode with motion vector difference), merge subblock mode, CIIP mode (combined inter-picture merge and intra-picture prediction mode), and partitioning mode (a mode in which the current block is divided into two partitions for prediction) as the inter-prediction mode to be applied to the current block, and can generate inter-prediction mode information representing this.

[0269] The encoding device can generate prediction samples for the current block based on the determined prediction mode (S1710). For example, the encoding device can generate a merge candidate list according to the determined inter-prediction mode.

[0270] For example, candidates can be inserted into the merge candidate list until the number of candidates in the merge candidate list reaches the maximum number of candidates. Here, a candidate can represent a candidate or a candidate block for deriving motion information (or a motion vector) of the current block. For example, a candidate block can be derived through a search for adjacent blocks of the current block. For example, adjacent blocks can include spatial adjacent blocks and / or temporal adjacent blocks of the current block, and spatial adjacent blocks can be preferentially searched (spatial merge) to derive candidates, and then temporal adjacent blocks can be searched (temporal merge) to be derived as candidates, and the derived candidates can be inserted into the merge candidate list. For example, even after inserting the candidates, if the number of candidates in the merge candidate list is less than the maximum number of candidates, additional candidates can be inserted into the merge candidate list. For example, the additional candidates can include at least one of history based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate (when the slice / tile group type of the current slice / tile group is of type B), and / or zero vector merge candidate.

[0271] As described above, the merge candidate list can include at least a part of spatial merge candidates, temporal merge candidates, pairwise candidates, or zero vector candidates, and for the inter prediction of the current block, one of such candidates can be selected.

[0272] For example, the selection information can include index information indicating one candidate among the merge candidates included in the merge candidate list. For example, the selection information can be called merge index information.

[0273] For example, the encoding device can generate a predicted sample of the current block based on the candidate pointed to by the merge index information. Or, for example, the encoding device can derive motion information based on the candidate pointed to by the merge index information, and generate a predicted sample of the current block based on the derived motion information.

[0274] On the other hand, according to one embodiment, the inter-prediction mode information includes a general merge flag indicating whether the merge mode is available for the current block, and the merge mode is available for the current block based on the general merge flag.

[0275] At this time, when the MMVD mode (merge mode with motion vector difference), the merge sub-block mode, the CIIP mode (combined inter-picture merge and intra-picture prediction mode), and the partitioning mode for dividing the current block into two partitions for prediction are not available, the regular merge mode can be applied.

[0276] For example, the inter-prediction mode information includes merge index information pointing to one candidate among the merge candidates included in the merge candidate list generated by applying the regular merge mode, and the predicted sample can be generated using the merge index information. That is, the motion information of the current block can be derived based on the candidate pointed to by the merge index information, and the predicted sample of the current block can be generated based on the derived motion information.

[0277] For example, the inter prediction mode information may include a first flag indicating whether the MMVD mode is applicable, a second flag indicating whether the merge sub-block mode is applicable, and a third flag indicating whether the CIIP mode is applicable.

[0278] For example, based on the case where the MMVD mode, the merge sub-block mode, the CIIP mode, and the partitioning mode are not available, the values of the first flag, the second flag, and the third flag can all be 0.

[0279] Also, for example, based on the case where the value of the general merge flag is 1, the merge mode is available for the current block, and based on the case where the value of the general merge flag is 1, the values of the first flag, the second flag, and the third flag can be signaled.

[0280] For example, a flag for enabling or disabling the partitioning mode is included in the SPS (Sequence Parameter Set) of the video information, and based on the case where the partitioning mode is disabled, the value of a fourth flag indicating whether the partitioning mode is applicable can be set to 0.

[0281] On the other hand, the inter prediction mode information may further include a fifth flag indicating whether the regular merge mode is applicable. Even when the value of the fifth flag is 0, based on the case where the MMVD mode, the merge sub-block mode, the CIIP mode, and the partitioning mode are not available, the regular merge mode can be applied to the current block.

[0282] In such a case, the motion information of the current block can be derived based on the first merge candidate among the merge candidates included in the merge candidate list of the current block. Further, the prediction sample can be generated based on the motion information of the current block derived based on the first merge candidate.

[0283] Alternatively, in such a case, the motion information of the current block can be derived based on the (0,0) motion vector, and the prediction sample can be generated based on the motion information of the current block derived based on the (0,0) motion vector.

[0284] The encoding device can generate residual information based on the residual sample for the current block (S1720). For example, the encoding device can derive a residual sample based on the prediction sample and the original sample. For example, the encoding device can generate residual information representing the quantized transform coefficients of the residual sample. The residual information can be generated by various encoding methods such as exponential Golomb, CAVLC, CABAC, etc.

[0285] The encoding device can encode video information including inter-prediction mode information and residual information (S1730). For example, the video information can also be referred to as video data. The video information can include various information according to the above-described embodiment(s) of this document. For example, the video information can include at least a part of prediction-related information or residual-related information. For example, the prediction-related information can include at least a part of the inter-prediction mode information, selection information, and inter-prediction type information. For example, the encoding device can encode video information including all or part of the above-described information (or syntax elements) to generate a bitstream or encoded information. Or it can be output in the form of a bitstream. Also, the bitstream or encoded information can be transmitted to a decoding device via a network or a storage medium.

[0286] Or, although not shown in FIG. 17, for example, the encoding device can generate restored samples based on the residual samples and the prediction samples. A restored block and a restored picture can be derived based on the restored samples. Or, for example, the encoding device can encode video information including residual-related information or prediction-related information.

[0287] For example, the encoding device can encode video information including all or part of the above-described information (or syntax elements) to generate a bitstream or encoded information. Or it can be output in the form of a bitstream. Also, the bitstream or encoded information can be transmitted to a decoding device via a network or a storage medium. Or, the bitstream or encoded information can be stored in a computer-readable storage medium, and the bitstream or the encoded information can be generated by the above-described video encoding method.

[0288] Figures 19 and 20 schematically show an example of a video / video decoding method and related components according to the embodiment(s) of this document.

[0289] The method disclosed in FIG. 19 can be performed by the decoding apparatus disclosed in FIG. 3 or FIG. 20. Specifically, for example, S1900 in FIG. 19 can be performed by the entropy decoding unit 310 of the decoding apparatus 300 in FIG. 20, and S1910 in FIG. 19 can be performed by the residue processing unit 320 of the decoding apparatus 300 in FIG. 20. Also, S1920 in FIG. 19 can be performed by the prediction unit 330 of the decoding apparatus 300 in FIG. 20, and S1930 in FIG. 19 can be performed by the addition unit 340 of the decoding apparatus 300 in FIG. 20. The method disclosed in FIG. 19 can include the embodiments described above in this document.

[0290] As shown in FIG. 19, the decoding apparatus can receive video information including inter prediction mode information and residual information via a bitstream (S1900). For example, the video information can also be referred to as video data. The video information can include various information according to the embodiment(s) described above in this document. For example, the video information can include at least a part of prediction-related information or residual-related information.

[0291] For example, the prediction-related information can include inter-prediction mode information or inter-prediction type information. For example, the inter-prediction mode information can include information representing at least a part of various inter-prediction modes. For example, various modes can be used, such as regular merge mode, skip mode, MVP (motion vector prediction) mode, MMVD mode (merge mode with motion vector difference), merge subblock mode, CIIP mode (combined inter-picture merge and intra-picture prediction mode), and partitioning mode for dividing the current block into two partitions for prediction. For example, the inter-prediction type information can include the inter_pred_idc syntax element. Alternatively, the inter-prediction type information can include information representing any one of L0 prediction, L1 prediction, or bi-prediction.

[0292] The decoding device can generate residual samples based on the residual information (S1910). The decoding device can derive quantized transform coefficients based on the residual information and can derive residual samples based on an inverse transform procedure for the transform coefficients.

[0293] The decoding device can generate prediction samples for the current block by applying the prediction mode determined based on the inter prediction mode information (S1510). For example, the decoding device can generate a merge candidate list according to the determined inter prediction mode among the regular merge mode, skip mode, MVP mode, MMVD mode, merge sub-block mode, CIIP mode, and partitioning mode in which the current block is divided into two partitions for prediction based on the inter prediction mode information.

[0294] For example, candidates can be inserted into the merge candidate list until the number of candidates in the merge candidate list reaches the maximum number of candidates. Here, a candidate can represent a candidate or a candidate block for deriving the motion information (or motion vector) of the current block. For example, a candidate block can be derived through searching for adjacent blocks of the current block. For example, adjacent blocks can include spatial adjacent blocks and / or temporal adjacent blocks of the current block, and spatial adjacent blocks are preferentially searched (spatial merge) to derive candidates, and then temporal adjacent blocks are searched (temporal merge) to be derived as candidates, and the derived candidates can be inserted into the merge candidate list. For example, if the number of candidates in the merge candidate list is less than the maximum number of candidates even after inserting the candidates, additional candidates can be inserted into the merge candidate list. For example, additional candidates can include at least one of history based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate (when the slice / tile group type of the current slice / tile group is of type B), and / or zero vector merge candidate.

[0295] As described above, the merge candidate list can include at least a part of spatial merge candidates, temporal merge candidates, pairwise candidates, or zero vector candidates, and for the inter prediction of the current block, one of such candidates can be selected.

[0296] For example, the selection information can include index information indicating one candidate among the merge candidates included in the merge candidate list. For example, the selection information can also be called merge index information.

[0297] For example, the decoding device can generate a prediction sample of the current block based on the candidate indicated by the merge index information. Or, for example, the decoding device can derive motion information based on the candidate indicated by the merge index information, and generate a prediction sample of the current block based on the motion information.

[0298] On the other hand, according to one embodiment, the inter prediction mode information includes a general merge flag indicating whether the merge mode is available for the current block, and based on the general merge flag, the merge mode is available for the current block.

[0299] At this time, when the MMVD mode (merge mode with motion vector difference), the merge subblock mode, the CIIP mode (combined inter-picture merge and intra-picture prediction mode), and the partitioning mode (which divides the current block into two partitions for prediction) are not available, the regular merge mode can be applied.

[0300] For example, the inter prediction mode information includes merge index information indicating one candidate among the merge candidates included in the merge candidate list generated by applying the regular merge mode, and the prediction sample can be generated using the merge index information. That is, the motion information of the current block can be derived based on the candidate indicated by the merge index information, and the prediction sample of the current block can be generated based on the derived motion information.

[0301] For example, the inter prediction mode information can include a first flag indicating whether the MMVD mode is applied, a second flag indicating whether the merge sub-block mode is applied, and a third flag indicating whether the CIIP mode is applied.

[0302] For example, based on the case where the MMVD mode, the merge sub-block mode, the CIIP mode, and the partitioning mode are not available, the values of the first flag, the second flag, and the third flag can all be 0.

[0303] Also, for example, based on the case where the value of the general merge flag is 1, the merge mode is available for the current block, and based on the case where the value of the general merge flag is 1, the values of the first flag, the second flag, and the third flag can be signaled.

[0304] For example, the flag for enabling or disabling the partitioning mode is included in the SPS (Sequence Parameter Set) of the video information, and based on the case where the partitioning mode is disabled, the value of the fourth flag indicating whether the partitioning mode is applied can be set to 0.

[0305] On the other hand, the inter prediction mode information may further include a fifth flag indicating whether the regular merge mode is applicable. Even when the value of the fifth flag is 0, the regular merge mode can be applied to the current block based on the case where the MMVD mode, the merge sub-block mode, the CIIP mode, and the partitioning mode are not available.

[0306] In such a case, the motion information of the current block can be derived based on the first merge candidate among the merge candidates included in the merge candidate list of the current block. Also, the predicted sample can be generated based on the motion information of the current block derived based on the first merge candidate.

[0307] Or, in such a case, the motion information of the current block can be derived based on the (0, 0) motion vector, and the predicted sample can be generated based on the motion information of the current block derived based on the (0, 0) motion vector.

[0308] The decoding device can generate a restored sample based on the predicted sample and the residual sample (S1930). For example, the decoding device can generate a restored sample based on the predicted sample and the residual sample, and a restored block and a restored picture can be derived based on the restored sample.

[0309] For example, the decoding device can decode a bitstream or encoded information to obtain video information including all or part of the above-described information (or syntax elements). Also, the bitstream or encoded information can be stored in a computer-readable storage medium, and the above-described decoding method can be executed.

[0310] In the foregoing embodiments, the method has been described based on a flowchart in a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with steps different from those described above. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps may be included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.

[0311] The method according to the embodiments of this document described above can be embodied in the form of software, and the encoding device and / or decoding device according to this document can be included in a device that performs video processing, such as a TV, computer, smartphone, set-top box, display device, etc.

[0312] In this document, when an embodiment is embodied in software, the method described above can be embodied by modules (processes, functions, etc.) that perform the functions described above. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), RAM (random access memory), flash memory, memory card, storage medium, and / or other storage devices. That is, the embodiments described in this document can be embodied and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be embodied and executed on a computer, processor, microprocessor, controller, or chip. In this case, information for embodiment (e.g., information on instructions) or an algorithm can be stored in a digital storage medium.

[0313] In addition, the decoding device and the encoding device to which the example(s) of this document is applied can be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a picture phone video device, a transportation means terminal (e.g., a vehicle (including an autonomous driving vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process a video signal or a data signal. For example, as an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.

[0314] In addition, the processing method to which the example(s) of this document is / are applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to the example(s) of this document can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (e.g., transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired or wireless communication network.

[0315] In addition, the example(s) of this document can be embodied as a computer program product by program code, and the program code can be executed by a computer according to the example(s) of this document. The program code is a computer

[0316] FIG. 21 shows an example of a content streaming system to which the example disclosed in this document can be applied.

[0317] As shown in FIG. 21, the content streaming system to which the example of this document is applied can include a large encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0318] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and plays the role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted.

[0319] The bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0320] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server plays the role of a medium to inform the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device within the content streaming system.

[0321] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0322] Examples of the user device include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, and a digital signage.

[0323] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.

[0324] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in a device, and the technical features of the device claims in this specification can be combined and implemented in a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a method.

Claims

1. A video decoding method performed by a decoding apparatus, comprising: obtaining video information including inter prediction mode information and residual information via a bitstream; generating residual samples based on the residual information; applying a prediction mode determined based on the inter prediction mode information to generate prediction samples of a current block; generating restored samples based on the prediction samples and the residual samples, wherein the inter prediction mode information includes a first flag related to whether a sub-block merge mode is applied, a second flag related to whether a regular merge mode is applied, and a third flag related to whether MMVD (merge mode with motion vector difference) is applied; based on the fact that the CIIP (combined inter-picture merge and intra-picture prediction) mode is not enabled, the partitioning prediction mode in which prediction is performed by dividing the current block into two partitions is not enabled, the value of the first flag is equal to 0, and the value of the third flag is equal to 0, the inter prediction mode information includes merge index information indicating one of the merge candidates included in the merge candidate list generated for the regular merge mode; the prediction samples are generated based on the merge index information; A method, wherein a regular merge mode is applied to the current block based on the fact that the value of a general merge flag is 1 and the MMVD is not enabled.

2. A video encoding method performed by an encoding apparatus, comprising: determining an inter prediction mode of a current block and generating inter prediction mode information representing the inter prediction mode; generating prediction samples of the current block based on the determined prediction mode; generating residual information based on residual samples for the current block; encoding video information including the inter prediction mode information and the residual information. The inter prediction mode information includes a first flag related to whether the sub-block merge mode is applied, a second flag related to whether the regular merge mode is applied, and a third flag related to whether MMVD (merge mode with motion vector difference) is applied. Based on the fact that the CIIP (combined inter-picture merge and intra-picture prediction) mode is not enabled, the partitioning prediction mode in which prediction is performed by dividing the current block into two partitions is not enabled, the value of the first flag is equal to 0, and the value of the third flag is equal to 0, the inter prediction mode information includes merge index information indicating one of the merge candidates included in the merge candidate list generated for the regular merge mode. A method in which the regular merge mode is applied to the current block based on the fact that the value of the general merge flag is 1 and the MMVD is not enabled.

3. A method for transmitting data for a video, comprising: A step of generating a bitstream for the video, the bitstream comprising: A step of determining an inter prediction mode of a current block and generating inter prediction mode information representing the inter prediction mode; A step of generating prediction samples of the current block based on the determined prediction mode; A step of generating residual information based on residual samples for the current block; Encoding video information including the inter prediction mode information and the residual information; and a step generated based thereon; A step of transmitting the data including the bitstream. The inter prediction mode information includes a first flag related to whether the sub-block merge mode is applied, a second flag related to whether the regular merge mode is applied, and a third flag related to whether MMVD (merge mode with motion vector difference) is applied. Based on the fact that the CIIP (combined inter-picture merge and intra-picture prediction) mode is not enabled, the partitioning prediction mode in which prediction is performed by splitting the current block into two partitions is not enabled, the value of the first flag is equal to 0, and the value of the third flag is equal to 0, the inter prediction mode information includes merge index information indicating one of the merge candidates included in the merge candidate list generated for the regular merge mode. A method in which the regular merge mode is applied to the current block based on the fact that the value of the general merge flag is 1 and the MMVD is not enabled.

Citation Information

Patent Citations

  • Image decoding method and apparatus for generating prediction samples by applying determined prediction mode

    JP7678069B2