Image encoding / decoding method, apparatus, and method for transmitting a bit stream for deriving a weight index for bidirectional prediction of merge candidates

The image encoding/decoding method addresses the challenge of efficiently compressing high-resolution images by using an affine merge candidate list and deriving a weight index for bidirectional prediction, resulting in enhanced encoding/decoding efficiency for effective image transmission, storage, and reproduction.

JP7691560B2Active Publication Date: 2025-06-11NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024122996
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-08-08
Filing Date
2024-07-30
Publication Date
2025-06-11
Estimated Expiration
2040-07-06

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images has led to a need for highly efficient image compression techniques to effectively transmit, store, and reproduce such images, as conventional methods face challenges in managing the increased amount of information.

Method used

An image encoding/decoding method and apparatus that improves encoding/decoding efficiency by constructing an affine merge candidate list, selecting an affine merge candidate, deriving motion information, generating a prediction block, and restoring the current block, while also deriving a weight index for bidirectional prediction of merge candidates.

Benefits of technology

The proposed method achieves improved encoding/decoding efficiency, enabling effective transmission, storage, and reproduction of high-resolution images by optimizing the image compression process through advanced prediction and motion information handling techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691560000004
    Figure 0007691560000004
  • Figure 0007691560000005
    Figure 0007691560000005
  • Figure 0007691560000006
    Figure 0007691560000006
Patent Text Reader

Abstract

To provide an image encoding / decoding method and device for improving encoding / decoding efficiency.SOLUTION: An image decoding method and device includes, if an inter-prediction mode of a current block is an affine merge mode, the steps for: selecting one affine merge candidate from an affine merge candidate list to the current block; deriving motion information of the current block on the basis of motion information of the selected affine merge candidate; generating a prediction block of the current block based on the motion information of the current block; and reconstructing the current block on the basis of the prediction block of the current block. A step for constructing the affine merge candidate list includes a step for deriving a combined affine merge candidate, and the step for deriving the combined affine merge candidate includes a step for deriving a weight index for bidirectional prediction of the combined affine merge candidate.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method, apparatus, and a method for transmitting a bitstream, and more particularly, to an image encoding / decoding method, apparatus, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure for inducing a weight index for bidirectional prediction of merge candidates.

Background Art

[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus for inducing a weight index for bidirectional prediction of merge candidates.

[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0007] Furthermore, an object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0008] Furthermore, an object of the present disclosure is to provide a recording medium storing a bitstream that is received by an image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0009] The technical problems to be solved by the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.

Means for Solving the Problems

[0010] An image decoding method according to an aspect of the present disclosure includes, when an inter prediction mode of a current block is an affine merge mode, a step of constructing an affine merge candidate list for the current block, a step of selecting one affine merge candidate from the affine merge candidate list, a step of deriving motion information of the current block based on motion information of the selected affine merge candidate, a step of generating a prediction block of the current block based on the motion information of the current block, and a step of restoring the current block based on the prediction block of the current block. The step of constructing the affine merge candidate list includes a step of deriving a combined affine merge candidate, and the step of deriving the combined affine merge candidate can include a step of deriving a weight index for bidirectional prediction of the combined affine merge candidate.

[0011] In the image decoding method according to the present disclosure, the step of deriving the combined affine merge candidate can be performed based on motion information for each CP included in a combination of already defined CPs among a plurality of CPs of the current block.

[0012] In the image decoding method according to the present disclosure, the motion information for the CP is derived based on the motion information of candidate blocks for the CP, and the candidate block can be an available candidate block among at least one candidate block for the CP.

[0013] In the image decoding method according to the present disclosure, the motion information for the CP includes a weight index for bidirectional prediction, and the weight index for bidirectional prediction can be derived when the CP is the upper left CP or the upper right CP of the current block.

[0014] In the image decoding method according to the present disclosure, the motion information for the CP includes a weight index for bidirectional prediction, and the weight index for bidirectional prediction cannot be derived when the CP is the lower left CP or the lower right CP of the block.

[0015] In the image decoding method according to the present disclosure, when there is no available candidate block among at least one candidate block for the CP, the CP can be determined to be unavailable.

[0016] In the image decoding method according to the present disclosure, the step of deriving the combined affine merge candidate can be performed when all the CPs included in the already defined combination of CPs are available.

[0017] In the image decoding method according to the present disclosure, the already defined CPs have a predetermined order within the combination, and the weight index of the combined affine merge candidate can be derived based on the order within the combination of the already defined CPs.

[0018] In the image decoding method according to the present disclosure, the weight index of the combined affine merge candidate is derived based on whether the prediction direction for the combination can be used, and whether the prediction direction for the combination can be used can be derived based on the motion information for the CP included in the combination.

[0019] In the image decoding method according to the present disclosure, when whether the prediction direction for the combination can be used is available for both the L0 direction and the L1 direction, the weight index of the combined affine merge candidate can be derived to the weight index of a predetermined CP in the combination.

[0020] In the image decoding method according to the present disclosure, the predetermined CP in the combination used to derive the weight index of the combined affine merge candidate can be the first CP among the CPs in the combination.

[0021] In the image decoding method according to the present disclosure, when whether the prediction direction for the combination can be used is not available for at least one of the L0 direction and the L1 direction, the weight index of the combined affine merge candidate can be derived to a predetermined weight index.

[0022] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor. When the inter prediction mode of the current block is the affine merge mode, the at least one processor configures an affine merge candidate list for the current block, selects one affine merge candidate from the affine merge candidate list, derives the motion information of the current block based on the motion information of the selected affine merge candidate, generates a predicted block of the current block based on the motion information of the current block, restores the current block based on the predicted block of the current block, and the configuration of the affine merge candidate list includes deriving a combined affine merge candidate, and the derivation of the combined affine merge candidate can include deriving a weight index for bidirectional prediction of the combined affine merge candidate.

[0023] An image encoding method according to another aspect of the present disclosure includes generating a predicted block of the current block based on the motion information of the current block, encoding the current block based on the predicted block, and encoding the motion information of the current block. When the inter prediction mode of the current block is the affine merge mode, the step of encoding the motion information of the current block includes configuring an affine merge candidate list for the current block and encoding the motion information of the current block based on the affine merge candidate list. The step of configuring the affine merge candidate list includes deriving a combined affine merge candidate, and the step of deriving the combined affine merge candidate can include deriving a weight index for bidirectional prediction of the combined affine merge candidate.

[0024] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or the image encoding apparatus of the present disclosure.

[0025] The features briefly summarized and described above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present disclosure.

Advantages of the Invention

[0026] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0027] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus for deriving a weight index for bidirectional prediction of merge candidates.

[0028] Also, according to the present disclosure, it is possible to provide a method for transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0029] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0030] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream received by an image decoding apparatus according to the present disclosure, decoded, and used for image restoration.

[0031] The effects obtained in the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.

Brief Description of the Drawings

[0032]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Best Mode for Carrying Out the Invention

[0033] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.

[0034] In describing the embodiments of the present disclosure, when it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure, a detailed description thereof will be omitted. In the drawings, parts not related to the description of the present disclosure are omitted, and the same reference numerals are given to the same parts.

[0035] In the present disclosure, when a certain component is "connected", "coupled", or "connected" to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Also, when a certain component "includes" or "has" another component, this means that, unless otherwise stated to the contrary, it does not exclude other components but can further include other components.

[0036] In the present disclosure, terms such as "first", "second", etc. are used only for the purpose of distinguishing one component from another and do not limit the order or importance between components, etc., unless otherwise specifically mentioned. Therefore, within the scope of the present disclosure, the first component of one embodiment may be referred to as the second component in another embodiment, and similarly, the second component of one embodiment may be referred to as the first component in another embodiment.

[0037] In the present disclosure, components that are distinguished from each other are for clearly explaining their respective features and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, without further mention, such integrated or distributed embodiments are also included in the scope of the present disclosure.

[0038] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments constituted by a subset of the components described in one embodiment are also included in the scope of the present disclosure. Also, embodiments that further include other components in addition to the components described in various embodiments are included in the scope of the present disclosure.

[0039] The present disclosure relates to the encoding and decoding of images, and the terms used in the present disclosure can have the ordinary meanings in the technical field to which the present disclosure belongs unless newly defined in the present disclosure.

[0040] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period, and a slice / tile is an encoding unit constituting a part of a picture, and one picture can be composed of one or more slices / tiles. Also, a slice / tile can include one or more CTUs (coding tree units).

[0041] In the present disclosure, "pixel" or "pel" can mean the smallest unit constituting one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0042] In the present disclosure, "unit" can indicate a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can be used interchangeably with terms such as "sample array", "block", or "area" as the case may be. In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0043] In the present disclosure, "current block" can mean any one of "current coding block", "current coding unit", "block to be coded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" can mean "current transformation block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".

[0044] In the present disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".

[0045] In the present disclosure, "or" can be interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" can mean "additionally or alternatively".

[0046] Overview of Video Coding System

[0047] FIG. 1 is a diagram showing a video coding system according to the present disclosure.

[0048] A video coding system according to an embodiment can include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0049] An encoding device 10 according to an embodiment can include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment can include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 can be referred to as a video / image encoding unit, and the decoding unit 22 can be referred to as a video / image decoding unit. The transmission unit 13 can be included in the encoding unit 12. The reception unit 21 can be included in the decoding unit 22. The rendering unit 23 can also include a display unit, and the display unit can be configured as a separate device or an external component.

[0050] The video source generation unit 11 can obtain a video / image through processes such as capture, synthesis, or generation of the video / image. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate a video / image. For example, a virtual video / image can be generated via a computer or the like, and in this case, the video / image capture process can be replaced by a process in which related data is generated.

[0051] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0052] The transmission unit 13 can transmit the encoded video / image information or data output in bitstream format to the receiving unit 21 of the decoding device 20 via a digital storage medium or a network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit 13 can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or the network and transmit it to the decoding unit 22.

[0053] The decoding unit 22 can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit 12.

[0054] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0055] Overview of Image Encoding Device

[0056] FIG. 2 is a diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure can be applied.

[0057] As shown in FIG. 2, the image encoding apparatus 100 can include an image division unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a "prediction unit". The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.

[0058] All or at least a part of the plurality of components constituting the image encoding device 100 can be realized by one hardware component (for example, an encoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.

[0059] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) in a QT / BT / TT (Quad-tree / Binary-tree / Ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units with a deeper depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and / or restoration described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). The prediction unit and the transform unit can be divided or partitioned from the final coding unit respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0060] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform prediction on a processing target block (current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy encoding unit 190. The information regarding the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0061] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of detail of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the setting. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0062] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU). The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on the peripheral blocks, and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the peripheral blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of the surrounding block is used as a motion vector predictor, and the motion vector difference and an indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0063] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described later. For example, the prediction unit can apply not only intra prediction or inter prediction for predicting the current block, but also apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for predicting the current block can be called CIIP (combined inter and intra prediction). In addition, the prediction unit can also perform intra block copy (IBC) for predicting the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block within the current picture at a position a predetermined distance away from the current block. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance.

[0064] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0065] The conversion unit 120 can apply a conversion technique to the residual signal to generate conversion coefficients (transform coefficients). For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen - Loeve Transform), GBT (Graph - Based Transform), or CNT (Conditionally Non - linear Transform). Here, GBT means the transform obtained from a graph when representing the relationship information between pixels as a graph. CNT means the transform obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can also be applied to pixel blocks having the same size of a square and can also be applied to blocks of variable size that are not square.

[0066] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bit stream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 130 can reorder the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0067] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 190 can also encode, together or separately, information necessary for video / image restoration (such as the values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (such as the encoded video / image information) can be transmitted or stored in the form of a bit stream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, the transmitted information, and / or the syntax elements mentioned in the present disclosure can be encoded through the above-described encoding procedure and included in the bit stream.

[0068] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.

[0069] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.

[0070] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next processing target block in the current picture and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0071] On the other hand, as will be described later, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding process.

[0072] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and can save the modified restored picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bit stream.

[0073] The modified restored picture transmitted to the memory 170 can be used as a reference picture by the inter prediction unit 180. When inter prediction is applied through this, the image encoding apparatus 100 can avoid prediction mismatches between the image encoding apparatus 100 and the image decoding apparatus, and can also improve the encoding efficiency.

[0074] The DPB in the memory 170 can save the modified restored picture for use as a reference picture by the inter prediction unit 180. The memory 170 can save the motion information of the block where the motion information in the current picture has been derived (or encoded) and / or the motion information of the block in the already restored picture. The saved motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 170 can save the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.

[0075] Overview of Image Decoding Device

[0076] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.

[0077] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transformation unit 230 can be included in a residual processing unit.

[0078] All or at least a part of the plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB and can be realized by a digital storage medium.

[0079] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 of FIG. 1 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be reproduced via a reproducing apparatus (not shown).

[0080] The image decoding apparatus 200 can receive the signal output from the image encoding apparatus of FIG. 1 in the form of a bitstream. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding apparatus can further use the information regarding the parameter set and / or the general constraint information for decoding an image. The signaling information, the received information, and / or the syntax elements referred to in the present disclosure can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values that have undergone entropy decoding in the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, of the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives the signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can also be provided as a component of the entropy decoding unit 210.

[0081] On the other hand, the image decoding device according to the present disclosure can be referred to as a video / image / picture decoding device. The image decoding device can also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.

[0082] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output the transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed in the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain the transform coefficients.

[0083] In the inverse conversion unit 230, the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0084] The prediction unit can perform prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0085] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described later, which is the same as described in the explanation of the prediction unit of the image encoding apparatus 100.

[0086] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can be similarly applied to the intra prediction unit 265.

[0087] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information regarding the prediction can include information indicating the mode (technique) of inter prediction for the current block.

[0088] The addition unit 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). The description of the addition unit 155 can be similarly applied to the addition unit 235.

[0089] On the other hand, as will be described later, LMCS (luma mapping with chroma scaling) can also be applied in the picture decoding process.

[0090] The filtering unit 240 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0091] The (modified) restored picture stored in the DPB of the memory 250 can be used as a reference picture in the inter prediction unit 260. The memory 250 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored pictures. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 250 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 265.

[0092] In this specification, the embodiments described in the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the image encoding apparatus 100 can be similarly or correspondingly applied to the filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of the image decoding apparatus 200.

[0093] Hereinafter, with reference to FIGS. 4 to 7, inter prediction encoding and inter prediction decoding will be described.

[0094] The image encoding / decoding device can derive a prediction sample by performing inter prediction in block units. Inter prediction can mean a prediction technique derived in a way that depends on data elements of pictures other than the current picture. When inter prediction is applied to the current block, a prediction block for the current block can be derived based on a reference block specified by a motion vector on a reference picture.

[0095] At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be derived based on the correlation of the motion information between the surrounding blocks and the current block, and the motion information can be derived in units of blocks, sub-blocks, or samples. At this time, the motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction type information. Here, the inter prediction type information can mean the directional information of the inter prediction. The inter prediction type information can indicate that the current block is predicted using any one of L0 prediction, L1 prediction, and Bi prediction.

[0096] When inter prediction is applied to the current block, the surrounding blocks of the current block can include spatial neighbouring blocks existing within the current picture and temporal neighbouring blocks existing in the reference picture. At this time, the reference picture including the reference block for the current block and the reference picture including the temporal neighbouring blocks may be the same or different. The temporal neighbouring blocks can be called collocated reference blocks, collocated coding units, etc. The reference picture including the temporal neighbouring blocks can be called a collocated picture (colPic).

[0097] On one hand, a motion information candidate list can be configured based on neighboring blocks of the current block. At this time, a flag or index information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block can be signaled.

[0098] The motion information can include L0 motion information and / or L1 motion information based on the inter prediction type. The motion vector in the L0 direction can be defined as the L0 motion vector or MVL0, and the motion vector in the L1 direction can be defined as the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be defined as L0 prediction, the prediction based on the L1 motion vector can be defined as L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector can be defined as bi prediction. Here, the L0 motion vector can mean the motion vector related to the reference picture list L0, and the L1 motion vector can mean the motion vector related to the reference picture list L1.

[0099] The reference picture list L0 can include, as reference pictures, pictures that are earlier than the current picture in the output order. The reference picture list L1 can include pictures that are later than the current picture in the output order. At this time, the earlier pictures can be defined as forward (reference) pictures, and the later pictures can be defined as backward (reference pictures). On the other hand, the reference picture list L0 can further include pictures that are later than the current picture in the output order. In this case, the earlier pictures can be indexed first within the reference picture list L0, and the later pictures can be indexed next. The reference picture list L1 can further include pictures that are earlier than the current picture in the output order. In this case, the later pictures can be indexed first within the reference picture list L1, and the earlier pictures can be indexed next. Here, the output order can correspond to the POC (picture order count) order.

[0100] FIG. 4 is a diagram exemplarily showing the configuration of an inter prediction unit that performs inter prediction coding according to the present disclosure.

[0101] For example, the inter prediction unit shown in FIG. 4 can correspond to the inter prediction unit 180 of the image coding apparatus in FIG. 2. The inter prediction unit 180 according to the present disclosure can include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The inter prediction unit 180 can receive, as inputs, the original picture to be coded and the reference pictures used for inter prediction. The prediction mode determination unit 181 can determine the prediction mode for the current block in the original picture. The motion information derivation unit 182 can derive the motion information for the current block. The prediction sample derivation unit 183 can derive a prediction sample by performing inter prediction on the current block. The prediction sample can be represented by the prediction block of the current block. The inter prediction unit 180 can output information regarding the prediction mode, information regarding the motion information, and the prediction sample.

[0102] For example, the inter-prediction unit 180 of the image encoding apparatus searches for a block similar to the current block within a certain area (search area) of the reference picture through motion estimation, and can derive a reference block whose difference from the current block is the minimum or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The image encoding apparatus can determine the mode applied to the current block among various prediction modes. The image encoding apparatus can compare the rate-distortion (RD) cost for the various prediction modes and determine the optimal prediction mode for the current block. However, the method by which the image encoding apparatus determines the prediction mode for the current block is not limited to the above example, and various methods can be used.

[0103] For example, the inter prediction mode for the current block can be determined to be at least one of a merge mode, a skip mode, a Motion Vector Prediction (MVP) mode, a Symmetric Motion Vector Difference (SMVD) mode, an affine mode, a Subblock-based merge mode, an Adaptive Motion Vector Resolution (AMVR) mode, a History-based Motion Vector Predictor (HMVP) mode, a Pair-wise average merge mode, a Merge mode with Motion Vector Differences (MMVD) mode, a Decoder side Motion Vector Refinement (DMVR) mode, a Combined Inter and Intra Prediction (CIIP) mode, and a Geometric Partitioning mode (GPM).

[0104] For example, when the skip mode or the merge mode is applied to the current block, the image encoding apparatus can derive merge candidates from the surrounding blocks of the current block and configure a merge candidate list using the derived merge candidates. Further, the image encoding apparatus can derive a reference block among the reference blocks pointed to by the merge candidates included in the merge candidate list, where the difference from the current block is the smallest or below a certain criterion. In this case, the merge candidate related to the derived reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the image decoding apparatus. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0105] As another example, when the MVP mode is applied to the current block, the image encoding apparatus can derive motion vector predictor (MVP) candidates from the surrounding blocks of the current block and configure an MVP candidate list using the derived MVP candidates. Further, the image encoding apparatus can use the motion vector of the selected MVP candidate among the MVP candidates included in the MVP candidate list as the MVP of the current block. In this case, for example, the motion vector indicating the reference block derived by the above-described motion estimation can be used as the motion vector of the current block, and among the MVP candidates, the MVP candidate having the motion vector with the smallest difference from the motion vector of the current block can be the selected MVP candidate. An MVD (motion vector difference), which is the difference obtained by subtracting the MVP from the motion vector of the current block, can be derived. In this case, the index information indicating the selected MVP candidate and the information regarding the MVD can be signaled to the image decoding apparatus. Also, when the MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and signaled to the image decoding apparatus separately.

[0106] FIG. 5 is a flowchart for explaining an encoding method based on inter prediction.

[0107] For example, the encoding method of FIG. 5 can be performed by the image encoding apparatus of FIG. 2. Specifically, step S510, step S520, and step S530 can be respectively performed by the inter prediction unit 180, the residual processing unit (for example, subtraction unit), and the entropy encoding unit 190. At this time, the prediction information and the residual information to be encoded can be derived by the inter prediction unit 180 and the residual processing unit respectively. The residual information can include information regarding the quantized transform coefficients for the residual samples. As described above, the residual samples are derived as transform coefficients via the conversion unit 120 of the image encoding apparatus, and the transform coefficients can be derived as quantized transform coefficients via the quantization unit 130. The information regarding the quantized transform coefficients can be encoded by the entropy encoding unit 190 via the residual coding procedure.

[0108] In step S510, the image encoding apparatus can perform an inter prediction on the current block. By performing the inter prediction, the image encoding apparatus can derive an inter prediction mode for the current block and motion information for the current block, and generate a prediction sample for the current block. Here, the determination of the inter prediction mode, the derivation of the motion information, and the generation procedure of the prediction sample may be performed simultaneously, or any one of the procedures may be performed prior to the other procedures.

[0109] In step S520, the image encoding apparatus can derive a residual sample based on the prediction sample. The image encoding apparatus can derive the residual sample based on the original sample of the current block and the prediction sample. For example, the residual sample can be derived by subtracting the corresponding prediction sample from the original sample.

[0110] In step S530, the image encoding device can encode image information including prediction information and residual information. The image encoding device can output the encoded image information in the form of a bitstream. The prediction information is information related to the prediction procedure and can include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information related to motion information. Among the prediction mode information, the skip flag is information indicating whether the skip mode is applied to the current block, and the merge flag is information indicating whether the merge mode is applied to the current block. Alternatively, the prediction mode information may be information indicating any one of a plurality of prediction modes, such as a mode index. When the skip flag and the merge flag are both 0, it can be determined that the MVP mode is applied to the current block. The information related to the motion information can include candidate selection information (e.g., merge index, mvp flag, or mvp index) that is information for deriving a motion vector. Among the candidate selection information, the merge index can be signaled when the merge mode is applied to the current block and can be information for selecting any one of the merge candidates included in the merge candidate list. Among the candidate selection information, the MVP flag or the MVP index can be signaled when the MVP mode is applied to the current block and can be information for selecting any one of the MVP candidates included in the MVP candidate list. Specifically, the MVP flag can be signaled using the syntax element mvp_l0_flag or mvp_l1_flag. Also, the information related to the motion information can include information related to the above-mentioned MVD and / or reference picture index information. Also, the information related to the motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information related to the residual samples.The residual information can include information regarding the quantized transform coefficients for the residual sample.

[0111] The output bitstream can be stored in a (digital) storage medium and transmitted to an image decoding device, or can also be transmitted to the image decoding device via a network.

[0112] On the other hand, as described above, the image encoding device can generate a reconstructed picture (a picture including reconstructed samples and reconstructed blocks) based on the reference sample and the residual sample. This is because the same prediction result as that performed by the image decoding device is derived by the image encoding device, and thereby the coding efficiency can be improved. Therefore, the image encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in the memory and utilize it as a picture for inter prediction. As described above, loop filtering procedures and the like can be further applied to the reconstructed picture.

[0113] FIG. 6 is a diagram exemplarily showing the configuration of an inter prediction unit that performs inter prediction decoding according to the present disclosure.

[0114] For example, the inter prediction unit shown in FIG. 6 can correspond to the inter prediction unit 260 of the image decoding apparatus in FIG. 3. The inter prediction unit 260 according to the present disclosure can include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. The inter prediction unit 260 can receive, as inputs, information regarding the prediction mode of the current block, information regarding the motion information of the current block, and a reference picture used for inter prediction. The prediction mode determination unit 261 can determine the prediction mode for the current block based on the information regarding the prediction mode. The motion information derivation unit 262 can derive motion information (such as a motion vector and / or a reference picture index) for the current block based on the information regarding the motion information. The prediction sample derivation unit 263 can derive a prediction sample by performing inter prediction on the current block. The prediction sample can be represented by the predicted block of the current block. The inter prediction unit 260 can output the derived prediction sample.

[0115] FIG. 7 is a flowchart for explaining a decoding method based on inter prediction.

[0116] For example, the decoding method in FIG. 7 can be performed by the image decoding apparatus in FIG. 3. The image decoding apparatus can perform operations corresponding to the operations performed by the image encoding apparatus. The image decoding apparatus can perform prediction on the current block based on the received prediction information and derive a prediction sample.

[0117] Specifically, steps S710 to S730 can be performed by the inter prediction unit 260, and the prediction information in step S710 and the residual information in step S740 can be obtained from the bitstream by the entropy decoding unit 210. Step S740 can be performed by the residual processing unit of the image decoding apparatus. Specifically, the inverse quantization unit 220 of the residual processing unit performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive the transform coefficients, and the inverse transform unit 230 of the residual processing unit can perform an inverse transform on the transform coefficients to derive the residual samples for the current block. Step S750 can be performed by the addition unit 235 or the restoration unit.

[0118] In step S710, the image decoding apparatus can determine the prediction mode for the current block based on the received prediction information. The image decoding apparatus can determine which inter prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0119] For example, based on the skip flag, it can be determined whether the skip mode is applied to the current block. Also, based on the merge flag, it can be determined whether the merge mode is applied to the current block or the MVP mode is determined. Alternatively, based on the mode index, any one of various inter prediction mode candidates can be selected. The inter prediction mode candidates can include the skip mode, the merge mode, and / or the MVP mode, or can include various inter prediction modes described later.

[0120] In step 720, the image decoding apparatus can derive the motion information of the current block based on the determined inter prediction mode. For example, when the skip mode or the merge mode is applied to the current block, the image decoding apparatus can construct a merge candidate list described later and select any one of the merge candidates included in the merge candidate list. The selection can be performed based on the candidate selection information (merge index) described above. The motion information of the current block can be derived using the motion information of the selected merge candidate. For example, the motion information of the selected merge candidate can be used as the motion information of the current block.

[0121] As another example, when the MVP mode is applied to the current block, the image decoding apparatus can construct an MVP candidate list and use the motion vector of the MVP candidate selected from among the MVP candidates included in the MVP candidate list as the MVP of the current block. The selection can be performed based on the candidate selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the MVP and the MVD of the current block. Also, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index in the reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.

[0122] In step S730, the image decoding apparatus can generate a prediction sample for the current block based on the motion information of the current block. In this case, the reference picture can be derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the samples of the reference block pointed to by the motion vector of the current block on the reference picture. Optionally, a prediction sample filtering procedure can be further performed on all or part of the prediction samples of the current block.

[0123] In step S740, the image decoding apparatus can generate a residual sample for the current block based on the received residual information.

[0124] In step S750, the image decoding apparatus can generate a restored sample for the current block based on the prediction sample and the residual sample, and generate a restored picture based on this. Thereafter, an in-loop filtering procedure or the like can be further applied to the restored picture.

[0125] Hereinafter, the motion information derivation step in the prediction mode will be described in more detail.

[0126] As described above, inter prediction can be performed using the motion information of the current block. The image encoding apparatus can derive the optimal motion information for the current block through a motion estimation procedure. The derived motion information can be signaled to the image decoding apparatus in various ways based on the inter prediction mode.

[0127] When the merge mode is applied to the current block, the motion information of the current block is not directly transmitted, and the motion information of the current block is derived using the motion information of neighboring blocks. Therefore, by transmitting flag information indicating that the merge mode is used and candidate selection information (e.g., merge index) indicating which neighboring blocks are used as merge candidates, the motion information of the current prediction block can be indicated. In the present disclosure, since the current block is a unit of prediction execution, the current block can be used in the same sense as the current prediction block, and the neighboring blocks can be used in the same sense as the neighboring prediction blocks.

[0128] The image encoding apparatus can search for a merge candidate block used to derive the motion information of the current block in order to perform the merge mode. For example, up to 5 merge candidate blocks can be used, but it is not limited thereto. The maximum number of the merge candidate blocks can be transmitted from a slice header or a tile group header, but it is not limited thereto. After finding the merge candidate blocks, the image encoding apparatus can generate a merge candidate list, and among these, the merge candidate block with the smallest RD cost can be selected as the final merge candidate block.

[0129] The present disclosure provides various examples for the merge candidate blocks constituting the merge candidate list. For example, 5 merge candidate blocks can be used for the merge candidate list. For example, 4 spatial merge candidates and 1 temporal merge candidate can be used.

[0130] FIG. 8 is a diagram illustrating neighboring blocks used as spatial merge candidates.

[0131] As shown in FIG. 8, the spatial peripheral blocks used as spatial merge candidates may include the lower left corner peripheral block A0, the left peripheral block A1, the upper right corner peripheral block B0, the upper peripheral block B1, and the upper left corner peripheral block B2 of the current block. However, this is merely an example, and in addition to the spatial peripheral blocks shown in FIG. 8, additional peripheral blocks such as the right peripheral block, the lower peripheral block, and the lower right corner peripheral block can also be used as the spatial peripheral blocks.

[0132] FIG. 9 is a diagram schematically showing a method for configuring a merge candidate list according to an example of the present disclosure.

[0133] The image encoding device / image decoding device can insert the spatial merge candidates derived by searching the spatial peripheral blocks of the current block into the merge candidate list (S910). The image encoding device / image decoding device can detect available blocks by searching the spatial peripheral blocks based on the priority, and derive the motion information of the detected blocks as the spatial merge candidates. For example, the image encoding device / image decoding device can search the five blocks shown in FIG. 8 in the order of A1, B1, B0, A0, B2, and sequentially index the available candidates to configure the merge candidate list.

[0134] Hereinafter, in the case of the merge mode and / or the skip mode, a method for inducing spatial candidates will be described more specifically. The spatial candidates can indicate the above-described spatial merge candidates.

[0135] The induction of spatial candidates can be performed based on spatially neighboring blocks. As an example, up to four spatial candidates can be induced from candidate blocks existing at the positions shown in FIG. 8. The order of inducing spatial candidates can be in the order of A1→B1→B0→A0→B2. However, the order of inducing spatial candidates is not limited to the above order. For example, it may be in the order of B1→A1→B0→A0→B2. The position that is last in order (position B2 in the above example) can be considered when at least one of the preceding four positions (in the above example, A1, B1, B0, and A0) is not available. At this time, the fact that a block at a predetermined position is not available can include the case where the corresponding block belongs to a different slice or a different tile from the current block, or the case where the block is an intra-predicted block. When spatial candidates are induced from the position that is first in order (A1 or B1 in the above example), a redundancy check can be performed on the spatial candidates at subsequent positions. For example, when the motion information of a subsequent spatial candidate is the same as the motion information of a spatial candidate already included in the merge candidate list, the subsequent spatial candidate can be excluded from the merge candidate list to improve the coding efficiency. The redundancy check performed on subsequent spatial candidates is not performed on as many candidate pairs as possible, but only on some candidate pairs, thereby reducing the computational complexity.

[0136] For example, when spatial candidates are induced in the order of B1 → A1 → B0 → A0 → B2, the redundancy check for the spatial candidate at the A1 position can be performed only for the spatial candidate at the B1 position. Also, the redundancy check for the spatial candidate at the B0 position can be performed only for the spatial candidate at the B1 position. Also, the redundancy check for the spatial candidate at the A0 position can be performed only for the spatial candidate at the A1 position. Finally, the redundancy check for the spatial candidate at the B2 position can be performed only for the spatial candidates at the B1 and A1 positions. However, it is not limited to this, and even when the order of inducing spatial candidates is changed, the redundancy check can be performed only for some candidate pairs as described above.

[0137] Referring to FIG. 9 again, the image encoding device / image decoding device can insert the temporal merge candidates derived by searching the temporal neighboring blocks of the current block into the merge candidate list (S920). The temporal neighboring blocks can be located on a reference picture that is a picture different from the current picture in which the current block is located. The reference picture on which the temporal neighboring blocks are located can be called a collocated picture or a col picture. The temporal neighboring blocks can be searched in the order of the lower right corner neighboring block (C0 block) and the lower right center block (C1 block) of the co-located block with respect to the current block on the col picture. That is, first, it is determined whether the C0 block is available. If the C0 block is available, temporal candidates can be derived based on the C0 block. If the C0 block is not available, temporal candidates can be derived based on the C1 block. For example, if the C0 block is an intra-predicted block or exists outside the current CTU row, it can be determined that the C0 block is not available. On the other hand, when motion data compression is applied to reduce the memory load, specific motion information can be stored as representative motion information for each certain storage unit for the col picture. In this case, it is not necessary to store the motion information for all blocks within the certain storage unit, and thus the motion data compression effect can be obtained. In this case, the certain storage unit can be predetermined, for example, in units of 16×16 samples, or 8×8 samples, or the size information for the certain storage unit can also be signaled from the image encoding device to the image decoding device. When the motion data compression is applied, the motion information of the temporal neighboring blocks can be replaced with the representative motion information of the certain storage unit in which the temporal neighboring blocks are located.That is, in this case, from the perspective of implementation, instead of the prediction block lock located at the coordinates of the time surrounding block, based on the coordinates (upper left sample position) of the time surrounding block, after arithmetic right shifting by a certain value, the motion information of the prediction block covering the position after arithmetic left shifting can be used to derive the time merge candidate. For example, the certain storage unit is 2. n ×2 n When it is a sample unit, if the coordinates of the time surrounding block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>n)<<n), (yTnb>>n)<<n)) can be used for the time merge candidate. Specifically, for example, when the certain storage unit is 16×16 sample units, if the coordinates of the time surrounding block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>4)<<4), (yTnb>>4)<<4)) can be used for the time merge candidate. Or, for example, when the certain storage unit is 8×8 sample units, if the coordinates of the time surrounding block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>3)<<3), (yTnb>>3)<<3)) can be used for the time merge candidate.

[0138] Hereinafter, in the case of the merge mode and / or skip mode, the method of inducing time candidates will be described more specifically. The time candidates can indicate the time merge candidates described above. Also, the motion vector of the time candidates can also correspond to the time candidates in the MVP mode.

[0139] Only one candidate can be included in the merge candidate list as a time candidate. In the process of deriving the time candidate, the motion vector of the time candidate can be scaled. For example, the scaling can be performed based on a collocated CU (colocated CU, hereinafter referred to as "col block") belonging to a collocated reference picture (hereinafter referred to as "col picture"). More specifically, the scaling can be performed based on the distance (tb) between the reference picture of the current block and the current picture and the distance (td) between the reference picture of the col block and the col picture. The tb and td can be represented by values corresponding to the difference in POC (Picture Order Count) between pictures. The reference picture list used for deriving the col block can be explicitly signaled in the slice header.

[0140] Referring to FIG. 9 again, the image encoding device / image decoding device can check whether the number of current merge candidates is less than the number of maximum merge candidates (S930). The number of the maximum merge candidates can be predefined or signaled from the image encoding device to the image decoding device. For example, the image encoding device can generate information regarding the number of the maximum merge candidates, encode it, and transmit it to the image decoding device in the form of a bit stream. When all of the number of the maximum merge candidates are satisfied, the subsequent candidate addition process (S940) cannot be performed.

[0141] If the verification result of step S930 indicates that the number of current merge candidates is less than the maximum number of merge candidates, the image encoding device / image decoding device can insert additional merge candidates into the merge candidate list after deriving them based on a predetermined method (S940). The additional merge candidates can include, for example, at least one of history based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate(s) (when the slice / tile group type of the current slice / tile group is of type B), and / or zero vector merge candidate(s).

[0142] If the verification result of step S930 indicates that the number of current merge candidates is not less than the maximum number of merge candidates, the image encoding device / image decoding device can terminate the configuration of the merge candidate list. In this case, the image encoding device can select the optimal merge candidate among the merge candidates constituting the merge candidate list based on the RD cost, and can signal candidate selection information (e.g., merge candidate index, merge index) indicating the selected merge candidate to the image decoding device. The image decoding device can select the optimal merge candidate based on the merge candidate list and the candidate selection information.

[0143] As described above, the motion information of the selected merge candidate can be used as the motion information of the current block, and the predicted sample of the current block can be derived based on the motion information of the current block. The image encoding device can derive the residual sample of the current block based on the predicted sample, and can signal the residual information regarding the residual sample to the image decoding device. As described above, the image decoding device can generate a restored sample based on the residual sample derived based on the residual information and the predicted sample, and can generate a restored picture based on this.

[0144] When a skip mode is applied to the current block, the motion information of the current block can be derived in the same way as when a merge mode was applied previously. However, when the skip mode is applied, the residual signal for the block is omitted. Thus, the predicted sample can be used immediately as a restored sample. The skip mode can be applied, for example, when the value of cu_skip_flag is 1.

[0145] Hereinafter, a method for deriving History-based candidates in the case of the merge mode and / or the skip mode will be described. The History-based candidates can be represented as History-based merge candidates.

[0146] The History-based candidates can be added to the merge candidate list after the spatial candidates and the temporal candidates are added to the merge candidate list. For example, the motion information of previously encoded / decoded blocks is stored in a table and can be used as History-based candidates for the current block. The table can store a plurality of History-based candidates during the encoding / decoding process. The table can be initialized when a new CTU row starts. Initializing the table can mean that all the History-based candidates stored in the table are deleted and the table becomes empty. Whenever an inter-predicted block exists, the related motion information can be added to the table as the last entry. At this time, the inter-predicted block can be a block predicted based on sub-blocks. The motion information added to the table can be used as a new History-based candidate.

[0147] The history-based candidate table can have a predetermined size. For example, the size can be 5. At this time, the table can store up to 5 history-based candidates. When a new candidate is added to the table, a restricted first-in-first-out (FIFO) rule can be applied where a redundancy check is first performed to see if the same candidate already exists in the table. If the same candidate already exists in the table, the same candidate is deleted from the table, and the positions of all subsequent history-based candidates can be moved forward.

[0148] History-based candidates can be used in the process of constructing the merge candidate list. At this time, the history-based candidates most recently included in the table are sequentially checked and can be included at positions after the time candidates in the merge candidate list. When a history-based candidate is included in the merge candidate list, a redundancy check can be performed with the spatial or temporal candidates already included in the merge candidate list. If a history-based candidate overlaps simultaneously with a spatial or temporal candidate already included in the merge candidate list, the history-based candidate can be excluded from the merge candidate list. The redundancy check can be simplified as follows to reduce the computational amount.

[0149] The number of history-based candidates used for generating the merge candidate list can be set to (N <= 4)? M : (8 - N). At this time, N indicates the number of candidates already included in the merge candidate list, and M indicates the number of available history-based candidates stored in the table. That is, when the merge candidate list contains 4 or fewer candidates, the number of history-based candidates used for generating the merge candidate list is M, and when the merge candidate list contains N candidates more than 4, the number of history-based candidates used for generating the merge candidate list can be set to (8 - N).

[0150] When the total number of available merge candidates reaches (the maximum allowable number of merge candidates - 1), the construction of the merge candidate list using History-based candidates can be terminated.

[0151] Hereinafter, in the case of the merge mode and / or the skip mode, a method for deriving Pair-wise average candidates will be described. Pair-wise average candidates can be expressed as Pair-wise average merge candidates or Pair-wise candidates.

[0152] Pair-wise average candidates can be generated by obtaining already defined candidate pairs from the candidates included in the merge candidate list and averaging them. The already defined candidate pairs are {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, and the numbers constituting each candidate pair can be the indices of the merge candidate list. That is, the already defined candidate pair (0,1) means a pair of the candidate at index 0 and the candidate at index 1 of the merge candidate list, and the Pair-wise average candidate can be generated by averaging the candidate at index 0 and the candidate at index 1. The derivation of Pair-wise average candidates can be performed in the order of the already defined candidate pairs. That is, after deriving the Pair-wise average candidate for the candidate pair (0,1), the process of deriving the Pair-wise average candidate can be performed in the order of the candidate pair (0,2) and the candidate pair (1,2). The process of deriving the Pair-wise average candidate can be performed until the construction of the merge candidate list is completed. For example, the process of deriving the Pair-wise average candidate can be performed until the number of merge candidates included in the merge candidate list reaches the maximum number of merge candidates.

[0153] Pair-wise average candidates can be calculated individually for each of the reference picture lists. If two motion vectors are available for one reference picture list (L0 list or L1 list), the average of these two motion vectors can be calculated. At this time, even if the two motion vectors point to different reference pictures from each other, the average of the two motion vectors can be performed. If only one motion vector is available for one reference picture list, the available motion vector can be used as the motion vector of the Pair-wise average candidate. If not all two motion vectors are available for one reference picture list, it can be determined that the reference picture list is not valid.

[0154] Even after the Pair-wise average candidate is included in the merge candidate list, if the configuration of the merge candidate list is not completed, zero vectors can be added to the merge candidate list until the maximum number of merge candidates is reached.

[0155] When the MVP mode is applied to the current block, a motion vector predictor (MVP) candidate list can be generated using the motion vectors of the restored spatial neighboring blocks (e.g., the neighboring blocks shown in FIG. 8) and / or the motion vectors corresponding to the temporal neighboring blocks (or Col blocks). That is, the motion vectors of the restored spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks can be used as candidates for the motion vector predictor of the current block. When dual prediction is applied, an MVP candidate list for L0 motion information derivation and an MVP candidate list for L1 motion information derivation can be individually generated and used. The prediction information (or information related to prediction) for the current block can include candidate selection information (e.g., an MVP flag or an MVP index) indicating the optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the MVP candidate list. At this time, the prediction unit can select the motion vector predictor of the current block from among the motion vector predictor candidates included in the MVP candidate list using the candidate selection information. The prediction unit of the image encoding apparatus can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can encode this and output it in the form of a bitstream. That is, the MVD can be obtained as the value obtained by subtracting the motion vector predictor from the motion vector of the current block. The prediction unit of the image decoding apparatus can obtain the motion vector difference included in the information related to the prediction, and can derive the motion vector of the current block through addition of the motion vector difference and the motion vector predictor. The prediction unit of the image decoding apparatus can obtain or derive a reference picture index indicating a reference picture, etc. from the information related to the prediction.

[0156] FIG. 10 is a diagram schematically showing a method for configuring a motion vector predictor candidate list according to an example of the present disclosure.

[0157] First, the spatial candidate blocks of the current block can be searched, and the available candidate blocks can be inserted into the MVP candidate list (S1010). Then, it is determined whether the number of MVP candidates included in the MVP candidate list is less than two (S1020). If it is two, the configuration of the MVP candidate list can be completed.

[0158] In step S1020, if the number of available spatial candidate blocks is less than two, the temporal candidate blocks of the current block can be searched, and the available candidate blocks can be inserted into the MVP candidate list (S1030). If the temporal candidate blocks are not available, the zero motion vector can be inserted into the MVP candidate list (S1040), thereby completing the configuration of the MVP candidate list.

[0159] On the other hand, when the MVP mode is applied, the reference picture index can be explicitly signaled. In this case, the picture index (refidxL0) for L0 prediction and the reference picture index (refidxL1) for L1 prediction can be separately signaled. For example, when the MVP mode is applied and bi-prediction is applied, the information regarding refidxL0 and the information regarding refidxL1 can both be signaled.

[0160] As described above, when the MVP mode is applied, the information regarding the MVD derived from the image encoding device can be signaled to the image decoding device. The information regarding the MVD can include, for example, the information indicating the x and y components of the MVD absolute value and the sign. In this case, whether the MVD absolute value is greater than 0 and greater than 1, the information indicating the rest of the MVD can be signaled step by step. For example, the information indicating whether the MVD absolute value is greater than 1 can be signaled only when the value of the flag information indicating whether the MVD absolute value is greater than 0 is 1.

[0161] The following provides a detailed description of the affine mode, which is an example of the inter prediction mode. In conventional video encoding / decoding systems, only one motion vector was used to represent the motion information of the current block. However, such a method had the problem that it could only represent the optimal motion information at the block level and could not represent the optimal motion information at the pixel level. To solve such a problem, an affine mode that defines the motion information of a block at the pixel level was proposed. According to the affine mode, the motion vectors for pixels / or sub-blocks of the block can be determined using two to four motion vectors associated with the current block.

[0162] Compared with the conventional motion information being represented using the translation (or displacement) of pixel values, in the affine mode, the motion information for each pixel can be represented using at least one of translation, scaling, rotation, and shear. Among them, the affine mode in which the motion information for each pixel is represented using displacement, scaling, and rotation can be defined as a similarity or simplified affine mode. The affine mode in the following description can mean a similarity or simplified affine mode.

[0163] The motion information in the affine mode can be represented using two or more CPMVs (Control Point Motion Vectors). The motion vector for a specific pixel position of the current block can be derived using CPMVs. At this time, the set of motion vectors for each pixel and / or sub-block of the current block can be defined as an Affine Motion Vector Field (Affine MVF).

[0164] FIG. 11 is a diagram for explaining the 4-parameter model of the affine mode.

[0165] FIG. 12 is a diagram for explaining the 6-parameter model of the affine mode.

[0166] When the affine mode is applied to the current block, the affine MVF can be derived using one of the 4-parameter model and the 6-parameter model. At this time, as shown in FIG. 11, the 4-parameter model can mean a model type in which two CPMVs (v0, v1) are used. Also, as shown in FIG. 12, the 6-parameter model can mean a model type in which three CPMVs (v0, v1, v2) are used.

[0167] When the position of the current block is defined as (x, y), the motion vector based on the pixel position can be derived according to the following Equation 1 or 2. For example, the motion vector by the 4-parameter model can be derived according to Equation 1, and the motion vector by the 6-parameter model can be derived according to Equation 2.

[0168]

Equation

[0169]

Equation

[0170] In Equation 1 and Equation 2, mv0 = {mv_0x, mv_0y} is the CPMV at the upper left corner position of the current block, v1 = {mv_1x, mv_1y} is the CPMV at the upper right corner position of the current block, and mv2 = {mv_2} can be the CPMV at the lower left position of the current block. Here, W and H respectively correspond to the width and height of the current block, and mv = {mv_x, mv_y} can mean the motion vector of the pixel position {x, y}.

[0171] In the symbolization / decryption process, the affine MVF can be determined in pixel units and / or predefined sub-block units. When the affine MVF is determined in pixel units, a motion vector can be derived based on each pixel value. On the other hand, when the affine MVF is determined in sub-block units, the motion vector of the block can be derived based on the central pixel value of the sub-block. The central pixel value can mean a virtual pixel existing at the center of the sub-block, or can mean the lower right pixel among the four pixels existing at the center. Also, the central pixel value can be a specific pixel within the sub-block that represents the sub-block. In the present disclosure, the case where the affine MVF is determined in 4×4 sub-block units will be described. However, this is for convenience of explanation, and the size of the sub-block can be variously changed.

[0172] That is, when Affine prediction is available, the motion models applicable to the current block can include three types: the Translational motion model (parallel translation motion model), the 4-parameter affine motion model, and the 6-parameter affine motion model. Here, the Translational motion model can indicate a model in which a conventional block unit motion vector is used, the 4-parameter affine motion model can indicate a model in which two CPMVs are used, and the 6-parameter affine motion model can indicate a model in which three CPMVs are used. The affine mode can be classified into detailed modes by the method of encoding / decrypting motion information. As an example, the affine mode can be subdivided into the affine MVP mode and the affine merge mode.

[0173] When the affine merge mode is applied to the current block, the CPMV can be derived from the neighboring blocks of the current block encoded / decoded in the affine mode. If at least one of the neighboring blocks of the current block is encoded / decoded in the affine mode, the affine merge mode can be applied to the current block. That is, when the affine merge mode is applied to the current block, the CPMV of the current block can be derived using the CPMV of the neighboring blocks. For example, the CPMV of the neighboring block can be determined as the CPMV of the current block, or the CPMV of the current block can be derived based on the CPMV of the neighboring block. When the CPMV of the current block is derived based on the CPMV of the neighboring block, at least one of the encoding parameters of the current block or the neighboring block can be used. For example, the CPMV of the neighboring block can be modified based on the size of the neighboring block and the size of the current block and used as the CPMV of the current block.

[0174] On the other hand, in the case of affine merge where the MV is derived in sub-block units, it can be called the sub-block merge mode, which can be indicated by a merge_subblock_flag having a first value (e.g., "1"). In this case, the affine merging candidate list described later can also be called the subblock merging candidate list. In this case, the subblock merging candidate list can further include candidates derived by SbTMVP described later. In this case, the candidate derived by the sbTMVP can be used as the candidate at the 0th index of the subblock merging candidate list. In other words, the candidate derived by the sbTMVP can be positioned ahead of the inherited affine candidates and the constructed affine candidates described later within the subblock merging candidate list.

[0175] As an example, an affine mode flag can be defined to indicate whether the affine mode can be applied to the current block. This can be signaled at at least one of the upper levels of the current block, such as sequence, picture, slice, tile, tile group, block, etc. For example, the affine mode flag can be named sps_affine_enabled_flag.

[0176] When the affine merge mode is applied, an affine merge candidate list can be constructed for CPMV derivation of the current block. At this time, the affine merge candidate list can include at least one of an inherited affine merge candidate, a combined affine merge candidate, and a zero merge candidate. The inherited affine merge candidate can mean a candidate derived using the CPMV of the surrounding block when the surrounding block of the current block is encoded / decoded in the affine mode. The combined affine merge candidate can mean a candidate in which each CPMV is derived based on the motion vectors of the surrounding blocks of each CP (Control Point). On the other hand, the zero merge candidate can mean a candidate consisting of a CPMV of size 0. In the following description, CP can mean a specific position of a block used to derive the CPMV. For example, the CP can be each vertex position of the block.

[0177] FIG. 13 is a diagram for explaining a method of generating an affine merge candidate list.

[0178] Referring to the flowchart of FIG. 13, affinity merge candidates can be added to the affinity merge candidate list in the order of inheritance affinity merge candidates (S1310), combination affinity merge candidates (S1320), and zero merge candidates (S1330). The zero merge candidate can be added when the number of candidates included in the candidate list does not satisfy the maximum number of candidates even though all the inheritance affinity merge candidates and combination affinity merge candidates have been added to the affinity merge candidate list. At this time, the zero merge candidate can be added until the number of candidates in the affinity merge candidate list satisfies the maximum number of candidates.

[0179] FIG. 14 is a diagram for explaining a method of inducing inheritance affinity merge candidates from peripheral blocks.

[0180] As an example, up to two inheritance affinity merge candidates can be induced, and each candidate can be induced based on at least one of the left peripheral block and the upper peripheral block. The peripheral blocks for inducing the inheritance affinity merge candidates will be described with reference to FIG. 8. The inheritance affinity merge candidate induced based on the left peripheral block is induced based on at least one of A0 and A1, and the inheritance affinity merge candidate induced based on the upper peripheral block can be induced based on at least one of B0, B1, and B2. At this time, the scan order of each peripheral block can be, but is not limited to, the order from A0 to A1, and the order from B0 to B1, B2. The inheritance affinity merge candidate can be induced based on the first available peripheral block in the scan order for each of the left and upper sides. In this case, redundancy checking can be omitted between the candidates induced from the left peripheral block and the upper peripheral block.

[0181] As an example, as shown in FIG. 14, when the left peripheral block A is encoded / decoded in the affine mode, at least one of the motion vectors v2, v3, and v4 corresponding to the CP of the peripheral block A can be derived. When the peripheral block A is encoded / decoded via a 4-parameter affine model, the inherited affine merge candidate can be derived using v2 and v3. On the other hand, when the peripheral block A is encoded / decoded via a 6-parameter affine model, the inherited affine merge candidate can be derived using v2, v3, and v4.

[0182] FIG. 15 is a diagram for explaining a peripheral block for deriving a combined affine merge candidate.

[0183] The combined affine candidate can mean a candidate in which the CPMV is derived using a combination of general motion information of the peripheral blocks. The motion information for each CP can be derived using the spatial or temporal peripheral blocks of the current block. In the following description, CPMVk can mean a motion vector representing the k-th CP. As an example, referring to FIG. 15, CPMV1 can be determined as the first available motion vector among the motion vectors of B2, B3, and A2, and the scan order at this time can be in the order of B2, B3, A2. CPMV2 can be determined as the first available motion vector among the motion vectors of B1 and B0, and the scan order at this time can be in the order of B1, B0. CPMV3 can be determined as the first available motion vector among the motion vectors of A1 and A0, and the scan order at this time can be in the order of A1, A0. When TMVP application is possible for the current block, CPMV4 can be determined as the motion vector of the temporal peripheral block T.

[0184] After four motion vectors for each CP are derived, combination affine merge candidates can be derived based on them. The combination affine merge candidates can be configured to include at least two motion vectors selected from among the four motion vectors for each derived CP. As an example, the combination affine merge candidates can be configured with at least one in the order of {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, and {CPMV1, CPMV3}. The combination affine candidates consisting of three motion vectors can be candidates for a 6-parameter affine model. In contrast, the combination affine candidates consisting of two motion vectors can be candidates for a 4-parameter affine model. To avoid the scaling process of the motion vectors, if the reference picture indexes of the CPs are different, the combinations of the related CPMVs can be ignored and not used for deriving the combination affine candidates.

[0185] When the affine MVP mode is applied to the current block, the image encoding device can derive two or more CPMV predictors and CPMV for the current block, and based on this, derive CPMV differences. At this time, the CPMV differences can be signaled from the encoding device to the decoding device. The image decoding device can derive a CPMV predictor for the current block, restore the signaled CPMV differences, and then derive the CPMV of the current block based on the CPMV predictor and the CPMV differences.

[0186] On the other hand, the affine MVP mode can be applied to the current block only when the affine merge mode or the sub-block-based TMVP is not applied to the current block. On the other hand, the affine MVP mode can also be expressed as the affine CP MVP mode.

[0187] When an affine MVP is applied to the current block, an affine mvp candidate list can be constructed for deriving the CPMV for the current block. Here, the affine MVP candidate list can include at least one of an inherited affine MVP candidate, a combined affine MVP candidate, a translational affine MVP candidate, and a zero MVP candidate.

[0188] At this time, the inherited affine MVP candidate can mean a candidate derived based on the CPMV of the surrounding block when the surrounding blocks of the current block are encoded / decoded in affine mode. The combined affine MVP candidate can mean a candidate derived by generating a CPMV combination based on the motion vectors of the surrounding blocks of the CP. The zero MVP candidate can mean a candidate consisting of a CPMV with a value of 0. Since the derivation methods and characteristics of the inherited affine MVP candidate and the combined affine MVP candidate are the same as those of the inherited affine candidate and the combined affine candidate described above, the description is omitted.

[0189] When the maximum number of candidates in the affine MVP candidate list is 2, the combined affine MVP candidate, the translational affine MVP candidate, and the zero MVP candidate can be added when the current number of candidates is less than 2. In particular, the translational affine MVP candidate can be derived according to the following order.

[0190] As an example, when the number of candidates included in the affine MVP candidate list is less than 2 and the CPMV0 of the combined affine MVP candidate is valid, the CPMV0 can be used as an affine MVP candidate. That is, an affine MVP candidate in which the motion vectors of CP0, CP1, and CP2 are all CPMV0 can be added to the affine MVP candidate list.

[0191] Next, if the number of candidates in the candidate list of affine MVP is less than 2 and CPMV1 of the combined affine MVP candidate is valid, CPMV1 can be used as an affine MVP candidate. That is, an affine MVP candidate in which the motion vectors of CP0, CP1, and CP2 are all CPMV1 can be added to the affine MVP candidate list.

[0192] Next, if the number of candidates in the candidate list of affine MVP is less than 2 and CPMV2 of the combined affine MVP candidate is valid, CPMV2 can be used as an affine MVP candidate. That is, an affine MVP candidate in which the motion vectors of CP0, CP1, and CP2 are all CPMV2 can be added to the affine MVP candidate list.

[0193] Despite the above conditions, if the number of candidates in the affine MVP candidate list is less than 2, the TMVP (temporal motion vector predictor) of the current block can be added to the affine MVP candidate list.

[0194] Despite the addition of translational affine MVP candidates, if the number of candidates in the affine MVP candidate list is less than 2, zero MVP candidates can be added to the affine MVP candidate list.

[0195] FIG. 16 is a diagram for explaining a method of generating an affine MVP candidate list.

[0196] Referring to the flowchart of FIG. 16, candidates can be added to the affine MVP candidate list in the order of inherited affine MVP candidates (S1610), combined affine MVP candidates (S1620), translational affine MVP candidates (S1630), and zero MVP candidates (S1640). As described above, steps S1620 to S1640 can be performed according to whether the number of candidates included in the affine MVP candidate list is less than 2 at each step.

[0197] The scanning order of the inherited affine MVP candidates can be the same as that of the inherited affine merge candidates. However, in the case of the inherited affine MVP candidates, only the neighboring blocks that refer to the same reference picture as the reference picture of the current block can be considered. When adding the inherited affine MVP candidates to the affine MVP candidate list, redundancy checking can be skipped.

[0198] To derive the combined affine MVP candidates, only the spatial neighboring blocks shown in FIG. 15 can be considered. Also, the scanning order of the combined affine MVP candidates can be the same as that of the combined affine merge candidates. Further, to derive the combined affine MVP candidates, the reference picture index of the neighboring blocks is checked, and the first neighboring block that is inter-coded and refers to the same reference picture as the reference picture of the current block in the said scanning order can be used.

[0199] Hereinafter, a sub-block based TMVP mode, which is an example of the inter-prediction mode, will be described in detail. According to the sub-block based TMVP mode, since a motion vector field (MVF) for the current block is derived, motion vectors can be derived in sub-block units.

[0200] Unlike the conventional TMVP mode which is performed in coding unit units, for the coding unit to which the sub-block based TMVP mode is applied, encoding / decoding of motion vectors can be performed in sub-coding unit units. Also, according to the conventional TMVP mode, a temporal motion vector is derived from a collocated block, whereas the sub-block based TMVP mode can derive a motion vector field from the reference block indicated by the motion vector derived from the neighboring blocks of the current block. Hereinafter, the motion vector derived from the neighboring blocks can be called the motion shift or representative motion vector of the current block.

[0201] FIG. 17 is a diagram for explaining the peripheral blocks of the sub-block based TMVP mode.

[0202] When the sub-block based TMVP mode is applied to the current block, the peripheral blocks for determining the motion shift can be determined. As an example, the scan of the peripheral blocks for determining the motion shift can be performed in the order of the A1, B1, B0, A0 blocks in FIG. 17. As another example, the peripheral blocks for determining the motion shift can be limited to specific peripheral blocks of the current block. For example, the peripheral blocks for determining the motion shift can always be determined as the A1 block. When the peripheral block has a motion vector that references the col picture, the motion vector can be determined as the motion shift. The motion vector determined as the motion shift can also be called the temporal motion vector. On the other hand, when the above-described motion vector cannot be derived from the peripheral block, the motion shift can be set to (0, 0).

[0203] FIG. 18 is a diagram for explaining a method of deriving a motion vector field according to the sub-block based TMVP mode.

[0204] Next, the reference block on the collocated picture indicated by the motion shift can be determined. For example, by adding the motion shift to the coordinates of the current block, sub-block based motion information (motion vector, reference picture index) can be obtained from the col picture. In the example shown in FIG. 18, assume that the motion shift is the motion vector of block A1. By applying the motion shift to the current block, the sub-blocks (col sub-blocks) in the col picture corresponding to each sub-block constituting the current block can be identified. Then, using the motion information of the corresponding sub-blocks (col sub-blocks) of the col picture, the motion information of each sub-block of the current block can be derived. For example, the motion information of the corresponding sub-block can be obtained from the central position of the corresponding sub-block. At this time, the central position can be the position of the lower right sample among the four samples located at the center of the corresponding sub-block. If the motion information of a specific sub-block of the col block corresponding to the current block is not available, the motion information of the central sub-block of the col block can be determined as the motion information of the said sub-block. When the motion information of the corresponding sub-block is derived, similar to the above-described TMVP process, the motion vector and reference picture index of the current sub-block can be switched. That is, when the sub-block based motion vector is derived, the scaling of the motion vector can be performed in consideration of the POC of the reference picture of the reference block.

[0205] As described above, the sub-block based TMVP candidates for the current block can be derived using the motion vector field or motion information of the current block derived based on sub-blocks.

[0206] Hereinafter, the merge candidate list composed of sub-blocks is defined as the sub-block unit merge candidate list. The above-described affine merge candidates and sub-block based TMVP candidates can be merged to form the sub-block unit merge candidate list.

[0207] On one hand, a sub-block based TMVP mode flag can be defined to indicate whether the sub-block based TMVP mode can be applied to the current block. This can be signaled at at least one of the upper levels of the current block, such as sequence, picture, slice, tile, tile group, block, etc. For example, the sub-block based TMVP mode flag can be named sps_sbtmvp_enabled_flag. If the sub-block based TMVP mode is applicable to the current block, the sub-block based TMVP candidates can be added to the sub-block unit merge candidate list first. Thereafter, the affine merge candidates can be added to the sub-block unit merge candidate list. On the other hand, the maximum number of candidates that can be included in the sub-block unit merge candidate list can be signaled. As an example, the maximum number of candidates that can be included in the sub-block unit merge candidate list can be 5.

[0208] The size of the sub-block used for deriving the sub-block unit merge candidate list can be signaled or already set to M×N. For example, M×N can be 8×8. Therefore, the affine mode or the sub-block based TMVP mode can be applied to the current block only when the size of the current block is 8×8 or more.

[0209] Hereinafter, an embodiment of the prediction execution method of the present disclosure will be described. The following prediction execution method can be performed in step S510 of FIG. 5 or step S730 of FIG. 7.

[0210] Based on the motion information derived according to the prediction mode, a predicted block for the current block can be generated. The predicted block (prediction block) can include the prediction sample (prediction sample array) of the current block. When the motion vector of the current block refers to the fractional sample unit, an interpolation procedure can be performed, whereby the prediction sample of the current block can be derived based on the reference sample of the fractional sample unit in the reference picture. When affine interpolation prediction is applied to the current block, prediction samples can be generated based on the sample / sub-block unit MV. When bi-prediction is applied, the prediction sample derived based on the L0 prediction (i.e., prediction using the reference picture in the reference picture list L0 and MVL0) and the prediction sample derived based on the L1 prediction (i.e., prediction using the reference picture in the reference picture list L1 and MVL1) can be used as the prediction sample of the current block by weighted sum or weighted average (by phase). When bi-prediction is applied, if the reference picture used for the L0 prediction and the reference picture used for the L1 prediction are located in different temporal directions with respect to the current picture (i.e., when it corresponds to bidirectional prediction while being bi-prediction), this can be called true bi-prediction.

[0211] In the image decoding device, a restored sample and a restored picture can be generated based on the derived prediction sample, and then procedures such as in-loop filtering can be performed. Also, in the image encoding device, a residual sample can be derived based on the derived prediction sample, and encoding of the image information including the prediction information and the residual information can be performed.

[0212] As described above, when dual prediction is currently applied to a block, a prediction sample can be derived based on a weighted average. In this case, the weights for performing the weighted average can be determined based on the weight index derived at the CU level. Conventionally, a dual prediction signal (i.e., a dual prediction sample) could be derived through a simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, the dual prediction sample was derived by averaging the L0 prediction sample based on the L0 reference picture and MVL0 and the L1 prediction sample based on the L1 reference picture and MVL1. However, according to the present disclosure, when dual prediction is applied, a dual prediction signal (dual prediction sample) can be derived through a weighted average of the L0 prediction signal and the L1 prediction signal as shown in Equation 3. Such dual prediction can be referred to as Bi-prediction with CU-level weight (BCW).

[0213] [Number]

[0214] In Equation 3 above, P bi-pred represents the dual prediction signal (dual prediction block) derived by weighted average, and P 0 and P 1 represent the L0 prediction sample (L0 prediction block) and the L1 prediction sample (L1 prediction block), respectively. Also, (8 - w) and w represent the weights applied to P 0 and P 1 , respectively.

[0215] In the generation of dual-prediction signals by weighted average, five weights are acceptable. For example, the weight w can be selected from {-2, 3, 4, 5, 10}. For each of the dual-predicted CUs, the weight w can be determined in one of two ways. As a first method of these two methods, when the current CU is not in the merge mode (non-merge CU), a weight index can be signaled together with the motion vector difference. For example, the bitstream can include information about the weight index after the information about the motion vector difference. As a second method of these two methods, when the current CU is in the merge mode (merge CU), the weight index can be derived from neighboring blocks based on the merge candidate index (merge index).

[0216] The generation of dual-prediction signals by weighted average can be restricted to be applied only to CUs of a size that includes 256 or more samples (luma component samples). That is, dual-prediction by weighted average can be performed only on CUs whose product of the width and height of the current block is 256 or more. Also, as described above, one of the five weights may be used for the weight w, or one of a different number of weights may be used. For example, depending on the characteristics of the current image, five weights can be used for low-delay pictures, and three weights can be used for non-low-delay pictures. At this time, the three weights can be {3, 4, 5}.

[0217] The image coding device can determine the weight index by applying a fast search algorithm without significantly increasing the complexity. In this case, the fast search algorithm can be summarized as follows. In the following, unequal weight means that the weights applied to P 0 and P 1 are not equal. Also, equal weight means that the weights applied to P 0 and P1 It can be meant that the weights applied thereto are equal.

[0218] - When an AMVR mode in which the resolution of motion vectors is adaptively changed is applied together, if the current picture is a low-delay picture, only unequal weights can be conditionally checked for each of the 1-pel motion vector resolution and the 4-pel motion vector resolution.

[0219] - When the affine mode is applied together and the affine mode is selected as the optimal mode of the current block, the image coding device can perform affine ME (motion estimation) for each of the unequal weights.

[0220] - When the two reference pictures used for bi-prediction are the same, only unequal weights can be conditionally checked.

[0221] - Unequal weights can be not checked when a predetermined condition is satisfied. The predetermined condition may be a condition based on the POC distance between the current picture and the reference picture, the quantization parameter (QP), the temporal level, etc.

[0222] The weight index of BCW can be coded using one context-coded bin and one or more subsequent bypass-coded bins. The first context-coded bin indicates whether equal weight is used. If unequal weight is used, additional bins can be bypass-coded and signaled. The additional bins can be signaled to indicate which weight is used.

[0223] Weighted prediction (WP) is a tool for efficiently encoding images including fading. According to weighted prediction, weighting parameters (weights and offsets) can be signaled for each reference picture included in each of the reference picture lists L0 and L1. Next, when motion compensation is performed, the weight(s) and offset(s) can be applied to the corresponding reference image(s). Weighted prediction and BCW can be used for different types of images. To avoid interaction between weighted prediction and BCW, for the CU using weighted prediction, the BCW weight index can be not signaled. In this case, the weight can be inferred as 4. That is, equal weight can be applied.

[0224] For the CU to which the merge mode is applied, the weight index can be inferred from the neighboring blocks based on the merge candidate index. This can be applied to both the normal merge mode and the inherited affine merge mode.

[0225] In the case of the combined affine merge mode, the affine motion information can be constructed based on the motion information of up to three blocks. In this case, the following process can be performed to derive the BCW weight index for the CU using the combined affine merge mode.

[0226] (1) First, the range of the BCW weight index {0, 1, 2, 3, 4} can be divided into three groups {0}, {1, 2, 3} and {4}. If the BCW weight index of all CPs is derived from the same group, the BCW weight index can be derived by the following step (2). Otherwise, the BCW weight index can be set to 2.

[0227] (2) If at least two CPs have the same BCW weight index, the same BCW weight index can be assigned as the weight index of the combined affine merge candidate. Otherwise, the weight index of the combined affine merge candidate can be set to 2.

[0228] The invention according to the present disclosure described below relates to weighted average-based bidirectional prediction, and specifically includes a method for deriving a BCW weight index when constructing a temporal candidate for the merge mode or a temporal candidate for the sub-block merge mode. Also, according to the present disclosure, a method for deriving a BCW weight index when deriving a pair-wise (merge) candidate is provided. Hereinafter, the BCW weight index may be simply referred to as the weight index. Also, according to the present disclosure, various embodiments for deriving the weight index of the combined affine merge candidate are provided. Each of the various embodiments included in the present disclosure may be used alone, or two or more embodiments may be combined and used.

[0229] According to an embodiment of the present disclosure, when deriving a merge candidate for the merge mode in sub-block units, the coding efficiency can be improved by efficiently deriving the weight index for the combined affine merge candidate.

[0230] As described with reference to FIG. 15, a representative motion vector CPMVk (k is an integer from 1 to 4) for the k-th CP can be derived. That is, CPMV1 is a motion vector representing the first CP (the upper left CP, CP0), CPMV2 is a motion vector representing the second CP (the upper right CP, CP1), CPMV3 is a motion vector representing the third CP (the lower left CP, CP2), and CPMV4 can be a motion vector representing the fourth CP (the lower right CP, RB or CP3). Combinations of CPs for inducing combined affine merge candidates include {CP0, CP1, CP2}, {CP0, CP1, CP3}, {CP0, CP2, CP3}, {CP1, CP2, CP3}, {CP0, CP1}, and {CP0, CP2}, and combined affine merge candidates can be induced in the order of the said combinations.

[0231] According to an example of this embodiment, the weight index for a combined affine merge candidate can be derived to the weight index of the block used to induce a predetermined CP among the CPs within the combination. For example, when the predetermined CP is the first CP or CP0, the weight index for the combined affine merge candidate can be derived to the weight index of the block used to induce CPMV1 among the candidate blocks of the first CP or CP0. For example, in FIG. 15, when it is determined that the motion vector of B2 among B2, B3, and A2 is CPMV1, the weight index for the combined affine merge candidate can be derived to the weight index of B2. However, this is only an example, and the predetermined CP may be the second CP or the third CP within the combination, or may be one CP among CP0, CP1, CP2, and RB. For example, when the predetermined CP is CP1, the weight index for the combined affine merge candidate can be derived to the weight index of the block used to induce CPMV2. For example, in FIG. 15, when it is determined that the motion vector of B0 among B1 and B0 is CPMV2, the weight index for the combined affine merge candidate can be derived to the weight index of B0.

[0232] According to another example of this embodiment, among the weight indexes of each CP, a weight index that is not the default index can be used as the weight index for the combined affine merge candidate. For example, if the combination of CPs for inducing the combined affine merge candidate is {CP0, CP1, CP2}, the weight indexes of CP0 and CP2 are the default indexes, and the weight index of CP1 is not the default index, the weight index of the combined affine merge candidate can be induced to the weight index of CP1.

[0233] According to another example of this embodiment, based on the occurrence frequency of the weight index, the weight index for the combined affine merge candidate can be determined. For example, among the weight indexes of the candidate blocks of a specific CP (for example, CP0), the weight index with a high occurrence frequency can be utilized to induce the weight index for the combined affine merge candidate. Alternatively, among the weight indexes of the CPs constituting the combined affine merge candidate, the weight index with a high occurrence frequency can also be utilized. For example, in the case of the combination of {CP0, CP2, CP3(RB)}, when the weight indexes of CP2 and RB are the same, the weight index for the combined affine merge candidate can be induced to the weight indexes of CP2 and RB with the highest occurrence frequency. At this time, the weight index of RB can be induced based on the method for inducing the weight index for the time candidate according to the present disclosure.

[0234] Hereinafter, with reference to FIGS. 19 to 22, a method for inducing a combined affine candidate according to another embodiment of the present disclosure will be described in detail.

[0235] The method for inducing combined affine merge candidates according to this embodiment can be composed of a step of inducing information for each CP (CP0 to CP3) of the current block, and a step of inducing combined affine merge candidates for each combination using the information of each induced CP.

[0236] FIG. 19 is a diagram for explaining a method of inducing combined affine merge candidates according to another embodiment of the present disclosure. The induction method in FIG. 19 includes a method of inducing a weight index for combined affine merge candidates.

[0237] As shown in FIG. 19, an image encoding device or an image decoding device can first induce information for each CP of the current block (S1910) in order to induce combined affine merge candidates. Each CP of the current block can include the above-described CP0, CP1, CP2, and CP3 (RB). The information for CPn (n is an integer from 0 to 3) can include at least one of a reference picture index (refIdxLXCorner[n]), a prediction direction information (predFlagLXCorner[n]), a motion vector (cpMvLXCorner[n]), the availability of CPn (availableFlagCorn), and a weight index of CPn (bcwIdxCorner[n]).

[0238] According to this embodiment, the weight index for the combined affinity merge candidate can be derived to the weight index of the first CP in each combination. As described above, the combination for deriving the combined affinity merge candidate can be any one of {CP0, CP1, CP2}, {CP0, CP1, CP3}, {CP0, CP2, CP3}, {CP1, CP2, CP3}, {CP0, CP1}, {CP0, CP2}. As described above, the first CP in each combination is CP0 or CP1. Therefore, among the information for the above-mentioned CPn, the weight index of CPn can be derived only for CP0 and CP1. That is, for CP2 and CP3, the process of deriving the weight index of CPn can be omitted. Because even if the weight index of CP2 or CP3 is derived, these weight indexes are not utilized as the weight index for the combined affinity merge candidate.

[0239] When information for each CP is derived in step S1910, the image encoding apparatus or the image decoding apparatus can derive combination affine merge candidates based on the information for each CP (S1920). Step S1920 can be performed for each of the combinations for deriving combination affine merge candidates. At this time, the image encoding apparatus or the image decoding apparatus can perform step S1920 based on information indicating whether a 6-parameter model can be used. For example, when the 6-parameter model can be used, step S1920 can be performed for all of the combinations including the three CPs and the combinations including the two CPs. Otherwise, when the 6-parameter model cannot be used, step S1920 can be performed only for the combinations including the two CPs. Since the 6-parameter model is a model that refers to three CPs, when the 6-parameter model cannot be used, it is not necessary to configure combination affine merge candidates for the combinations including the three CPs. Information indicating whether the 6-parameter model can be used can be signaled via a bitstream. For example, it can be signaled included in a sequence parameter set which is a higher level of a block.

[0240] FIG. 20 is a flowchart for explaining a method of deriving information for the current block's CP according to the embodiment of FIG. 19.

[0241] The current block's CP can include the above-described CP0, CP1, CP2, and CP3 (RB). The method of FIG. 20 can be executed for each of the current block's CPs. The execution order can be in the order of CP0, CP1, CP2, and CP3. However, it is not limited thereto, and it may be performed in a different order, or may be performed simultaneously for some or all of the CPs.

[0242] First, in step S2010, candidate blocks for the current CP can be identified. The candidate blocks for each CP are as described with reference to FIG. 15. For example, the candidate blocks for CP0 are B2, B3, and A2, the candidate blocks for CP1 are B1 and B0, the candidate blocks for CP2 are A1 and A0, and the candidate block for CP3 can be T. When the method of FIG. 20 is performed for each CP, the candidate blocks can be checked in the above-described order (scanning order). Therefore, when there are multiple candidate blocks, the first candidate block can be identified first. For example, B2, which is the first candidate block for CP0, can be identified first.

[0243] When the candidate blocks for the current CP are identified, in step S2020, it can be checked whether the identified candidate blocks are available. Whether a candidate block is available or not can be determined based on whether the candidate block exists within the current picture, whether the candidate block and the current block are within the same slice or the same tile, whether the prediction mode of the candidate block is the same as the prediction mode of the current block, etc. For example, if the candidate block exists outside the current picture, if the candidate block and the current block exist in different slices or different tiles, or if the prediction mode of the candidate block is different from the prediction mode of the current block, it can be determined that the candidate block is not available.

[0244] If the identified candidate block is not available (S2020-No), it can be determined whether there is a next candidate block for the current CP (S2030). For example, when the current CP is CP0 and the identified candidate block is B2, since the next candidate block B3 exists in the above-described order, it can be determined in step S2030 that there is a next candidate block. For example, when the current CP is CP0 and the identified candidate block is A2, since there is no next candidate block in the above-described order, it can be determined in step S2030 that there is no next candidate block.

[0245] If there is a next candidate block for the current CP (S2030 - Yes), steps S2010 to S2020 can be performed on the next candidate block. If there is no next candidate block for the current CP (S2030 - No), the availability (availableFlagCorner) of the current CP can be set to "unavailable" (S2060).

[0246] If the identified candidate block is available (S2020 - Yes), information for the current CP can be derived based on the information of the available candidate block (S2040). For example, the reference picture index, prediction direction information, motion vector, and weight index of the available candidate block can be utilized as information for the current CP. At this time, as described above, only when the current CP is CP0 or CP1, the weight index of the available candidate block can be utilized as the weight index of the current CP. Also, if the identified candidate block is available, the availability (availableFlagCorner) of the current CP can be set to "available" (S2050).

[0247] When information for each CP of the current block is derived, combined affine merge candidates can be derived based on this information.

[0248] FIG. 21 is a diagram showing a method of deriving combined affine merge candidates based on information for each CP according to the embodiment of FIG. 19.

[0249] The method of FIG. 21 can be executed for each of the combinations {CP0, CP1, CP2}, {CP0, CP1, CP3}, {CP0, CP2, CP3}, {CP1, CP2, CP3}, {CP0, CP1}, {CP0, CP2} for deriving combined affine merge candidates. The execution order can be, for example, the order listed above. Further, as described above, depending on the availability of the 6 - parameter model, the method of FIG. 21 can also be performed only for some of the combinations.

[0250] First, in step S2110, the CPs within the current combination for inducing combination affine merge candidates can be identified. Then, in step S2120, the availability of the CPs within the current combination can be determined. Step S2120 can be performed based on the availability information (availableFlagCorner) for each CP within the combination. For example, when the current combination is {CP0, CP1, CP2}, step S2120 can be performed based on availableFlagCorner[0], availableFlagCorner[1], and availableFlagCorner[2]. If all the CPs within the current combination are available (S2120-Yes), steps S2130 to S2150 can be performed separately for the prediction directions (L0 direction and L1 direction). Otherwise (S2120-No), since combination affine merge candidates cannot be induced for the current combination, the method of FIG. 21 is terminated.

[0251] If all the CPs within the combination are available (S2120-Yes), in step S2130, the availability (availableFlagL0 and availableFlagL1) for each of the L0 prediction direction and the L1 prediction direction can be induced. For example, when the current combination is {CP0, CP1, CP2}, for the prediction direction LX (X is 0 or 1), if the prediction direction information (predFlagLXCorner[n]) of CP0, CP1, and CP2 are all 1 and the reference picture indexes (refIdxLXCorner[n]) of CP0, CP1, and CP2 are all the same, the availability for the prediction direction LX is induced as "available", and otherwise, it can be induced as "unavailable".

[0252] Thereafter, in step S2140, it is determined whether the prediction direction LX can be used. If it is "usable" (S2140-Yes), the process can proceed to step S2150. In step S2150, combination affine merge candidates can be induced based on the information of CP for the prediction direction LX. For example, when the input combination is {CP0, CP1, CP2}, for the usable prediction direction LX (X is 0 or 1), the reference picture index (refIdxLXCorner[0]) of CP0, the motion vector (cpMvLXCorner[0]) of CP0, the motion vector (cpMvLXCorner[1]) of CP1, and the motion vector (cpMvLXCorner[2]) of CP2 can be respectively assigned to the reference picture index (refIdxLXConst1), CPMV1, CPMV2, and CPMV3 of the combination affine merge candidate. If the usability for the prediction direction LX is "unusable" (S2140-No), step S2150 is skipped and the process can proceed to step S2160.

[0253] Thereafter, in step S2160, the weight index of the combination affine merge candidate can be induced. Step S2160 can be performed based on the usability for the prediction direction LX. For example, when both the bidirectional of L0 and L1 are usable, the weight index of the combination affine merge candidate can be induced to the weight index of the first CP in the combination. For example, when the current combination is {CP0, CP1, CP2}, the weight index of CP0 can be utilized as the weight index of the combination affine merge candidate. If L0 or L1 is not usable, the weight index of the combination affine merge candidate can be induced to a predetermined index. The predetermined index can be a default index, and for example, it may be an index indicating equal weights.

[0254] Thereafter, in step S2170, it is possible to induce the availability and / or movement model for the current combination. Step S2170 can be performed based on the availability for the prediction direction LX and / or the number of CPs within the combination. For example, if L0 or L1 is available, the availability for the current combination can be induced to be "available". At this time, if the number of CPs within the current combination is three, the movement model of the current combination can be induced to be a 6-parameter affine model. If the number of CPs within the current combination is two, the movement model of the current combination can be induced to be a 4-parameter affine model.

[0255] If not (when both L0 and L1 are unavailable), the availability for the current combination can be induced to be "unavailable". Also, the movement model of the current combination can be induced to be a conventional Translational movement model.

[0256] FIG. 22 is a diagram exemplarily showing a method for inducing a weight index of combination affine merge candidates according to the embodiment of FIG. 21.

[0257] As described above, the weight index of combination affine merge candidates can be induced based on the availability for the prediction direction LX. For this purpose, the availability information for each prediction direction induced in step S2130 can be input (S2210).

[0258] Thereafter, in step S2220, it can be determined whether both the bidirections of L0 and L1 are available. If both the bidirections of L0 and L1 are available (S2220-Yes), the weight index of the combined affine merge candidate can be derived to the weight index of the first CP in the combination (S2230). For example, when the input combination is {CP0, CP1, CP2}, the weight index of CP0 can be used as the weight index of the combined affine merge candidate. For example, when the input combination is {CP1, CP2, CP3}, the weight index of CP1 can be used as the weight index of the combined affine merge candidate.

[0259] If L0 or L1 is not available (S2220-No), the weight index of the combined affine merge candidate can be derived to a predetermined index as described above (S2240). The predetermined index can be a default index, for example, an index indicating equal weights.

[0260] According to the embodiment described with reference to FIGS. 19 to 22, based on whether both the bidirections of L0 and L1 are available, the weight index of the first CP in the combination or the default index is set as the weight index of the combined affine merge candidate. According to another example of this embodiment, regardless of the availability of L0 and L1, the weight index of the first CP in the combination can be set as the weight index of the combined affine merge candidate. Alternatively, the weight index of the CP at a predetermined position in the combination can also be used. According to this, the process of determining the availability of L0 and L1 can be omitted, so that the computational complexity can be reduced and a rapid process can be expected.

[0261] Hereinafter, with reference to FIG. 23, a method for deriving a weight index for a combined affine merge candidate according to another embodiment of the present disclosure will be described.

[0262] According to this embodiment, the weight index for the combined affine merge candidate can be derived based on the weight index and / or weight index group of each CP. The weight W can be selected from five predefined weight sets (e.g., {-2, 3, 4, 5, 10}) based on the weight index of each CP. At this time, the weight index can have values from 0 to 4 and can be classified into three groups. For example, the weight index can be classified into three groups: {0}, {1, 2, 3}, and {4}. At this time, the weight index group indicating the group to which each weight index belongs can have values from 0 to 2. By classifying the weight index into three groups, the pairs of weights indicated by each index can also be classified into three groups. For example, the pairs of weights can be classified into three groups: {(-1 / 4, 5 / 4)}, {(1 / 4, 3 / 4), (2 / 4, 2 / 4), (3 / 4, 1 / 4)}, and {(5 / 4, -1 / 4)}.

[0263] FIG. 23 is a flowchart for explaining an example of a method for deriving a weight index for a combined affine merge candidate according to the present disclosure.

[0264] In FIG. 23, bcwIdxCorner0, bcwIdxCorner1, and bcwIdxCorner2 respectively indicate the weight index of the first CP in the combination, the weight index of the second CP in the combination, and the weight index of the third CP in the combination. Also, bcwIdxGroup0, bcwIdxGroup1, and bcwIdxGroup2 respectively indicate the weight index group of the first CP in the combination, the weight index group of the second CP in the combination, and the weight index group of the third CP in the combination. Also, bcwIdxConst indicates the weight index of the combined affine merge candidate.

[0265] As shown in FIG. 23, in step S2310, it is possible to check whether bcwIdxCorner0 and bcwIdxCorner1 are the same, and whether bcwIdxGroup0 and bcwIdxGroup2 are the same. If both are the same (S2310-Yes), bcwIdxConst can be derived from bcwIdxCorner0 (S2320).

[0266] If not (S2310-No), in step S2330, it is possible to check whether bcwIdxCorner0 and bcwIdxCorner2 are the same, and whether bcwIdxGroup0 and bcwIdxGroup1 are the same. If both are the same (S2330-Yes), bcwIdxConst can be derived from bcwIdxCorner0 (S2320).

[0267] If not (S2330-No), in step S2340, it is possible to check whether bcwIdxCorner1 and bcwIdxCorner2 are the same, and whether bcwIdxGroup1 and bcwIdxGroup0 are the same. If both are the same (S2340-Yes), bcwIdxConst can be derived from bcwIdxCorner2 (S2350).

[0268] If not (S2340-No), bcwIdxConst can be derived to the default index (S2360). The default index may be, for example, an index indicating equal weights.

[0269] According to the example shown in FIG. 23, when three CPs are used, the weight index for the combined affine merge candidates can be derived by a maximum of six comparison operations.

[0270] The method shown in FIG. 23 can be simplified as follows. As described above, the weight index for the time candidates is induced to the default index, and the time candidates are included as the last CP of each combination. Therefore, in the method shown in FIG. 23, the comparison for bcwIdxCorner2 can be omitted.

[0271] In this case, step S2330 can be performed by checking whether bcwIdxGroup0 and bcwIdxGroup1 are the same. Also, step S2340 can be performed by checking whether bcwIdxGroup0 and bcwIdxGroup1 are the same. In this case, since steps S2330 and S2340 check substantially the same conditions, they can be merged into one step.

[0272] FIG. 24 is a flowchart showing a method in which the comparison for bcwIdxCorner2 is omitted according to the present disclosure.

[0273] As shown in FIG. 24, in step S2410, it is possible to check whether bcwIdxCorner0 and bcwIdxCorner1 are the same, and whether bcwIdxGroup0 and bcwIdxGroup2 are the same. If both are the same (S2410-Yes), bcwIdxConst can be induced to bcwIdxCorner0 (S2420).

[0274] If not (S2410-No), it is possible to check whether bcwIdxGroup0 and bcwIdxGroup1 are the same (S2430). If they are the same (S2430-Yes), bcwIdxConst can be induced to bcwIdxCorner0 (S2420).

[0275] If not (S2430-No), bcwIdxConst can be induced to the default index (S2440).

[0276] The method shown in FIG. 23 can be simplified as follows. For example, three weights such as non-low-delay picture can be used and the three weights can belong to one group. In this case, in the method shown in FIG. 23, the comparison for the weight index group can be omitted.

[0277] FIG. 25 is a flowchart showing a method of omitting the comparison for the weight index group according to the present disclosure.

[0278] As shown in FIG. 25, in step S2510, it can be checked whether bcwIdxCorner0 and bcwIdxCorner1 are the same. If they are the same (S2510-Yes), bcwIdxConst can be derived to bcwIdxCorner0 (S2520).

[0279] If not (S2510-No), it can be checked whether bcwIdxCorner0 and bcwIdxCorner2 are the same (S2530). If they are the same (S2530-Yes), bcwIdxConst can be derived to bcwIdxCorner0 (S2520).

[0280] If not (S2530-No), it can be checked whether bcwIdxCorner1 and bcwIdxCorner2 are the same (S2540). If they are the same (S2540-Yes), bcwIdxConst can be derived to bcwIdxCorner2 (S2550).

[0281] If not (S2540-No), bcwIdxConst can be derived to the default index (S2560).

[0282] The method shown in Fig. 23 can be simplified as follows. For example, in the method shown in Fig. 23, all comparisons for bcwIdxCorner2 and for the weight index group can be omitted.

[0283] Fig. 26 is a flowchart showing a method in which the comparison for bcwIdxCorner2 and the comparison for the weight index group are omitted according to the present disclosure.

[0284] As shown in Fig. 26, in step S2610, it is checked whether bcwIdxCorner0 and bcwIdxCorner1 are the same. If they are the same (S2610 - Yes), bcwIdxConst can be derived to bcwIdxCorner0 (S2620). Otherwise (S2610 - No), bcwIdxConst can be derived to the default index (S2630).

[0285] According to another example of this embodiment, bcwIdxConst can also be derived to bcwIdxCorner0 without performing any comparison.

[0286] According to another embodiment of the present disclosure, when constructing motion vector candidates for the merge mode in sub-block units, a method for deriving a weight index when a representative prediction vector candidate uses bidirectional prediction can be provided. As described with reference to FIG. 18, sub-block-based TMVP can derive a col block corresponding to the current block based on a motion shift. At this time, as shown in FIG. 18, the motion shift can be derived from the left adjacent block A1 spatially adjacent to the current block. That is, there is a high probability that the weight index of the left block can be trusted. Therefore, in consideration of this, the weight index of the left block can be used as the weight index of the current block. That is, when the candidate derived by ATMVP uses bidirectional prediction, the weight index of the left block can be used as the weight index for the merge mode of the sub-block. Alternatively, when the motion shift is derived from a peripheral block other than the A1 block, the weight index of the peripheral block can be used as the weight index of the current block. According to this embodiment, when constructing motion vector candidates for the merge mode in sub-block units, by efficiently constructing the weight index for the representative prediction vector candidate, the complexity can be increased without increasing, and the coding efficiency can be improved.

[0287] According to another embodiment of the present disclosure, when a temporal merge candidate uses bidirectional prediction in deriving a merge candidate for the merge mode, the coding efficiency can be improved by deriving the weight index of the temporal candidate.

[0288] According to an example of this embodiment, the weight index of the temporal merge candidate can always be derived to a default index (for example, 0). In the present disclosure, the default index can be an index indicating that the weights for each prediction direction (that is, the L0 prediction direction and the L1 prediction direction in bi-prediction) are the same (equal weights).

[0289] According to another example of this embodiment, the weight index of the temporal merge candidate can be derived from the weight index of the collocated block.

[0290] According to another embodiment of the present disclosure, when deriving merge candidates for the merge mode in sub-block units, when the temporal candidate uses bidirectional prediction, by efficiently deriving the weight index of the temporal candidate, the coding efficiency can be improved.

[0291] According to an example of this embodiment, the weight index of the temporal candidate can always be derived to a default index (for example, 0). In this case, when a temporal candidate is selected from the sub-block merge candidate list based on the merge index, the weight index of the current block can also be set to the default index.

[0292] According to another example of this embodiment, the weight index of the temporal candidate can be derived from the weight index of the center block. The center block may be a block in the col picture including coordinates corresponding to the center position of the current block. The coordinates corresponding to the center position can be derived based on the upper left coordinates (x, y) of the current block and the width and height of the current block. For example, the coordinates corresponding to the center position may be (x + width / 2, y + height / 2).

[0293] According to another example of this embodiment, the weight index of the temporal candidate can be derived from the weight index of the col sub-block corresponding to the current sub-block. If the col sub-block is not available or the weight index of the col sub-block is not available, the weight index of the temporal candidate of the current sub-block can be derived from the weight index of the center block.

[0294] According to another embodiment of the present disclosure, when inducing merge candidates for the merge mode, the encoding efficiency can be improved by efficiently inducing a weight index for pairwise candidates. As described above, pairwise candidates can be induced based on predefined candidate pairs selected from the candidates included in the merge candidate list. At this time, the candidate pairs for inducing pairwise candidates can be represented by cand0 and cand1.

[0295] According to an example of this embodiment, when a pairwise candidate uses bidirectional prediction, the weight index of the pairwise candidate can be induced to the weight index of cand0. Or, the weight index of the pairwise candidate can also be induced to the weight index of cand0, and among the weight index of cand0 and the weight index of cand1, it can also be induced to a weight index that is not the default index (the index indicating a 1:1 weight).

[0296] According to another example of this embodiment, the weight index of the pairwise candidate can be induced by at least one of the following four methods.

[0297] - The weight index of cand0

[0298] - The weight index of the candidate predicted bidirectionally among cand0 and cand1

[0299] - When cand0 and cand1 have the same weight index, it can be set to that weight index, and otherwise, it can be set to the default index.

[0300] - When cand0 and cand1 have the same weight index, it can be set to that weight index, and otherwise, it can be set to a weight index that is not the default index among the weight index of cand0 and the weight index of cand1.

[0301] According to another example of this embodiment, consistency with the induction method for combined affinity candidates can be considered. This is because pairwise candidates and combined affinity candidates have similar characteristics in that a plurality of candidates are combined to generate. That is, when the weight index of cand0 is bcwIdx0 and the weight index of cand1 is bcwIdx1, at least one of the following two methods can be applied to induce the weight index of the pairwise candidate.

[0302] - Based on whether bcwIdx0 and bcwIdx1 are the same, the weight index can be set. First, it can be determined whether bcwIdx0 and bcwIdx1 are the same. If both are the same, the weight index of the pairwise candidate can be induced to bcwIdx0. If both are not the same, the weight index of the pairwise candidate can be induced to the default index.

[0303] - Simply, the weight index of the pairwise candidate can be set to the weight index of the first candidate (e.g., bcwIdx0).

[0304] The exemplary method of the present disclosure is represented in a series of operations for clarity of explanation, but this is not for limiting the order in which the steps are performed. If necessary, each step can also be performed simultaneously or in a different order. To implement the method according to the present disclosure, it can further include other steps in the exemplified steps, or include the remaining steps except for some steps, or include additional other steps except for some steps.

[0305] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of checking the execution conditions and status of the operation (step). For example, when it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation of checking whether the predetermined condition is satisfied.

[0306] The various embodiments of the present disclosure do not list all possible combinations, but are for explaining representative aspects of the present disclosure. The matters described in the various embodiments may be applied independently or in combination of two or more.

[0307] Also, the various embodiments of the present disclosure can be realized by hardware, firmware, software, or a combination thereof. In the case of realization by hardware, it can be realized by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.

[0308] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, an image phone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, as an over-the-top video (OTT) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0309] FIG. 27 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.

[0310] As shown in FIG. 27, the content streaming system to which the embodiments of the present disclosure are applied can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0311] The encoding server compresses the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream, and plays a role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a video camera directly generates a bitstream, the encoding server can be omitted.

[0312] The bitstream can be generated by an image encoding method and / or an image encoding apparatus to which the embodiments of the present disclosure are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0313] The streaming server transmits multimedia data to a user device based on a user request via a Web server, and the Web server can serve as a medium for informing the user of available services. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server can control commands / responses between each device in the content streaming system.

[0314] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0315] Examples of the user device may include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices, for example, smartwatches, smart glasses, head-mounted displays (HMDs), digital TVs, desktop computers, digital signage, and the like.

[0316] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received from each server can be distributedly processed.

[0317] The scope of the present disclosure includes software or machine-executable commands (for example, operating systems, applications, firmware, programs, etc.) that enable operations according to the methods of various embodiments to be executed on a device or computer, and non-transitory computer-readable media on which such software or commands are stored and can be executed on a device or computer.

Industrial Applicability

[0318] Examples according to the present disclosure can be used for encoding / decoding images.

Claims

1. An image decoding method performed by an image decoding device, comprising: The image decoding method includes: determining whether an inter prediction mode of a current block is a sub-block merging mode based on a first flag; constructing a sub-block merging candidate list for the current block based on the inter prediction mode of the current block being the sub-block merging mode; selecting a sub-block merging candidate from the sub-block merging candidate list; deriving motion information of the current block based on motion information of the selected sub-block merging candidate; generating a prediction block of the current block based on the motion information of the current block; deriving a residual block of the current block; reconstructing the current block based on the predicted block of the current block and the residual block of the current block; The step of constructing the sub-block merging candidate list includes a step of deriving combined sub-block merging candidates, and the step of deriving the combined sub-block merging candidates includes a step of deriving weight indexes for bi-prediction of the combined sub-block merging candidates; The step of deriving the combined sub-block merging candidate is performed based on motion information of each candidate control point (CP) included in a combination of predefined candidate control points among a plurality of candidate control points for the current block; The weight index of the combined sub-block merging candidate is derived as a weight index for bi-prediction of a first candidate CP among the candidate CPs included in the combination without performing a comparison between the weight indexes of the candidate CPs; The image decoding method, wherein the first candidate CP among the candidate CPs included in the combination is a top left candidate CP or a top right candidate CP with respect to the current block.

2. the motion information of the candidate CP is derived based on motion information of a candidate block relative to the candidate CP; The image decoding method according to claim 1 , wherein the candidate block is an available candidate block among at least one candidate block for the candidate CP.

3. the motion information of the candidate CP includes a weight index for bi-prediction; The image decoding method of claim 1 , wherein the weight index of the candidate CP is derived based on whether the candidate CP is the top-left candidate CP or the top-right candidate CP for the current block.

4. the motion information of the candidate CP includes a weight index for bi-prediction; The image decoding method of claim 1 , wherein the weight index of the candidate CP is not derived based on the candidate CP being a bottom-left candidate CP or a bottom-right candidate CP for the current block.

5. The image decoding method according to claim 2 , wherein the candidate CP is determined to be unavailable based on the absence of an available candidate block among at least one candidate block for the candidate CP.

6. The image decoding method of claim 4 , wherein the step of deriving the combined sub-block merging candidates is performed based on the availability of all candidate CPs included in the combination of the pre-defined candidate CPs.

7. the weight index for the combined sub-block merging candidate is derived based on whether a prediction direction for the combination is available; The image decoding method according to claim 1 , wherein whether the prediction direction for the combination is available is derived based on motion information of the candidate CPs included in the combination.

8. 8. The image decoding method of claim 7, wherein the weight index of the combined sub-block merging candidate is derived as the first candidate CP among the candidate CPs included in the combination based on whether the prediction direction for the combination is available for both L0 direction and L1 direction.

9. 8. The image decoding method of claim 7, wherein the weight index of the combined sub-block merging candidate is derived as a predetermined weight index based on whether the prediction direction for the combination is unavailable for at least one of an L0 direction and an L1 direction.

10. An image coding method performed by an image coding device, the image coding method comprising: generating a prediction block for the current block based on motion information of the current block; deriving a residual block of the current block based on the predicted block; reconstructing the current block based on the predicted block of the current block and the residual block of the current block; encoding a first flag indicating whether an inter prediction mode of the current block is a sub-block merge mode; encoding the current block based on the predicted block; encoding motion information of the current block; The step of encoding the motion information of the current block includes: constructing a sub-block merging candidate list for the current block based on the inter prediction mode of the current block being the sub-block merging mode; encoding the motion information of the current block based on the sub-block merging candidate list; The step of constructing the sub-block merging candidate list includes a step of deriving combined sub-block merging candidates, and the step of deriving the combined sub-block merging candidates includes a step of deriving weight indexes for bi-prediction of the combined sub-block merging candidates; The step of deriving the combined sub-block merging candidate is performed based on motion information of each candidate control point (CP) included in a combination of predefined candidate control points among a plurality of candidate control points for the current block; The weight index of the combined sub-block merging candidate is derived as a weight index for bi-prediction of a first candidate CP among the candidate CPs included in the combination without performing a comparison between the weight indexes of the candidate CPs; The image coding method, wherein the first candidate CP among the candidate CPs included in the combination is a top left candidate CP or a top right candidate CP with respect to the current block.

11. A method for transmitting a bitstream generated by an image coding method, comprising the steps of: The image encoding method includes: generating a prediction block for the current block based on motion information of the current block; deriving a residual block of the current block based on the predicted block; encoding a first flag indicating whether an inter prediction mode of the current block is a sub-block merge mode; encoding the current block based on the predicted block; encoding motion information of the current block; The step of encoding the motion information of the current block includes: constructing a sub-block merging candidate list for the current block based on the inter prediction mode of the current block being the sub-block merging mode; encoding the motion information of the current block based on the sub-block merging candidate list; The step of constructing the sub-block merging candidate list includes a step of deriving combined sub-block merging candidates, and the step of deriving the combined sub-block merging candidates includes a step of deriving weight indexes for bi-prediction of the combined sub-block merging candidates; The step of deriving the combined sub-block merging candidate is performed based on motion information of each candidate control point (CP) included in a combination of predefined candidate control points among a plurality of candidate control points for the current block; The weight index of the combined sub-block merging candidate is derived as a weight index for bi-prediction of a first candidate CP among the candidate CPs included in the combination without performing a comparison between the weight indexes of the candidate CPs; The method of claim 1, wherein the first candidate CP among the candidate CPs included in the combination is a top-left candidate CP or a top-right candidate CP for the current block.