Method and apparatus for encoding / decoding image on basis of picture header including information related to co-located picture, and method for transmitting bitstream

JP2025113499A5Active Publication Date: 2025-12-22NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025090754
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-06
Filing Date
2025-05-30
Publication Date
2025-12-22
Estimated Expiration
2040-12-04

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.SOLUTION: An image decoding method according to the present disclosure includes the steps of: deriving a temporal motion vector predictor for the current block on the basis of a co-located picture for the current block; deriving a motion vector of the current block on the basis of the temporal motion vector predictor; and generating a prediction block of the current block on the basis of the motion vector, wherein the co-located picture is determined on the basis of identification information on the co-located picture, which is included in a slice header of the current slice including the current block, and when the slice header does not include the identification information on the co-located picture, the co-located picture may be determined on the basis of identification information on the co-located picture, which is included in a picture header of the current picture including the current block.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method, apparatus, and method for transmitting a bitstream based on a picture header including information about a same-position picture, and more particularly, to an image encoding / decoding method and apparatus that perform inter prediction based on a picture header including identification information of a same-position picture, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure.

Background Art

[0002] Recently, demands for high-resolution and high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, have been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on a picture header including information about a same-position picture.

[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved signaling mechanism efficiency for TMVP.

[0007] Furthermore, an object of the present disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0008] Furthermore, an object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0009] Furthermore, an object of the present disclosure is to provide a recording medium storing a bitstream received by an image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0010] The technical problems to be solved by the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.

Means for Solving the Problems

[0011] An image decoding method according to an aspect of the present disclosure includes: deriving a temporal motion vector predictor for a current block based on a same-position picture for the current block; deriving a motion vector of the current block based on the temporal motion vector predictor; and generating a predicted block of the current block based on the motion vector, where the same-position picture is determined based on identification information of the same-position picture included in a slice header of a current slice including the current block, but when the slice header does not include the identification information of the same-position picture, the same-position picture can be determined based on the identification information of the same-position picture included in a picture header of a current picture including the current block.

[0012] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor. The at least one processor derives a temporal motion vector predictor for a current block based on a same-position picture for the current block, derives a motion vector of the current block based on the temporal motion vector predictor, generates a predicted block of the current block based on the motion vector, and the same-position picture is determined based on identification information of the same-position picture included in a slice header of a current slice including the current block. However, when the slice header does not include the identification information of the same-position picture, the same-position picture can be determined based on the identification information of the same-position picture included in a picture header of a current picture including the current block.

[0013] An image encoding method according to another aspect of the present disclosure includes generating a predicted block of a current block based on a motion vector of the current block, deriving a temporal motion vector predictor for the current block based on a same-position picture for the current block, and encoding the motion vector of the current block based on the temporal motion vector predictor. Identification information of the same-position picture is encoded in a slice header of a current slice including the current block. However, when the identification information of the same-position picture is not encoded in the slice header, the identification information of the same-position picture can be encoded in a picture header of a current picture including the current block.

[0014] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or the image encoding apparatus of the present disclosure.

[0015] The features briefly summarized and described above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present disclosure.

Advantages of the Invention

[0016] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0017] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on a picture header including information regarding a same-position picture.

[0018] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved signaling mechanism efficiency for TMVP.

[0019] Also, according to the present disclosure, it is possible to provide a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0020] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0021] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream received by an image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0022] The effects obtained in the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14a

Figure 14b

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

DETAILED DESCRIPTION OF THE INVENTION

[0024] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.

[0025] When explaining the embodiments of the present disclosure, if it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. And in the drawings, parts not related to the description of the present disclosure are omitted, and the same reference numerals are given to the same parts.

[0026] In the present disclosure, when a certain component is “connected”, “coupled” or “connected” to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Also, when a certain component “includes” or “has” another component, this means that, unless otherwise stated to the contrary, it does not exclude another component and can further include another component.

[0027] In the present disclosure, terms such as "first" and "second" are used only for the purpose of distinguishing one component from another, and do not limit the order or importance among the components unless otherwise specifically mentioned. Therefore, within the scope of the present disclosure, the first component of one embodiment may be referred to as the second component in another embodiment, and similarly, the second component of one embodiment may be referred to as the first component in another embodiment.

[0028] In the present disclosure, components that are distinguished from each other are for clearly explaining their respective features, and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, without further mention, such integrated or distributed embodiments are also included in the scope of the present disclosure.

[0029] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments constituted by a subset of the components described in one embodiment are also included in the scope of the present disclosure. In addition, embodiments that further include other components in the components described in various embodiments are also included in the scope of the present disclosure.

[0030] The present disclosure relates to image encoding and decoding, and the terms used in the present disclosure can have the ordinary meanings in the technical field to which the present disclosure belongs unless newly defined in the present disclosure.

[0031] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period, and a slice / tile is an encoding unit constituting a part of a picture, and one picture can be constituted by one or more slices / tiles. In addition, a slice / tile can include one or more CTUs (coding tree units).

[0032] In the present disclosure, "pixel" or "pel" can mean the smallest unit that constitutes a picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component, or can also indicate only the pixel / pixel value of the chroma component.

[0033] In the present disclosure, "unit" can indicate the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can be used interchangeably with terms such as "sample array", "block", or "area" as the case may be. In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0034] In the present disclosure, "current block" can mean any one of "current coding block", "current coding unit", "block to be coded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transform (inverse transform) / quantization (inverse quantization) is performed, "current block" can mean "current transform block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".

[0035] In the present disclosure, unless explicitly stated as a chroma block, the "current block" can mean a block that includes all luma component blocks and chroma component blocks, or the "luma block of the current block". The luma component block of the current block can be explicitly expressed as including an explicit description of the luma component block, such as "luma block" or "current luma block". Also, the chroma component block of the current block can be explicitly expressed as including an explicit description of the chroma component block, such as "chroma block" or "current chroma block".

[0036] In the present disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".

[0037] In the present disclosure, "or" can be interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" can mean "additionally or alternatively".

[0038] Overview of the video coding system

[0039] FIG. 1 is a diagram showing a video coding system according to the present disclosure.

[0040] A video coding system according to an embodiment can include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit the encoded video and / or image information or data in a file or streaming format to the decoding device 20 via a digital storage medium or a network.

[0041] An encoding device 10 according to an embodiment can include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment can include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 can be referred to as a video / image encoding unit, and the decoding unit 22 can be referred to as a video / image decoding unit. The transmission unit 13 can be included in the encoding unit 12. The reception unit 21 can be included in the decoding unit 22. The rendering unit 23 can also include a display unit, and the display unit can be configured as a separate device or an external component.

[0042] The video source generation unit 11 can obtain a video / image through processes such as capture, synthesis, or generation of the video / image. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate a video / image. For example, a virtual video / image can be generated via a computer or the like, and in this case, the video / image capture process can be replaced by a process in which related data is generated.

[0043] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0044] The transmission unit 13 can transmit the encoded video / image information or data output in bitstream format to the receiving unit 21 of the decoding device 20 via a digital storage medium or a network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark), HDD, SSD, etc. The transmission unit 13 can include elements for generating media files via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or the network and transmit it to the decoding unit 22.

[0045] The decoding unit 22 can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit 12.

[0046] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0047] Overview of the image encoding device

[0048] FIG. 2 is a diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure can be applied.

[0049] As shown in FIG. 2, the image encoding apparatus 100 can include an image dividing unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a "prediction unit". The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.

[0050] All or at least a part of the plurality of components constituting the image encoding device 100 can be realized by one hardware component (for example, an encoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.

[0051] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) in a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units at a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units at a lower depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and / or restoration, which will be described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit can be divided or partitioned from the final coding unit respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0052] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform prediction on a processing target block (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information related to the prediction of the current block and transmit it to the entropy encoding unit 190. The information related to the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0053] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of fineness of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the settings. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0054] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU). The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can configure a motion information candidate list based on the peripheral block, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Based on various prediction modes, inter prediction can be performed. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and the motion vector difference and the indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0055] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit can apply intra prediction or inter prediction for the prediction of the current block, and can also apply intra prediction and inter prediction simultaneously. The prediction method of applying intra prediction and inter prediction simultaneously for the prediction of the current block can be called CIIP (combined inter and intra prediction). In addition, the prediction unit can also perform intra block copy (IBC) for the prediction of the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block within the current picture at a position separated from the current block by a predetermined distance. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed in the same manner as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.

[0056] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0057] The conversion unit 120 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the transform obtained from a graph when representing the relationship information between pixels as a graph. CNT means the transform obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can also be applied to pixel blocks having the same size of a square and can also be applied to blocks of a variable size that are not square.

[0058] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 130 can reorder the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0059] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 190 can encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, transmitted information, and / or syntax elements mentioned in the present disclosure can be encoded via the above-described encoding procedure and included in the bitstream.

[0060] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.

[0061] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.

[0062] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed in the current picture and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0063] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and can save the modified restored picture in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bit stream.

[0064] The modified restored picture transmitted to the memory 170 can be used as a reference picture by the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve the encoding efficiency.

[0065] The DPB in the memory 170 can save the modified restored picture for use as a reference picture by the inter prediction unit 180. The memory 170 can save the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The saved motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 170 can save the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.

[0066] Overview of the image decoding device

[0067] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.

[0068] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in a residual processing unit.

[0069] All or at least a part of a plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB and can be realized by a digital storage medium.

[0070] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 of FIG. 2 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be reproduced via a reproducing apparatus (not shown).

[0071] The image decoding device 200 can receive the signal output from the image encoding device of FIG. 2 in the form of a bitstream. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding device can further use the information regarding the parameter set and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements referred to in the present disclosure can be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the information of the surrounding blocks and the decoded information of the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values that have undergone entropy decoding in the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, of the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives the signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can also be provided as a component of the entropy decoding unit 210.

[0072] On the other hand, the image decoding device according to the present disclosure can be called a video / image / picture decoding device. The image decoding device can also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.

[0073] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed by the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using quantization parameters (for example, quantization step size information) to obtain transform coefficients.

[0074] In the inverse conversion unit 230, the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0075] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0076] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, which is the same as described in the explanation of the prediction unit of the image encoding device 100.

[0077] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can also be similarly applied to the intra prediction unit 265.

[0078] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information related to the prediction can include information indicating the mode (technique) of the inter prediction for the current block.

[0079] The adder 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the restored block. The description of the adder 155 can be similarly applied to the adder 235. The adder 235 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed within the current picture and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0080] The filtering unit 240 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and save the modified restored picture in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.

[0081] The (corrected) restored picture stored in the DPB of the memory 250 can be used as a reference picture in the inter prediction unit 260. The memory 250 can store the motion information of the block from which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 250 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 265.

[0082] In this specification, the embodiments described in the filtering unit 160, the inter prediction unit 180, and the intra prediction unit 185 of the image encoding apparatus 100 can be similarly or correspondingly applied to the filtering unit 240, the inter prediction unit 260, and the intra prediction unit 265 of the image decoding apparatus 200 as well.

[0083] Overview of inter prediction

[0084] The image encoding / decoding apparatus can perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction can mean a prediction technique derived in a way that depends on the data elements of pictures other than the current picture. When inter prediction is applied to the current block, a prediction block for the current block can be induced based on the reference block specified by the motion vector on the reference picture.

[0085] At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be induced based on the correlation of the motion information between the neighboring blocks and the current block, and the motion information can be induced in units of blocks, sub-blocks, or samples. At this time, the motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction type information. Here, the inter prediction type information can mean the directional information of the inter prediction. The inter prediction type information can indicate that the current block is predicted using any one of L0 prediction, L1 prediction, and Bi prediction.

[0086] When inter prediction is applied to the current block, the neighboring blocks of the current block can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. At this time, the reference picture including the reference block for the current block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring block is a collocated reference block (sometimes called colCU, etc.), and the reference picture including the temporal neighboring block may sometimes be called a collocated picture (colPic).

[0087] On the other hand, a motion information candidate list can be constructed based on the neighboring blocks of the current block. At this time, flag or index information indicating which candidate is used can be signaled in order to derive the motion vector and / or reference picture index of the current block.

[0088] The motion information can include L0 motion information and / or L1 motion information based on the inter-prediction type. The motion vector in the L0 direction can be defined as the L0 motion vector or MVL0, and the motion vector in the L1 direction can be defined as the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be defined as L0 prediction, the prediction based on the L1 motion vector can be defined as L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector can be defined as bi-prediction. Here, the L0 motion vector can mean the motion vector related to the reference picture list L0, and the L1 motion vector can mean the motion vector related to the reference picture list L1.

[0089] The reference picture list L0 can include pictures previous to the current picture in the output order as reference pictures, and the reference picture list L1 can include pictures subsequent to the current picture in the output order. At this time, the previous picture can be defined as the forward (reference) picture, and the subsequent picture can be defined as the backward (reference) picture. On the other hand, the reference picture list L0 can further include pictures subsequent to the current picture in the output order. In this case, the previous pictures in the reference picture list L0 can be indexed first, and the subsequent pictures can be indexed next. The reference picture list L1 can further include pictures previous to the current picture in the output order. In this case, the subsequent pictures in the reference picture list L1 can be indexed first, and the previous pictures can be indexed next. Here, the output order can correspond to the POC (picture order count) order.

[0090] FIG. 4 is a flowchart showing an inter-prediction-based video / image encoding method.

[0091] FIG. 5 is a diagram exemplarily showing the configuration of the inter prediction unit 180 according to the present disclosure.

[0092] The encoding method of FIG. 4 can be performed by the image encoding apparatus of FIG. 2. Specifically, step S410 can be performed by the inter prediction unit 180, and step S420 can be performed by the residual processing unit. Specifically, step S420 can be performed by the subtraction unit 115. Step S430 can be performed by the entropy encoding unit 190. The prediction information of step S430 is derived by the inter prediction unit 180, and the residual information of step S430 can be derived by the residual processing unit. The residual information is information regarding the residual sample. The residual information can include information regarding the quantized transform coefficients for the residual sample. As described above, the residual sample is derived as a transform coefficient via the transform unit 120 of the image encoding apparatus, and the transform coefficient can be derived as a quantized transform coefficient via the quantization unit 130. The information regarding the quantized transform coefficients can be encoded by the entropy encoding unit 190 via the residual coding procedure.

[0093] The image encoding device can perform inter prediction on the current block (S410). The image encoding device can derive the inter prediction mode and motion information of the current block and generate a prediction sample of the current block. Here, the inter prediction mode determination, motion information derivation, and prediction sample generation procedures may be performed simultaneously, or any one of the procedures may be performed prior to the other procedures. For example, as shown in FIG. 5, the inter prediction unit 180 of the image encoding device may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 can determine the prediction mode for the current block, the motion information derivation unit 182 can derive the motion information of the current block, and the prediction sample derivation unit 183 can derive the prediction sample of the current block. For example, the inter prediction unit 180 of the image encoding device can search for a block similar to the current block within a certain region (search region) of the reference picture through motion estimation, and derive a reference block whose difference from the current block is the minimum or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The image encoding device can determine the mode to be applied to the current block among various inter prediction modes. The image encoding device can compare the rate-distortion cost (Rate-Distortion (RD) cost) for the various prediction modes and determine the optimal inter prediction mode for the current block. However, the method by which the image encoding device determines the inter prediction mode for the current block is not limited to the above example, and various methods can be used.

[0094] For example, the inter prediction mode for the current block can be determined to be at least one of a merge mode, a merge skip mode, a Motion Vector Prediction (MVP) mode, a Symmetric Motion Vector Difference (SMVD) mode, an affine mode, a subblock-based merge mode, an Adaptive Motion Vector Resolution (AMVR) mode, a History-based Motion Vector Predictor (HMVP) mode, a Pair-wise average merge mode, a Merge mode with Motion Vector Differences (MMVD) mode, a Decoder side Motion Vector Refinement (DMVR) mode, a Combined Inter and Intra Prediction (CIIP) mode, and a Geometric Partitioning (GPM) mode.

[0095] For example, when the skip mode or the merge mode is applied to the current block, the image encoding device can derive merge candidates from the surrounding blocks of the current block and construct a merge candidate list using the derived merge candidates. Further, the image encoding device can derive a reference block among the reference blocks pointed to by the merge candidates included in the merge candidate list, where the difference from the current block is the smallest or below a certain criterion. In this case, the merge candidate related to the derived reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the image decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0096] As another example, when the MVP mode is applied to the current block, the image encoding apparatus can derive motion vector predictor (MVP) candidates from the peripheral blocks of the current block and configure an MVP candidate list using the derived MVP candidates. Further, the image encoding apparatus can use the motion vector of the selected MVP candidate among the MVP candidates included in the MVP candidate list as the MVP of the current block. In this case, for example, the motion vector indicating the reference block derived by the above-described motion estimation can be used as the motion vector of the current block, and among the MVP candidates, the MVP candidate having the motion vector with the smallest difference from the motion vector of the current block can be the selected MVP candidate. An MVD (motion vector difference), which is the difference obtained by subtracting the MVP from the motion vector of the current block, can be derived. In this case, the index information indicating the selected MVP candidate and the information regarding the MVD can be signaled to the image decoding apparatus. Further, when the MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and separately signaled to the image decoding apparatus.

[0097] The image encoding apparatus can derive a residual sample based on the prediction sample (S420). The image encoding apparatus can derive the residual sample by comparing the original sample of the current block with the prediction sample. For example, the residual sample can be derived by subtracting the corresponding prediction sample from the original sample.

[0098] The image encoding device can encode image information including prediction information and residual information (S430). The image encoding device can output the encoded image information in the form of a bitstream. The prediction information is information related to the prediction procedure and can include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information regarding motion information. Among the prediction mode information, the skip flag is information indicating whether the skip mode is applied to the current block, and the merge flag is information indicating whether the merge mode is applied to the current block. Alternatively, the prediction mode information may be information indicating any one of a plurality of prediction modes, such as a mode index. When the skip flag and the merge flag are both 0, it can be determined that the MVP mode is applied to the current block. The information regarding the motion information can include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. Among the candidate selection information, the merge index can be signaled when the merge mode is applied to the current block and can be information for selecting any one of the merge candidates included in the merge candidate list. Among the candidate selection information, the MVP flag or MVP index can be signaled when the MVP mode is applied to the current block and can be information for selecting any one of the MVP candidates included in the MVP candidate list. Specifically, the MVP flag can be signaled using the syntax elements mvp_l0_flag or mvp_l1_flag. Also, the information regarding the motion information can include information regarding the MVD described above and / or reference picture index information. Also, the information regarding the motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information regarding the residual samples.The residual information can include information regarding the quantized transform coefficients for the residual samples.

[0099] The output bitstream can be stored in a (digital) storage medium and transmitted to an image decoding device, or can also be transmitted to the image decoding device via a network.

[0100] On the other hand, as described above, the image encoding device can generate a reconstructed picture (a picture including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the same prediction result as that performed by the image decoding device is derived by the image encoding device, and thereby the coding efficiency can be improved. Therefore, the image encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in a memory and utilize it as a picture for inter prediction. As described above, in-loop filtering procedures and the like can be further applied to the reconstructed picture.

[0101] FIG. 6 is a flowchart showing an inter prediction-based video / image decoding method.

[0102] FIG. 7 is a diagram exemplarily showing the configuration of the inter prediction unit 260 according to the present disclosure.

[0103] The image decoding device can perform operations corresponding to the operations performed by the image encoding device. The image decoding device can perform prediction on the current block based on the received prediction information and derive prediction samples.

[0104] The decoding method of FIG. 6 can be performed by the image decoding apparatus of FIG. 3. Steps S610 to S630 can be performed by the inter prediction unit 260, and the prediction information in step S610 and the residual information in step S640 can be obtained from the bit stream by the entropy decoding unit 210. The residual processing unit of the image decoding apparatus can derive the residual sample for the current block based on the residual information (S640). Specifically, the inverse quantization unit 220 of the residual processing unit performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive the transform coefficients, and the inverse transform unit 230 of the residual processing unit can perform an inverse transform on the transform coefficients to derive the residual sample for the current block. Step S650 can be performed by the addition unit 235 or the restoration unit.

[0105] Specifically, the image decoding apparatus can determine the prediction mode for the current block based on the received prediction information (S610). The image decoding apparatus can determine which inter prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0106] For example, based on the skip flag, it can be determined whether the skip mode is applied to the current block. Also, based on the merge flag, it can be determined whether the merge mode is applied to the current block or the MVP mode is determined. Alternatively, based on the mode index, any one of various inter prediction mode candidates can be selected. The inter prediction mode candidates can include the skip mode, the merge mode, and / or the MVP mode, or can include various inter prediction modes described later.

[0107] The image decoding device can derive the motion information of the current block based on the determined inter prediction mode (S620). For example, when the skip mode or the merge mode is applied to the current block, the image decoding device can construct a merge candidate list described later and select any one of the merge candidates included in the merge candidate list. The selection can be performed based on the candidate selection information (merge index) described above. The motion information of the current block can be derived using the motion information of the selected merge candidate. For example, the motion information of the selected merge candidate can be used as the motion information of the current block.

[0108] As another example, when the MVP mode is applied to the current block, the image decoding device can construct an MVP candidate list and use the motion vector of the MVP candidate selected from among the MVP candidates included in the MVP candidate list as the mvp of the current block. The selection can be performed based on the candidate selection information (mvp flag or mvp index) described above. In this case, based on the information regarding the MVD, the MVD of the current block can be derived, and based on the MVP and the MVD of the current block, the motion vector of the current block can be derived. Also, based on the reference picture index information, the reference picture index of the current block can be derived. The picture pointed to by the reference picture index within the related reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.

[0109] The image decoding apparatus can generate a prediction sample for the current block based on the motion information of the current block (S630). In this case, the reference picture can be derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the samples of the reference block pointed to by the motion vector of the current block on the reference picture. Optionally, a prediction sample filtering procedure can be further performed on all or part of the prediction samples of the current block.

[0110] For example, as shown in FIG. 7, the inter prediction unit 260 of the image decoding apparatus can include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. The inter prediction unit 260 of the image decoding apparatus determines a prediction mode for the current block based on the prediction mode information received from the prediction mode determination unit 261, and derives the motion information (such as a motion vector and / or a reference picture index) of the current block based on the information related to the motion information received from the motion information derivation unit 262, and the prediction sample derivation unit 263 can derive the prediction sample of the current block.

[0111] The image decoding apparatus can generate a residual sample for the current block based on the received residual information (S640). The image decoding apparatus can generate a restored sample for the current block based on the prediction sample and the residual sample, and generate a restored picture based on this (S650). As described above, an in-loop filter procedure and the like can be further applied to the restored picture.

[0112] As described above, the inter prediction procedure can include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction execution (generation of a prediction sample) step based on the derived motion information. The inter prediction procedure can be performed by the image encoding apparatus and the image decoding apparatus as described above.

[0113] Next, the motion information derivation step in the prediction mode will be described in more detail.

[0114] As described above, inter prediction can be performed using the motion information of the current block. The image encoding device can derive the optimal motion information for the current block through a motion estimation procedure. For example, the image encoding device can search for a highly correlated similar reference block within a defined search range in the reference picture in units of fractional pixels using the original block in the original picture for the current block, thereby deriving the motion information. The similarity of the blocks can be calculated based on the SAD (sum of absolute differences) between the current block and the reference block. In this case, the motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information can be signaled to the image decoding device in various ways based on the inter prediction mode.

[0115] When the merge mode is applied to the current block, the motion information of the current block is not directly transmitted, and the motion information of the current block is induced using the motion information of the surrounding blocks. Therefore, by transmitting the flag information indicating that the merge mode has been used and the candidate selection information (e.g., merge index) indicating which surrounding blocks are used as merge candidates, the motion information of the current prediction block can be indicated. In the present disclosure, since the current block is the unit of prediction execution, the current block can be used in the same meaning as the current prediction block, and the surrounding blocks can be used in the same meaning as the surrounding prediction blocks.

[0116] The image encoding device can search for merge candidate blocks used to derive the motion information of the current block for performing the merge mode. For example, up to 5 merge candidate blocks can be used, but it is not limited thereto. The maximum number of the merge candidate blocks can be transmitted from the slice header or the tile group header, but it is not limited thereto. After finding the merge candidate blocks, the image encoding device can generate a merge candidate list, and among these, the merge candidate block with the smallest RD cost can be selected as the final merge candidate block.

[0117] The present disclosure provides various embodiments for the merge candidate blocks constituting the merge candidate list. For example, 5 merge candidate blocks can be used for the merge candidate list. For example, 4 spatial merge candidates and 1 temporal merge candidate can be used.

[0118] FIG. 8 is a diagram illustrating peripheral blocks that can be used as spatial merge candidates.

[0119] FIG. 9 is a diagram schematically showing a method for configuring a merge candidate list according to an example of the present disclosure.

[0120] The image encoding device / image decoding device can insert a spatial merge candidate derived by searching the spatial neighboring blocks of the current block into the merge candidate list (S910). For example, as shown in FIG. 8, the spatial neighboring blocks can include the lower left corner neighboring block A0, the left neighboring block A1, the upper right corner neighboring block B0, the upper neighboring block B1, and the upper left corner neighboring block B2 of the current block. However, this is merely an example, and in addition to the aforementioned spatial neighboring blocks, additional neighboring blocks such as the right neighboring block, the lower neighboring block, and the lower right neighboring block can also be used as the spatial neighboring blocks. The image encoding device / image decoding device can detect available blocks by searching the spatial neighboring blocks based on the priority, and derive the motion information of the detected blocks as the spatial merge candidates. For example, the image encoding device / image decoding device can search the five blocks shown in FIG. 8 in the order of A1, B1, B0, A0, B2, and sequentially index the available candidates to construct the merge candidate list.

[0121] The image encoding device / image decoding device can insert a temporal merge candidate derived by searching for temporal neighboring blocks of the current block into the merge candidate list (S920). The temporal neighboring blocks can be located on a reference picture that is a picture different from the current picture in which the current block is located. The reference picture on which the temporal neighboring blocks are located can be called a collocated picture or a col picture. The temporal neighboring blocks can be searched in the order of the peripheral blocks of the lower right corner and the lower right center block of the collocated block with respect to the current block on the col picture. On the other hand, when motion data compression is applied to reduce the memory load, specific motion information can be stored as representative motion information for each fixed storage unit with respect to the col picture. In this case, it is not necessary to store the motion information for all blocks within the fixed storage unit, and thus the effect of motion data compression can be obtained. In this case, the fixed storage unit can be predetermined, for example, in units of 16×16 samples, or 8×8 samples, or the size information for the fixed storage unit can be signaled from the image encoding device to the image decoding device. When motion data compression is applied, the motion information of the temporal neighboring blocks can be replaced with the representative motion information of the fixed storage unit in which the temporal neighboring blocks are located. That is, in this case, from the perspective of implementation, instead of the prediction block lock located at the coordinates of the temporal neighboring blocks, based on the coordinates (upper left sample position) of the temporal neighboring blocks, after arithmetic right shift by a certain value, the temporal merge candidate can be derived based on the motion information of the prediction block covering the position after arithmetic left shift. For example, if the fixed storage unit is 2 n ×2 nWhen it is in sample units, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb >> n) << n), (yTnb >> n) << n)) can be used for the temporal merge candidate. Specifically, for example, when the fixed storage unit is 16×16 sample units, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb >> 4) << 4), (yTnb >> 4) << 4)) can be used for the temporal merge candidate. Or, for example, when the fixed storage unit is 8×8 sample units, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb >> 3) << 3), (yTnb >> 3) << 3)) can be used for the temporal merge candidate.

[0122] Referring to FIG. 9 again, the image encoding device / image decoding device can check whether the number of current merge candidates is smaller than the number of maximum merge candidates (S930). The number of the maximum merge candidates can be predefined or signaled from the image encoding device to the image decoding device. For example, the image encoding device can generate information regarding the number of the maximum merge candidates, encode it, and transmit it to the image decoding device in the form of a bit stream. When all of the number of the maximum merge candidates are satisfied, the subsequent candidate addition process (S940) can be not performed.

[0123] If the check result in step S930 indicates that the number of current merge candidates is smaller than the number of maximum merge candidates, the image encoding device / image decoding device can induce additional merge candidates based on a predetermined method and then insert them into the merge candidate list (S940). The additional merge candidates can include, for example, at least one of history based merge candidate(s), pair-wise average merge candidate(s), ATMVP, conbined bi-predictive merge candidate(s) (when the slice / tile group type of the current slice / tile group is of type B) or / and zero vector merge candidate(s).

[0124] If the check result in step S930 indicates that the number of current merge candidates is not smaller than the number of maximum merge candidates, the image encoding device / image decoding device can finish configuring the merge candidate list. In this case, the image encoding device can select the optimal merge candidate among the merge candidates that configure the merge candidate list based on the RD cost, and can signal candidate selection information (e.g., merge candidate index, merge index) indicating the selected merge candidate to the image decoding device. The image decoding device can select the optimal merge candidate based on the merge candidate list and the candidate selection information.

[0125] As described above, the motion information of the selected merge candidate can be used as the motion information of the current block, and the predicted sample of the current block can be derived based on the motion information of the current block. The image encoding device can derive the residual sample of the current block based on the predicted sample, and can signal the residual information regarding the residual sample to the image decoding device. As described above, the image decoding device can generate a restored sample based on the residual sample derived based on the residual information and the predicted sample, and can generate a restored picture based on this.

[0126] When a skip mode is applied to the current block, the motion information of the current block can be derived in the same way as when the merge mode was applied previously. However, when the skip mode is applied, the residual signal for the block is omitted. Therefore, the predicted sample can be immediately used as the restored sample. The skip mode can be applied, for example, when the value of cu_skip_flag is 1.

[0127] Hereinafter, a method for deriving a spatial candidate in the case of the merge mode and / or the skip mode will be described. The spatial candidate can indicate the spatial merge candidate described above.

[0128] The induction of spatial candidates can be performed based on spatially adjacent blocks. For example, up to four spatial candidates can be induced from the candidate blocks existing at the positions shown in FIG. 8. The order of inducing spatial candidates can be in the order of A1->B1->B0->A0->B2. However, the order of inducing spatial candidates is not limited to the above order, and for example, it may be in the order of B1->A1->B0->A0->B2. The last position in the order (in the above example, the B2 position) can be considered when at least one of the preceding four positions (in the above example, A1, B1, B0, and A0) is not available. At this time, the fact that the block at a predetermined position is not available can include the case where the block belongs to a different slice or a different tile from the current block, or the case where the block is an intra-predicted block. When spatial candidates are induced from the first position in the order (in the above example, A1 or B1), a redundancy check can be performed on the spatial candidates at the subsequent positions. For example, when the motion information of the subsequent spatial candidates is the same as the motion information of the spatial candidates already included in the merge candidate list, the subsequent spatial candidates can be excluded from the merge candidate list, thereby improving the coding efficiency. The redundancy check performed on the subsequent spatial candidates is not performed on all candidate pairs as much as possible, but only on some candidate pairs, thereby reducing the computational complexity.

[0129] FIG. 10 is a diagram illustrating candidate pairs for the redundancy check performed on spatial candidates.

[0130] In the example shown in FIG. 10, the redundancy check for the spatial candidate at the B0 position can be performed only for the spatial candidate at the A0 position. Also, the redundancy check for the spatial candidate at the B1 position can be performed only for the spatial candidate at the B0 position. Also, the redundancy check for the spatial candidate at the A1 position can be performed only for the spatial candidate at the A0 position. Finally, the redundancy check for the spatial candidate at the B2 position can be performed only for the spatial candidates at the A0 and B0 positions.

[0131] The example shown in FIG. 10 is an example when the order of inducing spatial candidates is in the order of A0 -> B0 -> B1 -> A1 -> B2. However, it is not limited to this, and even if the order of inducing spatial candidates is changed, as in the example shown in FIG. 10, the redundancy check can be performed only for some candidate pairs.

[0132] Hereinafter, in the case of the merge mode and / or the skip mode, a method for inducing time candidates will be described. The time candidates can indicate the time merge candidates described above. Also, the motion vector of the time candidates can also correspond to the time candidates in the MVP mode.

[0133] Only one candidate can be included in the merge candidate list for the time candidates. In the process of inducing the time candidates, the motion vector of the time candidates can be scaled. For example, the scaling can be performed based on the co-located Cu (hereinafter referred to as "col block") belonging to the collocated reference picture (hereinafter referred to as "col picture") at the same position. The reference picture list used for inducing the col block can be explicitly signaled in the slice header.

[0134] FIG. 11 is a diagram for explaining a method of scaling the motion vector of the time candidates.

[0135] In FIG. 11, curr_CU and curr_pic indicate the current block and the current picture, and col_CU and col_pic indicate the col block and the col picture. Also, curr_ref indicates the reference picture of the current block, and col_ref indicates the reference picture of the col block. Also, tb indicates the distance between the reference picture and the current picture of the current block, and td indicates the distance between the reference picture and the col picture of the col block. The tb and td can be represented by values corresponding to the difference in POC (Picture Order Count) between pictures. Scaling of the motion vectors of the temporal candidates can be performed based on tb and td. Also, the reference picture index of the temporal candidates can be set to 0.

[0136] FIG. 12 is a diagram for explaining the position for inducing temporal candidates.

[0137] In FIG. 12, the block with the thick solid line indicates the current block. The temporal candidates can be induced from the blocks in the col picture corresponding to the C0 position (lower right position) or the C1 position (central position) in FIG. 12. First, it is determined whether the C0 position is available. If the C0 position is available, the temporal candidates can be induced based on the C0 position. If the C0 position is not available, the temporal candidates can be induced based on the C1 position. For example, if the block in the col picture at the C0 position is an intra prediction block or exists outside the current CTU row (row), it can be determined that the C0 position is not available.

[0138] As described above, when motion data compression is applied, the motion vectors of the col blocks can be saved for each predetermined unit block. In this case, the C0 position or the C1 position can be modified in order to derive the motion vectors of the blocks covering the C0 position or the C1 position. For example, when the predetermined unit block is an 8×8 block and the C0 position or the C1 position is (xColCi, yColCi), the position for deriving the temporal candidate can be modified to ((xColCi>>3)<<3, (yColCi>>3)<<3).

[0139] Hereinafter, a method for deriving a History-based candidate in the case of the merge mode and / or the skip mode will be described. The History-based candidate can be expressed as a History-based merge candidate.

[0140] The History-based candidate can be added to the merge candidate list after the spatial candidate and the temporal candidate are added to the merge candidate list. For example, the motion information of the previously encoded / decoded block is saved in a table and can be used as the History-based candidate of the current block. The table can save a plurality of History-based candidates during the encoding / decoding process. The table can be initialized when a new CTU row starts. Initializing the table can mean that all the History-based candidates saved in the table are deleted and the table becomes empty. Each time an inter-predicted block exists, the related motion information can be added to the table as the last entry. At this time, the inter-predicted block does not have to be a block predicted based on sub-blocks. The motion information added to the table can be used as a new History-based candidate.

[0141] The history-based candidate table can have a predetermined size. For example, the size can be 5. At this time, the table can store up to 5 history-based candidates. When a new candidate is added to the table, first, a redundancy check is performed to see if the same candidate exists in the table, and a limited FIFO (first-in-first-out) rule can be applied. If the same candidate already exists in the table, the same candidate is deleted from the table, and the positions of all subsequent history-based candidates can be moved forward.

[0142] History-based candidates can be used in the process of constructing the merge candidate list. At this time, the history-based candidates recently included in the table are sequentially checked and can be included in positions after the time candidates in the merge candidate list. When a history-based candidate is included in the merge candidate list, a redundancy check can be performed with the space or time candidates already included in the merge candidate list. If a history-based candidate overlaps with a space or time candidate already included in the merge candidate list, the history-based candidate may not be included in the merge candidate list. The redundancy check can reduce the computational complexity by simplifying it as follows.

[0143] The number of history-based candidates used for generating the merge candidate list can be set to (N <= 4)? M : (8 - N). At this time, N indicates the number of candidates already included in the merge candidate list, and M indicates the number of available history-based candidates stored in the table. That is, when the merge candidate list contains 4 or fewer candidates, the number of history-based candidates used for generating the merge candidate list is M. When the merge candidate list contains N candidates more than 4, the number of history-based candidates used for generating the merge candidate list can be set to (8 - N).

[0144] When the total number of available merge candidates reaches (the maximum allowable number of merge candidates - 1), the construction of the merge candidate list using History-based candidates can be terminated.

[0145] Hereinafter, in the case of the merge mode and / or the skip mode, a method for deriving Pair-wise average candidates will be described. Pair-wise average candidates can be expressed as Pair-wise average merge candidates or Pair-wise candidates.

[0146] Pair-wise average candidates can be generated by obtaining defined candidate pairs from the candidates included in the merge candidate list and averaging them. The defined candidate pairs are {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, and the numbers constituting each candidate pair can be the indices of the merge candidate list. That is, the defined candidate pair (0,1) means a pair of the candidate at index 0 and the candidate at index 1 in the merge candidate list, and the Pair-wise average candidate can be generated by averaging the candidate at index 0 and the candidate at index 1. The derivation of Pair-wise average candidates can be performed in the order of the defined candidate pairs. That is, after deriving the Pair-wise average candidate for the candidate pair (0,1), the process of deriving the Pair-wise average candidate can be performed in the order of the candidate pairs (0,2) and (1,2). The process of deriving the Pair-wise average candidate can be performed until the construction of the merge candidate list is completed. For example, the process of deriving the Pair-wise average candidate can be performed until the number of merge candidates included in the merge candidate list reaches the maximum number of merge candidates.

[0147] Pair-wise average candidates can be calculated individually for each of the reference picture lists. If two motion vectors are available for one reference picture list (L0 list or L1 list), the average of these two motion vectors can be calculated. At this time, even if the two motion vectors point to different reference pictures from each other, the average of the two motion vectors can be performed. If only one motion vector is available for one reference picture list, the available motion vector can be used as the motion vector of the Pair wise average candidate. If not all two motion vectors are available for one reference picture list, the reference picture list can be determined to be invalid.

[0148] Even after the Pair-wise average candidate is included in the merge candidate list, if the configuration of the merge candidate list is not completed, zero vectors can be added to the merge candidate list until the maximum number of merge candidates is reached.

[0149] When the MVP mode is applied to the current block, a motion vector predictor (MVP) candidate list can be generated using the motion vectors of the restored spatial neighboring blocks (e.g., the neighboring blocks shown in FIG. 8) and / or the motion vectors corresponding to the temporal neighboring blocks (or Col blocks). That is, the motion vectors of the restored spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks can be used as candidates for the motion vector predictor of the current block. When dual prediction is applied, an MVP candidate list for L0 motion information derivation and an MVP candidate list for L1 motion information derivation can be generated and used separately. The prediction information (or information related to prediction) for the current block can include candidate selection information (e.g., an MVP flag or an MVP index) indicating the optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the MVP candidate list. At this time, the prediction unit can select the motion vector predictor of the current block from among the motion vector predictor candidates included in the MVP candidate list using the candidate selection information. The prediction unit of the image encoding apparatus can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can encode this and output it in the form of a bitstream. That is, the MVD can be obtained by subtracting the motion vector predictor from the motion vector of the current block. The prediction unit of the image decoding apparatus can obtain the motion vector difference included in the information related to the prediction, and can derive the motion vector of the current block through addition of the motion vector difference and the motion vector predictor. The prediction unit of the image decoding apparatus can obtain or derive a reference picture index indicating a reference picture, etc. from the information related to the prediction.

[0150] FIG. 13 is a diagram schematically showing a method for configuring a motion vector predictor candidate list according to an example of the present disclosure.

[0151] First, the spatial candidate blocks of the current block can be searched, and the available candidate blocks can be inserted into the MVP candidate list (S1310). Then, it is determined whether the number of MVP candidates included in the MVP candidate list is less than two (S1320). If the number is two, the configuration of the MVP candidate list can be completed.

[0152] In step S1320, if the number of available spatial candidate blocks is less than two, the temporal candidate blocks of the current block can be searched, and the available candidate blocks can be inserted into the MVP candidate list (S1330). If the temporal candidate blocks are not available, the zero motion vector can be inserted into the MVP candidate list (S1340), thereby completing the configuration of the MVP candidate list.

[0153] On the other hand, when the MVP mode is applied, the reference picture index can be explicitly signaled. In this case, the picture index (refidxL0) for L0 prediction and the reference picture index (refidxL1) for L1 prediction can be separately signaled. For example, when the MVP mode is applied and bi-prediction is applied, both the information regarding refidxL0 and the information regarding refidxL1 can be signaled.

[0154] As described above, when the MVP mode is applied, the information regarding the MVD derived from the image encoding device can be signaled to the image decoding device. The information regarding the MVD can include, for example, the MVD absolute value and the information indicating the x and y components of the sign. In this case, whether the MVD absolute value is greater than 0 and greater than 1, and the remaining information of the MVD can be signaled step by step. For example, the information indicating whether the MVD absolute value is greater than 1 can be signaled only when the value of the flag information indicating whether the MVD absolute value is greater than 0 is 1.

[0155] Overview of TMVP (temporal motion vector predictor) signaling

[0156] As described above, for the inter prediction of a current block, a temporal motion vector predictor (TMVP) that is used as a temporal merge candidate or a temporal MVP candidate can be derived. The TMVP can be derived based on temporal neighboring blocks within a collocated (reference) picture (colPic) at the same position. Here, the temporal neighboring blocks can include a collocated reference block (colCb) of the current block at the same position.

[0157] Information / syntax elements indicating the collocated reference block (colCb) can be signaled via a high level syntax (HLS) such as a picture header. A TMVP coding tool can be used to encode / decrypt a bitstream. The signaling mechanism for the TMVP is as follows.

[0158] The TMVP can be used for picture coding within a CLVS (coded layer video sequence). Whether the TMVP is available for a picture within the CLVS can be determined based on predetermined flag information (e.g., sps_temporal_mvp_enabled_flag) within an SPS (sequence parameter set). For example, a sps_temporal_mvp_enabled_flag having a first value (e.g., 0) indicates that the TMVP is not available for a picture within the CLVS, and a sps_temporal_mvp_enabled_flag having a second value (e.g., 1) can indicate that the TMVP is available for a picture within the CLVS.

[0159] When TMVP is available for a picture in CLVS, whether TMVP is available for each picture can be determined based on predetermined flag information (e.g., pic_temporal_mvp_enabled_flag) in the picture header. For example, a pic_temporal_mvp_enabled_flag having a first value (e.g., 0) can indicate that TMVP is not available for inter prediction for slices in the picture related to the picture header. In this case, the syntax elements for slices in the picture related to the picture header can be restricted so that TMVP is not used for slice decoding. In contrast, a pic_temporal_mvp_enabled_flag having a second value (e.g., 1) can indicate that TMVP is available for inter prediction for slices in the picture related to the picture header. On the other hand, when the pic_temporal_mvp_enabled_flag is not signaled, the value of the pic_temporal_mvp_enabled_flag can be inferred as the first value (e.g., 0). On the other hand, when there is no reference picture having the same spatial resolution as the current picture in the reference picture buffer (DPB), the value of the pic_temporal_mvp_enabled_flag can be restricted to the first value (e.g., 0).

[0160] When TMVP is available for a predetermined picture, for each slice in the picture, information regarding TMVP, such as identification information of the same-position picture (colPic), can be signaled.

[0161] FIG. 14a is a diagram showing an example of a picture header including information regarding TMVP.

[0162] Referring to FIG. 14a, whether TMVP is available at the sequence level can be determined based on predetermined flag information (e.g., sps_temporal_mvp_enabled_flag) at the sequence level. For example, sps_temporal_mvp_enabled_flag having a first value (e.g., 0) can indicate that TMVP is not available at the sequence level. In contrast, sps_temporal_mvp_enabled_flag having a second value (e.g., 1) can indicate that TMVP is available at the sequence level.

[0163] When TMVP is available at the sequence level, pic_temporal_mvp_enabled_flag can be signaled via the picture header. For example, pic_temporal_mvp_enabled_flag having a first value (e.g., 0) can indicate that TMVP is not available at the picture level. In contrast, pic_temporal_mvp_enabled_flag having a second value (e.g., 1) can indicate that TMVP is available at the picture level.

[0164] FIG. 14b is a diagram showing an example of a slice header including information related to TMVP.

[0165] Referring to FIG. 14b, num_ref_idx_active_override_flag can indicate the existence of num_ref_idx_active_minus1[i]. For example, num_ref_idx_active_override_flag having a first value (e.g., 0) can indicate that num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] do not exist. In contrast, num_ref_idx_active_override_flag having a second value (e.g., 1) can indicate that num_ref_idx_active_minus1[0] exists for P slices and B slices, and num_ref_idx_active_minus1[1] exists for B slices. On the other hand, when num_ref_idx_active_override_flag is not signaled, the value of num_ref_idx_active_override_flag can be inferred as the second value (e.g., 1).

[0166] num_ref_idx_active_minus1[i] can be used in the derivation of the variable NumRefIdxActive[i]. Here, the value obtained by subtracting 1 from the variable NumRefIdxActive[i] can indicate the maximum reference index of the i-th (where i is 0 or 1) reference picture list (RPL) used for decoding the current slice. In one example, the value of num_ref_idx_active_minus1[i] can be between 0 and 14 inclusive.

[0167] When the current slice is a B slice, the num_ref_idx_active_override_flag has a second value (e.g., 1), and num_ref_idx_active_minus1[i] does not exist, the value of num_ref_idx_active_minus1[i] can be inferred as the first value (e.g., 0). Or, when the current slice is a P slice, the num_ref_idx_active_override_flag has a second value (e.g., 1), and num_ref_idx_active_minus1[0] does not exist, the value of num_ref_idx_active_minus1[0] can be inferred as the first value (e.g., 0).

[0168] Based on the value of num_ref_idx_active_minus1[i], the variable NumRefIdxActive[i] can be derived as shown in Table 1 below.

[0169] [Table 1]

[0170] In Table 1, NumRefIdxActive[i] having the first value (e.g., 0) can indicate that the current slice cannot be decoded based on the reference index in the i-th (where i is 0 or 1) reference picture list.

[0171] In one example, when the current slice is a P slice, the value of NumRefIdxActive[0] can be greater than 0. Also, when the current slice is a B slice, the respective values of NumRefIdxActive[0] and NumRefIdxActive[1] can be greater than 0.

[0172] Continuing to refer to FIG. 14b, the collocated_from_l0_flag can indicate from which reference picture list among the reference picture list L0 and the reference picture list L1 the collocated picture (colPic) for TMVP is derived (i.e., the direction information of the collocated picture (colPic)). For example, the collocated_from_l0_flag having a first value (e.g., 0) can indicate that the collocated picture (colPic) is derived from the reference picture list L1. In contrast, the collocated_from_l0_flag having a second value (e.g., 1) can indicate that the collocated picture (colPic) is derived from the reference picture list L0. On the other hand, when the collocated_from_l0_flag is not signaled and the slice type is not a B slice, the value of the collocated_from_l0_flag can be inferred as the second value (e.g., 1). Or, when the collocated_from_l0_flag is not signaled and the slice type is a B slice, the value of the collocated_from_l0_flag can be inferred as the value obtained by subtracting 1 from pps_collocated_from_l0_idc. Here, pps_collocated_from_l0_idc can indicate whether the collocated_from_l0_flag exists in the slice header. For example, when the collocated_from_l0_flag exists in the slice header, pps_collocated_from_l0_idc can have the first value (e.g., 0). In contrast, when the collocated_from_l0_flag does not exist in the slice header, pps_collocated_from_l0_idc can have the second value (e.g., 1) or the third value (e.g., 2).

[0173] The collocated_ref_idx can indicate the reference picture index of the co-located picture (colPic) for TMVP. For example, when the slice type is a P slice or a B slice (i.e., the value of NumRefIdxActive[0] is greater than 0) and the collocated_from_l0_flag has a second value (e.g., 1), the collocated_ref_idx can point to a reference picture within the reference picture list L0. Here, the value of collocated_ref_idx can be 0 or greater and less than or equal to the value obtained by subtracting 1 from NumRefIdxActive[0]. Or, when the slice type is a B slice (i.e., the value of NumRefIdxActive[1] is greater than 0) and the collocated_from_l0_flag has a first value (e.g., 0), the collocated_ref_idx can point to a reference picture within the reference picture list L1. Here, the value of collocated_ref_idx can be 0 or greater and less than or equal to the value obtained by subtracting 1 from NumRefIdxActive[1]. On the other hand, when collocated_ref_idx is not signaled, the value of collocated_ref_idx can be inferred as a first value (e.g., 0).

[0174] In one example, it may be a restriction for bitstream compliance that the reference pictures specified by collocated_ref_idx are identical to each other for all slices related to the coded picture. In another example, it may be a restriction for bitstream compliance that the resolutions of the reference picture specified by collocated_ref_idx and the current picture are identical to each other and the value of RefPicIsScaled[collocated_from_l0_flag?0:1][collocated_ref_idx] is a first value (e.g., 0).

[0175] In the signaling mechanism for TMVP described above with reference to FIGS. 14a and 14b, information regarding the same-position picture (colPic), for example, identification information of the same-position picture (colPic) (e.g., collocated_from_l0_flag and collocated_ref_idx) can be signaled only via the slice header. However, when the same-position picture (colPic) that is the same for all slices within a picture is applied, according to the signaling mechanism for TMVP described above, there is a possibility of a problem that the signaling overhead increases because information regarding the same-position picture (colPic) has to be signaled for each slice.

[0176] To solve such a problem, according to an embodiment of the present disclosure, information regarding the same-position picture (colPic) for TMVP can be signaled via the slice header or can also be signaled via a higher-level syntax, for example, the picture header.

[0177] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0178] FIGS. 15 to 18 are diagrams showing an example of a picture header according to an embodiment of the present disclosure.

[0179] First, referring to FIG. 15, the picture header can include information regarding the same-position picture (colPic) for TMVP, including pic_collocated_from_l0_flag and pic_collocated_ref_idx.

[0180] pic_collocated_from_l0_flag and pic_collocated_ref_idx can only be signaled when TMVP is available at the picture level (e.g., pic_temporal_mvp_enabled_flag == 1).

[0181] pic_collocated_from_l0_flag can indicate from which reference picture list among reference picture list 0 (RPL0) and reference picture list 1 (RPL1) the collocated picture (colPic) for TMVP is derived (i.e., the direction information of the collocated picture (colPic)). For example, pic_collocated_from_l0_flag having a first value (e.g., 0) can indicate that the collocated picture (colPic) is derived from reference picture list 1 (RPL1). In contrast, pic_collocated_from_l0_flag having a second value (e.g., 1) can indicate that the collocated picture (colPic) is derived from reference picture list 0 (RPL0). On the other hand, when pic_collocated_from_l0_flag is not signaled and the number of entries in reference picture list 1 (RPL1) (e.g., num_ref_entries[1][RplsIdx[1]]) is 0, the value of pic_collocated_from_l0_flag can be inferred as the second value (e.g., 1). That is, the collocated picture (colPic) can be derived from reference picture list 0 (RPL0).

[0182] If pic_collocated_from_l0_flag has the second value (e.g., 1) and the number of entries in reference picture list 0 (RPL0) is greater than 1 (e.g., num_ref_entries[0][RplsIdx[0]] > 1), pic_collocated_ref_idx can be signaled. Alternatively, if pic_collocated_from_l0_flag has the first value (e.g., 0) and the number of entries in reference picture list 1 (RPL1) is greater than 1 (e.g., num_ref_entries[1][RplsIdx[1]] > 1), pic_collocated_ref_idx can be signaled.

[0183] pic_collocated_ref_idx can indicate the reference picture index of the co-located picture (colPic) for TMVP. For example, if pic_collocated_from_l0_flag has the second value (e.g., 1) and pic_collocated_ref_idx indicates one of the entries in reference picture list 0 (RPL0), the value of pic_collocated_ref_idx can be greater than or equal to 0 and less than or equal to the value obtained by subtracting 1 from the number of entries in reference picture list 0 (RPL0). In contrast, if pic_collocated_from_l0_flag has the first value (e.g., 0) and pic_collocated_ref_idx indicates one of the entries in reference picture list 1 (RPL1), the value of pic_collocated_ref_idx can be greater than or equal to 0 and less than or equal to the value obtained by subtracting 1 from the number of entries in reference picture list 1 (RPL1). On the other hand, if pic_collocated_ref_idx is not signaled, the value of pic_collocated_ref_idx can be inferred as the first value (e.g., 0).

[0184] In one example, information regarding the collocated picture (colPic) can be signaled in both the picture header and the slice header. For example, information regarding the collocated picture (colPic) is signaled via the picture header, but for at least a part of the slices related to the picture header, it can also be signaled via the slice header. In this case, the collocated picture (colPic) for the current block can be determined based on the information regarding the collocated picture (colPic) signaled via the slice header (e.g., collocated_from_l0_flag and collocated_ref_idx).

[0185] In other examples, information regarding the collocated picture (colPic) can be selectively signaled via the picture header or the slice header. In this case, whether the information regarding the collocated picture (colPic) is signaled via the picture header (or the slice header) can be determined based on predetermined signaling information (e.g., pps_rpl_info_in_ph_flag) within the picture header (and / or slice header). For example, when the signaling information has a second value (e.g., 1), the information regarding the collocated picture (colPic) can be signaled via the picture header. In contrast, when the signaling information has a first value (e.g., 0), the information regarding the collocated picture (colPic) can be signaled via the slice header. The signaling information can constitute the signaling conditions for pic_collocated_from_l0_flag together with pic_temporal_mvp_enabled_flag. Also, the signaling information can constitute the signaling conditions for collocated_from_l0_flag together with pic_temporal_mvp_enabled_flag within the slice header described above with reference to FIG. 14b.

[0186] When information about the collocated picture (colPic) for the same position is signaled via the slice header, the collocated picture (colPic) for the current block can be determined based on the information about the collocated picture (colPic) obtained via the slice header. In contrast, when information about the collocated picture (colPic) is signaled only via the picture header, the collocated picture (colPic) for the current block can be determined based on the information about the collocated picture (colPic) obtained via the picture header. For example, when collocated_ref_idx is not signaled via the slice header as described above with reference to FIG. 14b (e.g., pps_rpl_info_in_ph_flag == 1), the value of the collocated_ref_idx can be inferred to be the same as pic_collocated_ref_idx obtained via the picture header. When the pic_collocated_ref_idx is not signaled via the picture header as well as the slice header, the value of the collocated_ref_idx can be inferred to be a first value (e.g., 0).

[0187] On the other hand, that the collocated picture (colPic) for TMVP is an already decoded picture different from the current picture can be a limitation for bitstream compliance.

[0188] Another example of a picture header according to an embodiment of the present disclosure is as shown in FIG. 16.

[0189] Referring to FIG. 16, the picture header can be information about the collocated picture (colPic) for TMVP and can include col_ref_delta_poc_val.

[0190] col_ref_delta_poc_val can be signaled only if the TMVP is available for the current picture (e.g., pic_temporal_mvp_enabled_flag == 1).

[0191] col_ref_delta_poc_val can indicate the POC (picture order count) difference between the current picture and the co-located picture (colPic) for TMVP. For example, col_ref_delta_poc_val can have a value obtained by subtracting the POC of the co-located picture (colPic) from the POC of the current picture. Alternatively, col_ref_delta_poc_val can also have a value obtained by subtracting the POC of the current picture from the POC of the co-located picture (colPic). col_ref_delta_poc_val has a positive sign (+) or a negative sign (-) and can be represented in a signed integer type.

[0192] In one example, the information regarding the co-located picture (colPic) can be signaled in both the picture header and the slice header. For example, the information regarding the co-located picture (colPic) is signaled via the picture header, but for at least a part of the slices related to the said picture header, it can also be signaled via the slice header. In this case, the co-located picture (colPic) for the current block can be determined based on the information regarding the co-located picture (colPic) signaled via the slice header (e.g., collocated_from_l0_flag and collocated_ref_idx).

[0193] In other examples, information regarding the co-located picture (colPic) can be selectively signaled via the picture header or the slice header. In this case, whether the information regarding the co-located picture (colPic) is signaled via the picture header (or the slice header) can be determined based on predetermined signaling information (e.g., pps_rpl_info_in_ph_flag) within the picture header (and / or slice header). For example, when the signaling information has a second value (e.g., 1), the information regarding the co-located picture (colPic) can be signaled via the picture header. In contrast, when the signaling information has a first value (e.g., 0), the information regarding the co-located picture (colPic) can be signaled via the slice header.

[0194] When the information regarding the co-located picture (colPic) is signaled via the slice header, in the image decoding stage, the co-located picture (colPic) for the current block can be determined based on the information regarding the co-located picture (colPic) obtained via the slice header (e.g., collocated_from_l0_flag and collocated_ref_idx). In contrast, when the information regarding the co-located picture (colPic) is signaled only via the picture header, the co-located picture (colPic) for the current block can be determined based on the information regarding the co-located picture (colPic) obtained via the picture header (e.g., col_ref_delta_poc_val).

[0195] On the other hand, the co-located picture (colPic) for TMVP being a previously decoded picture different from the current picture can be a limitation for bitstream compliance.

[0196] Another example of a picture header according to an embodiment of the present disclosure is as shown in FIG. 17.

[0197] Referring to FIG. 17, the picture header can be information about the co-located picture (colPic) for TMVP, and can include col_ref_abs_delta_poc_val and col_ref_abs_delta_poc_sign_flag.

[0198] col_ref_abs_delta_poc_val and col_ref_abs_delta_poc_sign_flag can be signaled only when TMVP is available at the picture level (e.g., pic_temporal_mvp_enabled_flag == 1).

[0199] col_ref_abs_delta_poc_val can indicate the absolute value of the POC (picture order count) difference between the current picture and the co-located picture (colPic) for TMVP. For example, col_ref_abs_delta_poc_val can have the absolute value of the value obtained by subtracting the POC of the co-located picture (colPic) from the POC of the current picture. Alternatively, col_ref_abs_delta_poc_val can also have the absolute value of the value obtained by subtracting the POC of the current picture from the POC of the co-located picture (colPic). col_ref_abs_delta_poc_val indicates the absolute value of the POC difference between the current picture and the co-located picture (colPic), and different from col_ref_delta_poc_val in FIG. 16, it can be represented in unsigned integer type.

[0200] The col_ref_abs_delta_poc_sign_flag can indicate the sign of the POC difference between the current picture and the co-located picture (colPic) for the TMVP. For example, if the POC of the current picture is greater than the POC of the co-located picture (colPic) (or vice versa), the col_ref_abs_delta_poc_sign_flag can have a second value (e.g., 1) indicating a positive sign (+). In contrast, if the POC of the current picture is the same as or less than the POC of the co-located picture (colPic) (or vice versa), the col_ref_abs_delta_poc_sign_flag can have a first value (e.g., 0) indicating a negative sign (-).

[0201] In one example, the information regarding the co-located picture (colPic) can be signaled in both the picture header and the slice header. For example, the information regarding the co-located picture (colPic) is signaled via the picture header, but for at least a part of the slices associated with the picture header, it can also be signaled via the slice header. In this case, the co-located picture (colPic) for the current block can be determined based on the information regarding the co-located picture (colPic) signaled via the slice header (e.g., collocated_from_l0_flag and collocated_ref_idx).

[0202] In other examples, information regarding the co-located picture (colPic) can be selectively signaled via a picture header or a slice header. In this case, whether the information regarding the co-located picture (colPic) is signaled via the picture header (or the slice header) can be determined based on predetermined signaling information within the picture header (and / or the slice header). For example, when the signaling information has a second value (e.g., 1), the information regarding the co-located picture (colPic) can be signaled via the picture header. In contrast, when the signaling information has a first value (e.g., 0), the information regarding the co-located picture (colPic) can be signaled via the slice header.

[0203] When the information regarding the co-located picture (colPic) is signaled via the slice header, the co-located picture (colPic) for the current block can be determined based on the information regarding the co-located picture (colPic) obtained via the slice header (e.g., collocated_from_l0_flag and collocated_ref_idx). In contrast, when the information regarding the co-located picture (colPic) is signaled only via the picture header, the co-located picture (colPic) for the current block can be determined based on the information regarding the co-located picture (colPic) obtained via the picture header (e.g., col_ref_abs_delta_poc_val and col_ref_abs_delta_poc_sign_flag).

[0204] On the other hand, the co-located picture (colPic) for TMVP being a previously decoded picture different from the current picture can be a limitation for bitstream compliance.

[0205] Another example of a picture header according to an embodiment of the present disclosure is as shown in FIG. 18.

[0206] Referring to FIG. 18, the picture header can be information about the co-located picture (colPic) for TMVP, including col_ref_abs_delta_poc_val_minus1 and col_ref_abs_delta_poc_sign_flag.

[0207] col_ref_abs_delta_poc_val_minus1 and col_ref_abs_delta_poc_sign_flag can be signaled only when the TMVP mode is available at the picture level (e.g., pic_temporal_mvp_enabled_flag == 1).

[0208] col_ref_abs_delta_poc_val_minus1 can indicate a value obtained by subtracting 1 from the absolute value of the POC (picture order count) difference between the current picture and the co-located picture (colPic) for TMVP.

[0209] col_ref_delta_poc_sign_flag can indicate whether the value obtained by adding 1 to col_ref_abs_delta_poc_val_minus1 is greater than 0. For example, col_ref_abs_delta_poc_sign_flag having a first value (e.g., 0) can indicate that the value obtained by adding 1 to col_ref_abs_delta_poc_val_minus1 is less than 0. In contrast, col_ref_abs_delta_poc_sign_flag having a second value (e.g., 1) can indicate that the value obtained by adding 1 to col_ref_abs_delta_poc_val_minus1 is greater than 0.

[0210] In one example, information regarding the collocated picture (colPic) can be signaled in both the picture header and the slice header. For example, information regarding the collocated picture (colPic) is signaled via the picture header, but for at least a part of the slices related to the picture header, it can also be signaled via the slice header. In this case, the collocated picture (colPic) for the current block can be determined based on the information regarding the collocated picture (colPic) signaled via the slice header (e.g., collocated_from_l0_flag and collocated_ref_idx).

[0211] In other examples, information regarding the collocated picture (colPic) can be selectively signaled via the picture header or the slice header. In this case, whether the information regarding the collocated picture (colPic) is signaled via the picture header (or the slice header) can be determined based on predetermined signaling information (e.g., pps_rpl_info_in_ph_flag) within the picture header (and / or the slice header). For example, when the signaling information has a second value (e.g., 1), the information regarding the collocated picture (colPic) can be signaled via the picture header. In contrast, when the signaling information has a first value (e.g., 0), the information regarding the collocated picture (colPic) can be signaled via the slice header.

[0212] When information regarding the co-located picture (colPic) is signaled via a slice header, the co-located picture (colPic) for the current block can be determined based on the information regarding the co-located picture (colPic) obtained via the slice header (e.g., collocated_from_l0_flag and collocated_ref_idx). In contrast, when information regarding the co-located picture is signaled only via a picture header, the co-located picture (colPic) for the current block can be determined based on the information regarding the co-located picture (colPic) obtained via the picture header (e.g., col_ref_abs_delta_poc_val_minus1 and col_ref_abs_delta_poc_sign_flag).

[0213] On the other hand, the fact that the co-located picture (colPic) for TMVP is an already decoded picture different from the current picture can be a limitation for bitstream compliance. The fact that the co-located picture (colPic) is an already decoded picture different from the current picture can be achieved by signaling col_ref_abs_delta_poc_val_minus1 that always makes the POC difference between the current picture and the co-located picture (colPic) greater than 0.

[0214] As described above with reference to FIGS. 15 to 18, the picture header according to the embodiment of the present disclosure can include information regarding the co-located picture (colPic) for TMVP, for example, identification information of the co-located picture (colPic). Thereby, for a plurality of slices that refer to co-located pictures (colPic) that are identical to each other, the information regarding the co-located picture (colPic) can be signaled only once via the picture header, so that the signaling overhead for TMVP can be reduced and the efficiency of the signaling mechanism can be improved.

[0215] Hereinafter, with reference to FIGS. 19 to 21, an image encoding / decoding method according to an embodiment of the present disclosure will be described in detail.

[0216] FIG. 19 is a flowchart showing an image encoding method according to an embodiment of the present disclosure.

[0217] The image encoding method of FIG. 19 can be performed by the image encoding apparatus of FIG. 2. For example, steps S1910 and S1920 can be performed by the inter prediction unit 180. Also, step S1930 can be performed by the entropy encoding unit 190.

[0218] Referring to FIG. 19, when an inter prediction mode is applied to the current block, the image encoding apparatus can generate a prediction block of the current block based on the motion vector of the current block (S1910).

[0219] The inter prediction mode for the current block can be determined to be any one of various inter prediction modes (for example, merge mode, skip mode, Motion Vector Prediction (MVP) mode, Symmetric Motion Vector Difference (SMVD) mode, affine mode, etc.). For example, the image encoding apparatus can compare the rate-distortion (RD) costs for various inter prediction modes to select an optimal inter prediction mode, and determine the selected optimal inter prediction mode as the inter prediction mode for the current block.

[0220] The image encoding device can search for a block similar to the current block within a certain area (search area) of the reference picture for the current block through motion estimation, and derive a reference block whose difference from the current block is the smallest or below a certain criterion. The image encoding device can derive the motion vector of the current block based on the positional difference between the derived reference block and the current block.

[0221] For example, when the merge mode or skip mode is applied to the current block, the image encoding device can derive merge candidates from the surrounding blocks of the current block, and can configure a merge candidate list using the derived merge candidates. Here, the surrounding blocks of the current block can include spatial surrounding blocks and / or temporal surrounding blocks. The image encoding device can derive a reference block whose difference from the current block is the smallest or below a certain criterion among the reference blocks estimated by the motion information of the merge candidates included in the merge candidate list. In this case, the motion vector of the current block can be derived using the motion vector of the merge candidate related to the derived reference block.

[0222] In another example, when the MVP mode is applied to the current block, the image encoding device can derive motion vector predictor (MVP) candidates from the surrounding blocks of the current block, and can configure an MVP candidate list using the derived MVP candidates. Here, the surrounding blocks of the current block can include spatial surrounding blocks and / or temporal surrounding blocks. In this case, for example, the motion vector pointing to the reference block derived by the above-mentioned motion estimation can be used as the motion vector of the current block, and among the MVP candidates, the MVP candidate having the smallest difference from the motion vector of the current block can be selected as the MVP for the current block.

[0223] The picture symbolization device can derive a temporal motion vector predictor (TMVP) for the current block based on the same-position picture (colPic) (S1920).

[0224] The same-position picture (colPic) can be determined at the slice level or the picture level. For example, when the same-position picture (colPic) is determined at the slice level, different same-position pictures (colPic) can be selected for at least a part of the slices in the current picture. In this case, information about the same-position picture (colPic), for example, the identification information of the same-position picture (colPic), can be signaled to the image decoding device via the slice header. In contrast, when the same-position picture (colPic) is determined at the picture level, the same same-position picture (colPic) that is identical to each other can be selected for all slices in the current picture. In this case, information about the same-position picture (colPic), for example, the identification information of the same-position picture (colPic), can be signaled to the image decoding device via the picture header.

[0225] In one example, the information about the same-position picture (colPic) can be selectively signaled via the picture header or the slice header based on predetermined signaling information (for example, pps_rpl_info_in_ph_flag). For example, when the signaling information has a first value (for example, 1), the information about the same-position picture (colPic) can be signaled via the picture header. In contrast, when the signaling information has a second value (for example, 0), the information about the same-position picture (colPic) can be signaled via the slice header.

[0226] The TMVP for the current block can be derived based on the temporal neighboring blocks within the same-position picture (colPic). The temporal neighboring blocks can include the co-located block (colCb) for the current block. Here, the co-located block (colCb) can mean a block having the same position and / or the same size as the current block within the same-position picture (colPic).

[0227] In one example, the co-located block (colCb) can be determined as a luma coding block that covers the modified position from the first position (xColBr, yColBr) corresponding to the right-bottom corner of the current block within the same-position picture (colPic). Here, the first position (xColBr, yColBr) can mean the position (xCb + cbWidth, yCb + cbHeight) shifted by the width (cbWidth) and height (cbHeight) of the current block from the position (xCb, yCb) corresponding to the upper-left corner of the current block within the same-position picture. The first position (xColBr, yColBr) can be modified using an arithmetic shift operation for motion data compression. For example, the first position (xColBr, yColBr) can be modified to ((xColBr >> n) << n, (yColBr >> n) << n). Here, n can be an integer greater than or equal to 0.

[0228] In another example, the same position block (colCb) can be determined as a luma coding block covering a modified position from the second position (xColCtr, yColCtr) corresponding to the bottom - right sample (central right - bottom sample) among the 4 samples at the center of the current block within the same position picture. Here, the second position (xColCtr, yColCtr) can mean a position (xCb+(cbWidth>>1), yCb+(cbHeight>>1)) moved by half of the width (cbWidth) and height (cbHeight) of the current block respectively from the position (xCb, yCb) corresponding to the upper - left corner of the current block within the same position picture. The second position (xColCtr, yColCtr) can be modified using an arithmetic shift operation for motion data compression. For example, the second position (xColCtr, yColCtr) can be modified to ((xColCtr>>n)<<n, (yColCtr>>n)<<n). Here, n can be an integer greater than or equal to 0.

[0229] In the above - described example, n used in the arithmetic shift operation can mean a storage unit for storing the motion information of the temporal neighboring blocks. For example, when n is 3, the storage unit for storing the motion information of the temporal neighboring blocks can be 8×8 sample units. Alternatively, when n is 4, the storage unit for storing the motion information of the temporal neighboring blocks can be 16×16 sample units.

[0230] Alternatively, n used in the arithmetic shift operation can mean a read - out unit for reading out the motion information of the temporal neighboring blocks. For example, when n is 3, the read - out unit for reading out the motion information of the temporal neighboring blocks can be 8×8 sample units. In this case, the image coding device can identify an 8×8 sample unit including the first position ((xColBr>>n)<<n, (yColBr>>n)<<n) or the second position ((xColCtr>>n)<<n, (yColCtr>>n)<<n) modified based on the arithmetic shift operation, and read out the motion information of the identified 8×8 sample unit.

[0231] On the one hand, in FIG. 19, step S1920 is illustrated as being performed after step S1910, but the embodiments of the present disclosure are not limited thereto. For example, step S1920 may be performed before step S1910, or step S1920 may be performed simultaneously with step S1910.

[0232] The image encoding device can encode the motion vector of the current block based on the TMVP for the current block (S1930).

[0233] For example, when the merge mode or skip mode is applied to the current block, the TMVP for the current block can be included in the merge candidate list as a temporal merge candidate. The temporal merge candidate can also include a plurality of candidates including the TMVP. When the temporal merge candidate is selected for the current block, the image encoding device can encode the motion vector of the current block by encoding the merge index information indicating the temporal merge candidate.

[0234] In another example, when the MVP mode is applied to the current block, the TMVP for the current block can be included in the MVP candidate list as a temporal MVP candidate. The temporal MVP candidate can also include a plurality of candidates including the TMVP. When the temporal MVP candidate is selected for the current block, the image encoding device can encode the motion vector of the current block based on the TMVP of the temporal MVP candidate. For example, the image encoding device can derive an MVD (motion vector difference), which is the difference obtained by subtracting the TMVP from the motion vector of the current block, and encode the information regarding the MVD and the MVP index information indicating the temporal MVP candidate, thereby encoding the motion vector of the current block.

[0235] As described above, according to one embodiment of the present disclosure, information regarding the same-position picture (colPic) can be selectively signaled via a picture header or a slice header. For example, information regarding the same-position picture (colPic) can be signaled only once via a picture header, or can be adaptively signaled via a slice header. Alternatively, information regarding the same-position picture (colPic) can be signaled by both a picture header and a slice header. For example, information regarding the same-position picture (colPic) is signaled via a picture header, but can also be signaled via a slice header for at least a part of the slices related to the picture header. In this case, the same-position picture (colPic) for the current block can be determined based on the information regarding the same-position picture (colPic) signaled via the slice header. Thereby, the signaling overhead for TMVP can be reduced and the efficiency of the signaling mechanism can be improved.

[0236] FIG. 20 is a flowchart showing an image decoding method according to one embodiment of the present disclosure.

[0237] The image decoding method of FIG. 20 can be performed by the image decoding apparatus of FIG. 3. For example, step S2010 can be performed by the entropy decoding unit 210. Also, steps S2020 and S2030 can be performed by the inter prediction unit 260.

[0238] Referring to FIG. 20, when an inter prediction mode is applied to the current block, the image decoding apparatus can derive a temporal motion vector predictor (TMVP) for the current block based on the same-position picture (colPic) (S2010).

[0239] The same-position picture (colPic) can be determined based on information regarding the same-position picture (colPic) obtained from the picture header or the slice header. The specific method for determining the same-position picture (colPic) is as shown in FIG. 21.

[0240] FIG. 21 is a flowchart showing a method for determining the same-position picture according to an embodiment of the present disclosure.

[0241] Referring to FIG. 21, the image decoding apparatus can determine whether information regarding the same-position picture (colPic) is obtained from the slice header (S2110).

[0242] When information regarding the same-position picture (colPic), for example, identification information of the same-position picture (colPic), is obtained from the slice header (\"Yes\" in S2110), the image decoding apparatus can determine the same-position picture (colPic) for the current block based on the slice header (S2120). For example, based on the identification information of the same-position picture (colPic) obtained from the slice header described above with reference to FIG. 14b (for example, collocated_from_l0_flag and collocated_ref_idx), the same-position picture (colPic) for the current block can be determined.

[0243] In contrast, when information regarding the same-position picture (colPic) is not obtained from the slice header (No in S2110), the image decoding apparatus can determine the same-position picture (colPic) for the current block based on the picture header (S2130). For example, based on the identification information of the same-position picture (colPic) obtained from the picture header described above with reference to FIG. 15 (for example, pic_collocated_from_l0_flag and pic_collocated_ref_idx), the same-position picture (colPic) for the current block can be determined. Alternatively, based on the identification information of the same-position picture (colPic) obtained from the picture header described above with reference to FIGS. 16 to 18 (for example, col_ref_delta_poc_val, col_ref_abs_delta_poc_val, col_ref_abs_delta_poc_sign_flag, etc.), the same-position picture (colPic) for the current block can be determined.

[0244] In one example, whether information regarding the same-position picture (colPic) is obtained via the picture header (or slice header) can be determined based on predetermined signaling information (for example, pps_rpl_info_in_ph_flag) in the picture header (and / or slice header). For example, when the signaling information has a first value (for example, 0), information regarding the same-position picture (colPic) can be obtained via the slice header. In this case, different same-position pictures (colPic) can be applied to at least some of the slices in the current picture. In contrast, when the signaling information has a second value (for example, 1), information regarding the same-position picture (colPic) can be obtained via the picture header. In this case, the same same-position picture (colPic) that is identical to each other can be applied to all slices in the current picture.

[0245] Referring back to FIG. 20, the image decoding apparatus can derive a temporal motion vector predictor (TMVP) for the current block based on the same-position picture (colPic).

[0246] The TMVP can be derived based on temporal neighboring blocks within the same-position picture (colPic). The temporal neighboring blocks can include collocated blocks (colCb) having the same position and / or the same size as the current block within the same-position picture (colPic).

[0247] In one example, the collocated block (colCb) can be determined as a luma coding block that covers a modified position from a first position (xColBr, yColBr) corresponding to the right-bottom corner of the current block within the same-position picture (colPic). Here, the first position (xColBr, yColBr) can mean a position (xCb + cbWidth, yCb + cbHeight) that is moved by the width (cbWidth) and height (cbHeight) of the current block from a position (xCb, yCb) corresponding to the upper-left corner of the current block within the same-position picture (colPic). On the other hand, the first position (xColBr, yColBr) can be changed using an arithmetic shift operation for motion data compression. For example, the first position (xColBr, yColBr) can be modified to ((xColBr >> 3) << 3, (yColBr >> 3) << 3).

[0248] In another example, the same-position block (colCb) can be determined as a luma coding block covering a position modified from the second position (xColCtr, yColCtr) corresponding to the bottom-right sample (centeral right-bottom sample) among the four samples at the center of the current block within the same-position picture (colPic). Here, the second position (xColCtr, yColCtr) can mean a position moved by half of the width (cbWidth) and height (cbHeight) of the current block from the position (xCb, yCb) corresponding to the upper-left corner of the current block within the same-position picture (colPic), that is, (xCb+(cbWidth>>1), yCb+(cbHeight>>1)). On the other hand, the second position (xColCtr, yColCtr) can be changed using an arithmetic shift operation for motion data compression. For example, the second position (xColCtr, yColCtr) can be modified to ((xColCtr>>3)<<3, (yColCtr>>3)<<3).

[0249] The image decoding device can derive the motion vector of the current block based on the TMVP for the current block (S2020).

[0250] For example, when the merge mode or skip mode is applied to the current block, the image decoding device can configure a merge candidate list and derive the motion vector of the time merge candidate including the TMVP for the current block within the merge candidate list as the motion vector of the current block.

[0251] In another example, when the MVP mode is applied to the current block, the image decoding device can construct an MVP candidate list and derive, as the MVP of the current block, the motion vector of the temporal MVP candidate including the TMVP for the current block within the MVP candidate list. In this case, the image decoding device can derive the MVD (motion vector difference) of the current block based on the information regarding the MVD obtained from the bitstream, and can derive the motion vector of the current block by adding the MVP to the MVD.

[0252] The image decoding device can generate a prediction block of the current block based on the motion vector of the current block (S2030). For example, the image decoding device can derive the reference picture of the current block based on the reference picture index information obtained from the bitstream, and can generate the prediction block of the current block by using the samples of the reference block pointed to by the motion vector of the current block on the reference picture. In one example, a prediction sample filtering procedure can be further performed on all or part of the samples within the prediction block of the current block.

[0253] On the other hand, the image decoding device can generate a residual block of the current block based on the residual information obtained from the bitstream, and can restore the current block by adding the residual block to the prediction block. In one example, an in-loop filtering procedure or the like can be further performed on the restored image.

[0254] As described above, according to one embodiment of the present disclosure, information regarding the same-position picture (colPic) can be selectively signaled via a picture header or a slice header. For example, information regarding the same-position picture (colPic) can be signaled only once via a picture header, or can be adaptively signaled via a slice header. Alternatively, information regarding the same-position picture (colPic) can be signaled by both a picture header and a slice header. For example, information regarding the same-position picture (colPic) is signaled via a picture header, but for at least a part of the slices related to the picture header, it can also be signaled via a slice header. In this case, the same-position picture (colPic) for the current block can be determined based on the information regarding the same-position picture (colPic) signaled via the slice header. Thereby, the signaling overhead for TMVP can be reduced and the efficiency of the signaling mechanism can be improved.

[0255] The exemplary method of the present disclosure is represented in a series of operations for clarity of explanation, but this is not for limiting the order in which the steps are performed, and each step can be performed simultaneously or in a different order if necessary. To implement the method according to the present disclosure, it can further include other steps in the exemplified steps, or include the remaining steps except for some steps, or include additional other steps except for some steps.

[0256] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of checking the execution conditions and situations of the operation (step). For example, when it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation of checking whether the predetermined condition is satisfied.

[0257] The various embodiments of the present disclosure do not list all possible combinations, but are for explaining representative aspects of the present disclosure. The matters described in the various embodiments may be applied independently or in combinations of two or more.

[0258] In addition, the various embodiments of the present disclosure can be realized by hardware, firmware, software, or a combination thereof. In the case of realization by hardware, it can be realized by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0259] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied can be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, picture phone video devices, and medical video devices, etc., and can be used for processing video signals or data signals. For example, as an over-the-top (OTT) video device, it can include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.

[0260] FIG. 22 is a diagram exemplarily showing a content streaming system to which an embodiment according to the present disclosure can be applied.

[0261] As shown in FIG. 22, a content streaming system to which an embodiment of the present disclosure is applied can generally include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.

[0262] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a video camera directly generates a bitstream, the encoding server can be omitted.

[0263] The bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0264] The streaming server transmits multimedia data to a user device based on a user's request via a Web server, and the Web server can serve as a medium for informing the user of what services are available. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server can play a role of controlling commands / responses between each device in the content streaming system.

[0265] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0266] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device, for example, a smartwatch, smart glass, an HMD (head mounted display), a digital TV, a desktop computer, a digital signage, and the like.

[0267] Each server in the content streaming system can be operated as a distributed server. In this case, the data received from each server can be processed distributively.

[0268] The scope of the present disclosure includes software or machine-executable commands (for example, an operating system, an application, firmware, a program, etc.) that enable operations according to methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium on which such software or commands are stored and can be executed on a device or a computer.

Industrial Applicability

[0269] Examples according to the present disclosure are available for encoding / decoding images.

Claims

1. An image decoding method performed by an image decoding device, comprising: deriving a temporal motion vector predictor for the current block based on a co-located picture for the current block; deriving a motion vector differential for the current block based on information about the motion vector differential; deriving a motion vector for the current block based on the temporal motion vector predictor and the motion vector differential; generating a prediction block of the current block based on the motion vector; the co-located picture is determined based on identification information of the co-located picture, the identification information of the co-located picture including reference index information pointing to the co-located picture in a reference picture list for the current block; the identification information of the co-located picture further includes orientation information of the reference picture list that includes the co-located picture; determining whether the identification information of the co-located picture is obtained from a picture header based on predetermined signaling information indicating whether the temporal motion vector predictor is valid for a current picture including the current block; 10. A method for decoding an image, wherein the identification information of the co-located picture is obtained from a picture header for the current picture based on the temporal motion vector predictor being valid for a current picture that includes the current block.

2. An image coding method performed by an image coding device, comprising: generating a predicted block for the current block based on the motion vector of the current block; deriving a temporal motion vector predictor for the current block based on a co-located picture for the current block; deriving a motion vector differential for the current block based on the motion vector and the temporal motion vector predictor; encoding the motion vector differential of the current block; the identification information of the co-located picture is encoded, and the identification information of the co-located picture includes reference index information that points to the co-located picture in a reference picture list for the current block; the identification information of the co-located picture further includes orientation information of the reference picture list that includes the co-located picture; determining whether the identification information of the co-located picture is coded in a picture header based on whether the temporal motion vector predictor is valid for a current picture that includes the current block; signaling information indicating whether the temporal motion vector predictor is valid for a current picture that includes the current block; 10. A method for encoding an image, wherein the identification information of the co-located picture is encoded in a picture header for the current picture based on the temporal motion vector predictor being valid for the current picture that contains the current block.

3. A method for transmitting a bitstream generated by an image coding method, comprising: The image encoding method includes: generating a predicted block for the current block based on the motion vector of the current block; deriving a temporal motion vector predictor for the current block based on a co-located picture for the current block; deriving a motion vector differential for the current block based on the motion vector and the temporal motion vector predictor; encoding the motion vector differential of the current block; the identification information of the co-located picture is encoded, and the identification information of the co-located picture includes reference index information that points to the co-located picture in a reference picture list for the current block; the identification information of the co-located picture further includes orientation information of the reference picture list that includes the co-located picture; determining whether the identification information of the co-located picture is coded in a picture header based on whether the temporal motion vector predictor is valid for a current picture that includes the current block; signaling information indicating whether the temporal motion vector predictor is valid for a current picture that includes the current block; 10. The method of claim 9, wherein the identification of the co-located picture is encoded in a picture header for the current picture based on the temporal motion vector predictor being valid for the current picture that contains the current block.