Image encoding / decoding method based on bidirectional inter-prediction, method for transmitting a bitstream, and recording medium storing the bitstream

The proposed image encoding/decoding method addresses the challenge of high data volume in high-resolution images by utilizing AMVP and merge modes to enhance encoding/decoding efficiency and compression, achieving accurate prediction blocks and reduced transmission/storage costs.

JP2025521457APending Publication Date: 2025-07-10LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024573455
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-07-05
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images leads to higher transmission and storage costs due to increased data volume, necessitating a more efficient image compression technique.

Method used

An image encoding/decoding method that includes advanced motion vector prediction (AMVP) modes and merge modes, with methods for configuring and reordering merge candidate lists, and correcting motion information to improve encoding/decoding efficiency.

Benefits of technology

Enhances encoding/decoding efficiency, allows for more accurate prediction blocks, and improves compression efficiency through slice-level processing, particularly for low-delay slices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521457000001_ABST
    Figure 2025521457000001_ABST
Patent Text Reader

Abstract

An image encoding / decoding method, a bitstream transmission method, and a computer-readable recording medium for storing a bitstream are provided. The image decoding method according to the present disclosure is an image decoding method performed by an image decoding apparatus, and includes a step of restoring a picture from a bitstream, and a step of converting the restored picture based on any one of at least one candidate transformation, wherein the at least one candidate transformation is a rotation transformation or a symmetry transformation with respect to the picture, and the method can be an image decoding method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method, a method for transmitting a bitstream, and a recording medium storing a bitstream, and more particularly to a method for generating a prediction block in an AMVP mode and a merge mode for performing bidirectional inter prediction.

Background Art

[0002] Recently, demands for high-resolution and high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, have been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to conventional image data. The increase in the amount of information or bits to be transmitted brings about an increase in transmission costs and storage costs.

[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to propose an amvpMerge mode that predicts an arbitrary prediction direction in the AMVP mode and predicts another prediction direction in the merge mode.

[0006] Another object of the present disclosure is to propose a method for configuring a reference picture list for the amvpMerge mode.

[0007] Another object of the present disclosure is to propose a method for configuring a merge candidate list for the amvpMerge mode.

[0008] Furthermore, an object of the present disclosure is to propose a method for reordering a merge candidate list for the amvpMerge mode.

[0009] Furthermore, an object of the present disclosure is to propose a method for correcting motion information for the amvpMerge mode.

[0010] Furthermore, an object of the present disclosure is to propose a method for deriving bidirectional motion information for the amvpMerge mode.

[0011] Furthermore, an object of the present disclosure is to provide a non-transitory computer-readable recording medium that stores a bitstream generated by the image encoding method according to the present disclosure.

[0012] Furthermore, an object of the present disclosure is to provide a non-transitory computer-readable recording medium that stores a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0013] Furthermore, an object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method according to the present disclosure.

[0014] The technical problems to be solved in the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Means for Solving the Problems

[0015] An image decoding method according to one aspect of the present disclosure is an image decoding method performed by an image decoding apparatus, and includes a step of determining a prediction mode applied to each of the prediction directions of a current block, where the prediction mode includes an AMVP (advanced motion vector prediction) mode and a merge mode; a step of generating a prediction block for each of the prediction directions based on at least one reference picture set, at least one AMVP candidate, and at least one merge candidate; and a step of deriving a prediction block of the current block based on the prediction block, and the prediction mode of a certain prediction direction among the prediction directions is determined as the AMVP mode, and the prediction mode of another prediction direction is determined as the merge mode.

[0016] An image encoding method according to another aspect of the present disclosure is an image encoding method performed by an image encoding apparatus, and includes a step of determining a prediction mode applied to each of the prediction directions of a current block, where the prediction mode includes an AMVP (advanced motion vector prediction) mode and a merge mode; a step of generating a prediction block for each of the prediction directions based on at least one reference picture set, at least one AMVP candidate, and at least one merge candidate; and a step of deriving a prediction block of the current block based on the prediction block, and the prediction mode of a certain prediction direction among the prediction directions is determined as the AMVP mode, and the prediction mode of another prediction direction is determined as the merge mode.

[0017] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or apparatus of the present disclosure.

[0018] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding method or apparatus of the present disclosure.

[0019] The features briefly summarized and described above about the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present disclosure.

Advantages of the Invention

[0020] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0021] Also, according to the present disclosure, by predicting an arbitrary prediction direction in the AMVP mode and predicting other prediction directions in the merge mode, it is possible to generate a more accurate prediction block.

[0022] Also, according to the present disclosure, it is possible to improve the compression efficiency through slice-level processing and low-level processing for a low delay slice.

[0023] Also, according to the present disclosure, it is possible to provide a non-transitory computer-readable recording medium that stores a bitstream generated by the image encoding method according to the present disclosure.

[0024] Also, according to the present disclosure, it is possible to provide a non-transitory computer-readable recording medium that stores a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0025] Also, according to the present disclosure, it is possible to provide a method for transmitting a bitstream generated by an image encoding method.

[0026] The effects obtained in the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.

Brief Description of the Drawings

[0027]

Figure 1

[0028]

Figure 2

[0029]

Figure 3

[0030]

Figure 4

[0031]

Figures 5a - 5b

[0032]

Figures 6a - 6c

[0033]

Figure 7

[0034]

Figure 8

[0035]

Figures 9a - 9b

[0036]

Figure 10

[0037]

Figure 11

[0038]

Figures 12a - 12b

[0039]

Figures 13a - 13b

[0040]

Figure 14

[0041]

Figure 15

[0042]

Figure 16

[0043]

Figure 17

[0044]

Figures 18a - 18b

[0045]

Figure 19

[0046]

Figures 20a - 20b

[0047]

Figure 21

[0048]

Figure 22

[0049]

Figure 23

[0050]

Figure 24

[0051]

Figure 25

[0052]

Figure 26

Embodiments for Carrying Out the Invention

[0053] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.

[0054] When it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure in explaining the embodiments of the present disclosure, the detailed description thereof will be omitted. And in the drawings, parts not related to the description of the present disclosure are omitted, and the same reference numerals are given to the same parts.

[0055] In the present disclosure, when a component is "coupled", "connected", or "joined" to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Further, when a component "includes" or "has" another component, this means that, unless otherwise stated to the contrary, it does not exclude other components but can further include other components.

[0056] In the present disclosure, terms such as "first" and "second" are used only for the purpose of distinguishing one component from another and do not limit the order or importance between components, etc., unless otherwise specifically mentioned. Thus, within the scope of the present disclosure, the first component of one embodiment may be referred to as the second component in another embodiment, and similarly, the second component of one embodiment may be referred to as the first component in another embodiment.

[0057] In the present disclosure, components that are distinguished from each other are for clearly explaining their respective features and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, without further mention, such integrated or distributed embodiments are also included within the scope of the present disclosure.

[0058] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments constituted by a subset of the components described in one embodiment are also included within the scope of the present disclosure. Further, embodiments that include still other components in addition to the components described in various embodiments are also included within the scope of the present disclosure.

[0059] The present disclosure relates to the encoding and decoding of images. The terms used in the present disclosure can have the ordinary meanings in the technical field to which the present disclosure pertains, unless newly defined in the present disclosure.

[0060] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period, and a slice / tile is an encoding unit constituting a part of a picture, and one picture can be composed of one or more slices / tiles. Also, a slice / tile can include one or more CTUs (coding tree units).

[0061] In the present disclosure, "pixel" or "pel" can mean the smallest unit constituting one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0062] In the present disclosure, "unit" can indicate the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can be used interchangeably with terms such as "sample array", "block", or "area" as the case may be. In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients composed of M columns and N rows.

[0063] In the present disclosure, "current block" can mean any one of "current coding block", "current coding unit", "block to be coded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" can mean "current transformation block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".

[0064] Also, in the present disclosure, unless explicitly stated as a chroma block, "current block" can mean a block that includes all luma component blocks and chroma component blocks or "luma block of the current block". The luma component block of the current block can be explicitly expressed including an explicit description of the luma component block such as "luma block" or "current luma block". Also, the chroma component block of the current block can be explicitly expressed including an explicit description of the chroma component block such as "chroma block" or "current chroma block".

[0065] In the present disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".

[0066] In the present disclosure, "or" can be interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" can mean "additionally or alternatively".

[0067] Overview of Video Coding System

[0068] FIG. 1 is a diagram schematically showing a video coding system to which an embodiment according to the present disclosure can be applied.

[0069] A video coding system according to an embodiment can include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0070] The encoding device 10 according to an embodiment can include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding device 20 according to an embodiment can include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 can be called a video / image encoding unit, and the decoding unit 22 can be called a video / image decoding unit. The transmission unit 13 can be included in the encoding unit 12. The reception unit 21 can be included in the decoding unit 22. The rendering unit 23 can also include a display unit, and the display unit can be configured as a separate device or an external component.

[0071] The video source generation unit 11 can obtain a video / image through a process such as capture, synthesis, or generation of the video / image. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, and the like. The video / image generation device can include, for example, a computer, a tablet, a smartphone, and the like, and can (electronically) generate a video / image. For example, a virtual video / image can be generated via a computer or the like, and in this case, the video / image capture process can be replaced with a process in which related data is generated.

[0072] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0073] The transmission unit 13 can acquire the encoded video / image information or data output in the form of a bitstream, and transmit this to the receiving unit 21 of the decoding device 20 or other external objects via a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark), HDD, SSD, etc. The transmission unit 13 can include elements for generating a media file via a predetermined file format, and can include elements for transmission via a broadcast / communication network. The transmission unit 13 can be provided as a transmission device separate from the encoding device 12. In this case, the transmission device can include at least one processor that acquires the encoded video / image information or data output in the form of a bitstream, and a transmission unit that transmits this in file or streaming form. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0074] The decoding unit 22 can perform a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding unit 12 to decode the video / image.

[0075] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0076] Overview of Image Encoding Device

[0077] FIG. 2 is a diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure can be applied.

[0078] As shown in FIG. 2, the image encoding apparatus 100 can include an image dividing unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a "prediction unit". The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.

[0079] All or at least a part of a plurality of components constituting the image encoding apparatus 100 can be realized by one hardware component (for example, an encoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.

[0080] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units with a deeper depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and / or restoration described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit can be divided or partitioned from the final coding unit respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0081] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform prediction on a processing target block (current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information related to the prediction of the current block and transmit it to the entropy encoding unit 190. The information related to the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bit stream.

[0082] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to or away from the periphery of the current block according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of detail of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the settings. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0083] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU). The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and the motion vector difference and an indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0084] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of the current block, but also apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for the prediction of the current block can be called CIIP (combined inter and intra prediction). In addition, the prediction unit can also perform intra block copy (IBC) for the prediction of the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block within the current picture at a position a predetermined distance away from the current block. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed in the same manner as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.

[0085] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0086] The conversion unit 120 can apply a conversion technique to the residual signal to generate conversion coefficients (transform coefficients). For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen - Loeve Transform), GBT (Graph - Based Transform), or CNT (Conditionally Non - linear Transform). Here, GBT means the transform obtained from a graph when representing the relationship information between pixels as a graph. CNT means the transform obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can also be applied to pixel blocks having the same size of a square or can be applied to blocks of variable size that are not square.

[0087] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 130 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0088] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 190 can also encode, together or separately, information necessary for video / image restoration (such as values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (such as encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, transmitted information, and / or syntax elements referred to in the present disclosure can be encoded via the above-described encoding procedure and included in the bitstream.

[0089] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.

[0090] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.

[0091] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after being filtered as described later.

[0092] The filtering unit 160 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bit stream.

[0093] The modified restored picture transmitted to the memory 170 can be used as a reference picture by the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve the encoding efficiency.

[0094] The DPB in the memory 170 can store the modified restored picture for use as a reference picture by the inter prediction unit 180. The memory 170 can store the motion information of the blocks in which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of spatial neighboring blocks or temporal neighboring blocks. The memory 170 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.

[0095] Overview of Image Decoding Device

[0096] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.

[0097] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in a residual processing unit.

[0098] All or at least a part of the plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB and can be realized by a digital storage medium.

[0099] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 in FIG. 2 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be reproduced via a reproducing apparatus (not shown).

[0100] The image decoding device 200 can receive the signal output from the image encoding device of FIG. 2 in bitstream format. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding device can further use the information regarding the parameter set and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements referred to in the present disclosure can be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values entropy decoded by the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives the signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can be provided as a component of the entropy decoding unit 210.

[0101] On the other hand, the image decoding device according to the present disclosure can be called a video / image / picture decoding device. The image decoding device can also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.

[0102] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output the transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed by the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain the transform coefficients.

[0103] In the inverse conversion unit 230, the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0104] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0105] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described later, which is the same as described in the explanation of the prediction unit of the image encoding device 100.

[0106] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can be similarly applied to the intra prediction unit 265.

[0107] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can construct a motion information candidate list based on neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information regarding the prediction can include information indicating the prediction mode (technique) for the current block.

[0108] The adder 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block. The description of the adder 155 can be similarly applied to the adder 235. The adder 235 may also be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and can also be used for inter prediction of the next picture through filtering as described later.

[0109] Filtering unit 240 can apply filtering to the restored signal to improve subjective / objective image quality. For example, filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in memory 250, specifically in the DPB of memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0110] The (modified) restored picture stored in the DPB of memory 250 can be used as a reference picture by inter prediction unit 260. Memory 250 can store the motion information of the block where the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to inter prediction unit 260 for utilization as the motion information of spatial neighboring blocks or temporal neighboring blocks. Memory 250 can store the restored samples of the restored blocks in the current picture and transmit them to intra prediction unit 265.

[0111] In this specification, the embodiments described in filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of image encoding apparatus 100 can be applied to filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of image decoding apparatus 200 in the same or corresponding manner.

[0112] Inter - Prediction

[0113] The prediction units of the image encoding device 100 and the image decoding device 200 can derive prediction samples by performing inter prediction in block units. Inter prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values or motion information) of picture(s) other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can also be called collocated picture (colPic).For example, a motion information candidate list can be configured based on neighboring blocks of a current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block can be signaled. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In the case of skip mode, unlike merge mode, the residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference. The MVP mode may also be referred to as the advanced motion vector prediction (AMVP) mode.

[0114] The motion information can include L0 motion information and / or L1 motion information according to an inter prediction type (such as L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction can be called the L0 motion vector or MVL0, and the motion vector in the L1 direction can be called the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be called L0 prediction, the prediction based on the L1 motion vector can be called L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector can be called bi (Bi) prediction. Here, the L0 motion vector can represent a motion vector related to the reference picture list L0 (L0), and the L1 motion vector can represent a motion vector related to the reference picture list L1 (L1). The reference picture list L0 can include previous pictures as reference pictures in the output order before the current picture, and the reference picture list L1 can include subsequent pictures in the output order after the current picture. The previous picture can be called a forward (reference) picture, and the subsequent picture can be called a backward (reference) picture. The reference picture list L0 can further include subsequent pictures as reference pictures in the output order after the current picture. In this case, the previous pictures can be indexed first within the reference picture list L0, and the subsequent pictures can be indexed next. The reference picture list L1 can further include previous pictures as reference pictures in the output order before the current picture. In this case, the subsequent pictures can be indexed first within the reference picture list 1, and the previous pictures can be indexed next. Here, the output order can correspond to the POC (picture order count) order (order).

[0115] Template matching(TM)

[0116] Template Matching (TM) is a method for deriving motion vectors performed at the decoder end. By finding the template in the reference picture that is most similar to the template (hereinafter referred to as the "current template") adjacent to the current block (e.g., current coding unit, current CU), the motion information of the current block can be refined. The current template may be the upper adjacent block and / or the left adjacent block of the current block, or a part of these adjacent blocks. Also, the reference template can be determined to be the same size as the current template.

[0117] Once the initial motion vector of the current block is derived, a search for a better motion vector can be performed in the peripheral area of the initial motion vector. For example, the range of the peripheral area where the search is performed can be within the [-8, +8]-pel search area centered on the initial motion vector. Also, the size of the search step for performing the search can be determined based on the AMVR mode of the current block. Also, template matching may be performed continuously with the bilateral matching process in the merge mode.

[0118] When the prediction mode of the current block is the AMVP mode, the motion vector predictor candidate (MVP candidate) can be determined based on the template matching error. For example, a motion vector predictor candidate (MVP candidate) that minimizes the error between the current template and the reference template can be selected. Then, template matching for improving the motion vector can be performed for the selected motion vector predictor candidate. At this time, template matching for improving the motion vector may not be performed for the motion vector predictor candidates that are not selected.

[0119] More specifically, the improvement for the selected motion vector predictor candidate can start from full-pel (integer-pel) accuracy within the [-8, +8]-pel search region using iterative diamond search. Or, in the case of the 4-pel AMVR mode, it can start from 4-pel accuracy. Thereafter, depending on the AMVR mode, half-pel and / or quarter-pel accuracy search can continue. According to the search process, the motion vector predictor candidate can maintain the same motion vector accuracy as that indicated by the AMVR mode even after the template matching process. In the iterative search process, if the difference between the previous minimum cost and the current minimum cost is smaller than any threshold value, the search process ends. The threshold value can be the same as the area of the block, i.e., the number of samples within the block. Table 1 is an illustration of search patterns for the AMVR mode and the merge mode with AMVR.

[0120]

Table 1

[0121] When the prediction mode of the current block is the merge mode, a similar search method can be applied to the merge candidates indicated by the merge index. As shown in Table 1 above, template matching can be performed up to 1 / 8-pel accuracy, or accuracy below half-pel can be skipped, which can be determined dependently on whether an Alternative Interpolation Filter is used based on the merge motion information. At this time, the Alternative Interpolation Filter may be a filter used when AMVR is in the half-pel mode. Also, when template matching is available, depending on whether bilateral matching (BM) is available, the template matching can operate as an independent process, or can operate as an additional motion vector improvement process between block-based bilateral matching and sub-block-based bilateral matching. Whether the template matching is available and / or whether the bilateral matching is available can be determined by checking the available conditions. The accuracy of the motion vector in the above can mean the accuracy of the motion vector difference (MVD).

[0122] Examples

[0123] This application relates to inter prediction, and in the process of generating a prediction block through two or more reference pictures, any prediction direction constitutes an AMVP prediction candidate to induce prediction motion information, and the opposite prediction direction constitutes a MERGE prediction candidate to induce prediction motion information, thereby generating a final prediction block for the current block in the amvpMerge mode.

[0124] When the motion information is guided to the AMVP prediction candidate in an arbitrary prediction direction that guides the motion information in the process of guiding the motion information, when specific conditions are satisfied, the motion information can be guided without explicitly signaling the MVD, and when the specific conditions are not satisfied, the MVD can be explicitly signaled to correct the motion information in the AMVP prediction direction and then the motion information can be guided. The present application proposes a method for the case where the reference picture list of the slice to be encoded / decoded is composed only of POCs with values smaller than the POC of the picture (current picture) including the current slice, or the reference picture list of the slice is composed only of POCs with values larger than the POC of the current picture including the current slice.

[0125] The embodiments described below may be performed alone or in combination of two or more. Also, the embodiments described below can be performed by the image encoding device 100 and the image decoding device 200.

[0126] Example 1 - 1

[0127] Embodiment 1-1 proposes a method for determining whether to execute the amvpMerge mode at the SPS, PPS, Picture header, or Slice header level.

[0128] The flag for determining the execution status can be signaled via a bitstream. When the value of the flag is 1, motion information can be derived through the amvpMerge mode at the level (such as picture, slice, coding unit, prediction unit, etc.) that references the level at which the flag is signaled. The flag can be signaled via at least one level among SPS, PPS, Picture header, or Slice header. The flag can be defined as sps_amvpMerge_enabled_flag, pps_amvpMerge_enabled_flag, ph_amvpMerge_enabled_flag, or sh_amvpMerge_enabled_flag, etc.

[0129] Example 1 - 2

[0130] Examples 1-2 propose a method for constructing a reference picture list for the amvpMerge mode.

[0131] When the reference picture list of the slice to be encoded / decoded meets the conditions proposed, amvpMerge prediction candidates for the coding units or prediction units within the slice can be constructed and motion information can be derived.

[0132] The method of Examples 1-2 can be performed on a slice-by-slice basis as shown in FIG. 4. That is, when one picture is composed of multiple slices, the execution status of the amvpMerge mode can be determined for each slice unit.

[0133] Referring to FIG. 4, sliceIdx indicating any one of a plurality of slices can be set to 0, and numSliceInPic indicating the number of slices in the picture can be set to N (S402). If the reference picture list construction process has not been performed for a part of the slice (S404), the reference picture list construction process can be performed for the slice (sliceIdx) (S406).

[0134] It can be considered whether all of the reference pictures in the reference picture list of the slice have a POC (picture order count) prior to the current picture (forward) (S408). All of the reference pictures in the reference picture list that are prior to the current picture can correspond to low delay slices. If all of the reference pictures in the reference picture list of the slice have a POC prior to the current picture, bLDCFlag indicating whether it is a low delay slice is set to true (S412), and if even a part of the reference pictures in the reference picture list of the slice has a POC that is later than the current picture (backward), bLDCFlag can be set to false (S410).

[0135] It can be determined whether the type of the slice is a B (bi-predictive) type (S414). If the type of the slice is a B type, a process of constructing a valid reference picture set in the amvpMerge mode for the reference picture list in the L0 direction (REF_PIC_LIST_0) and the reference picture list in the L1 direction (REF_PIC_LIST_1) can be performed (S416). If the type of the slice is not a B type, the process of S416 may not be performed.

[0136] Proceed to the next slice after the said slice (sliceIdx = sliceIdx + 1), and the processes from S404 to S416 can be performed. By repeatedly performing these processes, if the reference picture list construction process for all slices is completed (sliceIdx = numSliceInPic), the reference picture list construction process can end.

[0137] According to FIG. 4, when the slice type allows two or more motion prediction information (B slice), a process of considering the existence of a reference picture for which the proposed amvpMerge mode is executable (available for amvpMerge), and a process of determining whether to enable the amvpMerge mode considering the reference picture for which the merge mode is executable (available for the merge mode) can be performed.

[0138] The process of considering the existence of a reference picture available for amvpMmerge can be performed as shown in FIGS. 5a and 5b for each direction (L0, L1) of each reference picture.

[0139] Referring to FIGS. 5a and 5b, the initial values of variables can be set (S502). CurPOC indicates the POC of the current picture included in the current slice, bAmvpMergeEnabledFlag indicates whether to enable the amvpMerge mode in the said slice, amvpMergeValldRefPair and amvpMergeValidRefIdx respectively indicate whether a pair of a reference picture in a reference picture list in a certain prediction direction and a reference picture in a reference picture list in another prediction direction is available for the amvpMerge mode and the index of the said picture, and amvpMergeUniqValidRefIdx indicates whether a certain reference picture in REF_PIC_LIST_X is the only reference picture available for the amvpMerge mode.

[0140] The determination of whether the reference picture set of the amvpMerge mode can be configured can be performed for any reference picture (refIdxInListX) in the reference picture list X (X is 0 or 1) and all reference pictures (refIdxListY) in the reference picture list Y (Y is 1 - X) (S520, S528 - S536). After the determination for all reference pictures (refIdxInListY) is completed (S520), the process proceeds to the next reference picture (S526, S520, S528 - S536), so that the determination can be performed for all reference pictures in the reference picture list X and the reference picture list Y.

[0141] If all the reference pictures in the reference picture list of the current slice are forward, that is, if the POC of all reference pictures is smaller than the POC of the current picture, then bLDCFlag is set to true. Only when the value of bLDCFlag is false (S504), the determination of whether the reference picture set of the amvpMerge mode can be configured for the reference picture list L0 and the reference picture list L1 is made (S506 - S536), and the reference picture set for the amvpMerge mode can be configured only if possible.

[0142] REF_PIC_LIST_X can be represented by REF_PIC_LIST_0 and REF_PIC_LIST_1. REF_PIC_LIST_0 indicates the reference picture list in the L0 direction, and REF_PIC_LIST_1 indicates the reference picture list in the L1 direction.

[0143] Also, (1 - REF_PIC_LIST_X) indicates the reference picture list in the opposite direction of REF_PIC_LIST_X. For example, when REF_PIC_LIST_X is REF_PIC_LIST_0, (1 - REF_PIC_LIST_X) can mean REF_PIC_LIST_1, and the opposite case can also hold.

[0144] When amvpMergeValidRefPair[refIdxInListX][refIdxInListY] is true, the reference pictures corresponding to refIdxInListX and refIdxInListY form a set, and this reference picture set corresponds to the reference pictures that can induce the prediction information in the amvpMerge mode.

[0145] When amvpMergeUniqValidRefIdx[REF_PIC_LIST_X] is not NOT_VALID, this indicates the index of the reference picture in REF_PIC_LIST_X that can uniquely induce the motion information in the amvpMerge mode. That is, when the value is not NOT_VALID, the reference picture index for the amvpMerge mode in the coding unit or prediction unit can be induced through the reference picture index of amvpMergeUniqValidRefIdx[REF_PIC_LIST_X] without explicitly signaling the reference picture index.

[0146] amvpMergeValidRefIdx[REF_PIC_LIST_X][refIdxInListX] can indicate whether the reference picture corresponding to refIdxInListX in REF_PIC_LIST_X is a reference picture that can induce the motion information in the amvpMerge mode. When the value of amvpMergeValidRefIdx[REF_PIC_LIST_X][refIdxInListX] is true, the reference picture corresponding to refIdxInListX in REF_PIC_LIST_X can correspond to the reference pictures that can induce the motion information in the amvpMerge mode.

[0147] Condition 1 (condition1, S528) is a condition with respect to the state of a reference picture, and can include whether the reference pictures corresponding to refIdxInListX and refIdxInListY are encoded / decoded by reference picture resampling (RPR), and / or whether they are long-term reference pictures, and / or whether there is a weight for weighted prediction.

[0148] That is, as one example, Condition 1 can be specified as the case where all or one or more of the following six conditions are satisfied.

[0149] 1. The reference picture corresponding to refIdxInListX is not RPR

[0150] 2. The reference picture corresponding to refIdxInListY is not RPR

[0151] 3. The reference picture corresponding to refIdxInListX is not a long-term reference picture

[0152] 4. The reference picture corresponding to refIdxInListY is not a long-term reference picture

[0153] 5. The weight for weighted prediction for the reference picture corresponding to refIdxInListX is not signaled

[0154] 6. The weight for weighted prediction for the reference picture corresponding to refIdxInListY is not signaled

[0155] If it is assumed that all of the Condition 1 is satisfied, it is determined whether the reference pictures corresponding to refIdxLX and refIdxLY can be used as a reference picture set for the amvpMerge mode (S532). If the query of S532 is not satisfied, a determination for the next reference picture is made (S536).

[0156] Condition 2 (condition2, S532) indicates a condition for determining whether it is suitable as a reference picture for the amvpMerge mode in consideration of the POCs of the reference pictures corresponding to refIdxLX and refIdxLY.

[0157] If the POC of the reference picture corresponding to refIdxLX is defined as pocLX, the POC of the reference picture corresponding to refIdxLY is defined as pocLY, and the POC of the current picture is defined as curPOC (S530), Condition 2 can be satisfied if one reference picture has a POC smaller than that of the current picture and one reference picture has a POC larger than that of the current picture.

[0158] That is, Condition 2 is represented by Equation 1.

[0159]

Equation

[0160] Example 1 - 3

[0161] Examples 1-3 propose a method for determining whether to apply the amvpMerge mode in a coding unit or a prediction unit, and information explicitly signaled in the bitstream therefor.

[0162] Referring to FIGS. 6A to 6C, a skip_flag indicating whether skip mode is applied is signaled (S602). If skip mode is not applied (S604), a pred_mode indicating prediction mode is signaled (S606). When pred_mode indicates an inter prediction mode (S608), a merge_flag indicating whether merge mode is applied is signaled (S610). If merge mode is not applied (S612), it is determined whether amvpMerge mode is activated (S614).

[0163] If amvpMerge mode is not applied (S618), information indicating the prediction direction (S622), information indicating whether affine prediction mode is applied (S624), information indicating whether smvd (symmetric MVD) mode is applied (S626), information regarding the bi - directional weight index (S628), information indicating a pair of reference pictures (S630), and information indicating the reference pictures in the reference picture list (S632) can be signaled.

[0164] If amvpMerge mode is applied (S618), it is determined whether the prediction direction is bi - directional (S634). Among the AMVP mode and the merge mode, the mode applied to each direction is determined (S636). The reference picture index and the motion information predictor for the AMVP mode and the determined prediction direction are signaled (S638, S640).

[0165] If the affine mode is not applied (S642), it is determined whether merge mode is applied in any direction and whether the value of the motion information predictor index in the opposite direction is less than 2 (S644). If it is true, the MVD is set to zero (S648). If it is false, the MVD in the opposite direction is signaled (S646). When the affine mode is applied (S642), the MVDs for the control points cp0, cp1, cp2 are signaled (S650 - S654).

[0166] It is determined whether the prediction direction is unidirectional (S656), it is determined whether the smvd mode is applied (S658), it is determined which mode of the AMVP mode and the merge mode is applied to each direction (S660), and the reference picture index and the motion information predictor for the AMVP mode and the determined prediction direction are signaled (S662, S664).

[0167] If the affine mode is not applied (S666), it is determined whether the merge mode is applied in any direction and whether the value of the motion information predictor index in the opposite direction is less than 2 (S668). If it is true, the MVD is set to zero (S672). If it is false, the MVD in the opposite direction is signaled (S670). When the affine mode is applied (S666), the MVDs for the control points cp0, cp1, and cp2 are signaled (S674 - S678).

[0168] When bAmvpMergeEnabled induced at the slice level is true (1), the information required for the amvpMerge mode can be signaled through the process of amvpMerge_mode() in FIG. 7.

[0169] Referring to FIG. 7, it is determined whether bidirectional prediction is possible (S702). If it is possible, the amvpMergeFlag is signaled (S704), and it can be determined whether the amvpMerge mode is applied based on the amvpMergeFlag (S706). When the prediction direction is bidirectional (S708), in order to distinguish an arbitrary prediction direction that induces motion information in the AMVP prediction candidate and an arbitrary prediction direction that induces motion information in the merge prediction candidate, the enable of the adaptive reference list (ARL) defined at the SPS / PPS / PH / SH level (S710) and the mvdL1ZeroFlag induced from the PH are determined (S710, S712).

[0170] When useARL() is disabled with reference to the condition and mvdL1ZeroFlag is 0, a 1-bit flag (mergeDir) is signaled (S714, S718), and any prediction direction that induces motion prediction information into the merge prediction candidate can be determined. Conversely, when useARL() is enabled or mvdL1ZeroFlag is 1, motion information can be induced from the merge prediction candidate in the L1 prediction direction without separate signaling (S716, S718).

[0171] The amvpMergeModeFlag[] array specifies how to induce motion information for any prediction direction. When amvpMergeModeFlag[0] is true, it means that L0 induces motion information from the merge prediction candidate and L1 induces motion information from the AMVP prediction candidate. Conversely, when amvpMergeModeFlag[1] is true, it means that L1 induces motion information from the merge prediction candidate and L0 induces motion information from the AMVP prediction candidate.

[0172] Example 1 - 4

[0173] Examples 1-4 propose a process for determining a reference picture list and syntax information therefor to induce motion information using the amvpMerge mode in a coding unit or a prediction unit.

[0174] The proposed method selectively determines one of two methods of reference picture index signaling to transmit the reference picture index for the amvpMerge mode. The two proposed methods can be diagrammed in FIGS. 8, 9a, 9b, and 10, and each reference picture index method follows the flowcharts of FIGS. 6a-6c.

[0175] The first method follows the process shown in FIG. 8. numCombinedListAmvpMerge in FIG. 8 indicates the number of reference picture sets for which prediction in amvpMerge mode can be executed. Also, useARL() indicates whether or not an adaptive reference list construction method signaled by SPS / PPS / PH / SH or implicitly induced is enabled.

[0176] Referring to FIG. 8, if useARL() is true and the merge mode is not applied without applying the IBC mode (S802), it is determined whether the prediction direction is bidirectional (S804). If the prediction direction is not bidirectional (S804), or if the amvpMerge mode is not applied (S806), a value obtained by subtracting 1 from the number of reference pictures is set to be the same as a value obtained by subtracting 1 from the number of combined reference pictures (S808). If the prediction direction is bidirectional (S804) and the amvpMerge mode is applied (S806), a value obtained by subtracting 1 from the number of reference pictures is set to be the same as a value obtained by subtracting 1 from the number of combined amvpMerge reference picture sets (S810).

[0177] If the set value is greater than 0 (S812), the reference picture index is signaled (S814). However, if the set value is 0, that is, if the combined amvpMerge reference picture set is unique (when numRefMinus1 is 0), the reference picture index is implicitly induced without separate signaling (S816).

[0178] numCombineListAmvpMerge indicates the number of sets of reference picture sets (combinedListAmvpMerge) for amvpMerge induced at the slice level, and the reference picture sets (combinedListAmvpMerge) can be induced through the processes shown in FIGS. 9a and 9b.

[0179] Referring to FIGS. 9a and 9b, the reference picture set (refIdxcombinedListAmvpMerge) value is initialized (S902), the reference picture index (refIdx), the number of reference pictures in the reference picture list L0 is set as the value of numRefL0, the number of reference pictures in the reference picture list L1 is set as the value of numRefL1, and the larger value of numRefL0 and numRefL1 is set as the value of numRefLX (S904).

[0180] While the value of refIdx is less than the value of numRefLX (S906), if the value of refIdx is less than the value of numRefL0 (S908), a reference picture set for the L0 direction is induced (S910 - S914). Specifically, it is determined whether the reference picture corresponding to refIdx is available for the amvpMerge mode (S910). If it is available, the reference picture is added to the reference picture set (S912), and the same process is repeated for the next reference picture (S914). Even when the process of S910 is false, the same process is repeated for the next reference picture.

[0181] While the value of refIdx is less than the value of numRefLX (S906l), if the value of refIdx is not less than the value of numRefL0 (S908) and the value of refIdx is less than the value of numRefL1 (S916), a reference picture set for the L1 direction is induced (S918 - S924). Specifically, it is determined whether refIdx is not valid for each prediction direction (S918). If it is not valid, it is determined whether the reference picture corresponding to refIdx is available for the amvpMerge mode (S920). If it is available, the reference picture is added to the reference picture set (S922), and the same process is repeated for the next reference picture (S924). Even when the processes of S916, S918, and S920 are false, the same process is repeated for the next reference picture (S924).

[0182] If the value of refIdx is not less than the value of numRefLX (S906), it returns to the first reference picture again (S926). If the value of refIdx is less than the value of numRefL1, it is determined whether refIdx is valid for each prediction direction (S930). If it is not valid, it is determined whether the reference picture corresponding to refIdx is available for the amvpMerge mode (S932). If it is available, the reference picture is added to the reference picture set (S934), and the same process is repeated for the next reference picture (S936). Even when the processes of S930 and S932 are false, the same process is repeated for the next reference picture (S936).

[0183] The proposed second method performs each process shown in FIG. 10. It is determined whether the smvd mode is applied (S1002). If it is applied, the reference index for the smvd mode is signaled (S1016). If it is not applied, it is determined whether useARL() is applied (S1004). If useARL() is not applied, it is determined whether amvpMerge is applied (S1006). If amvpMerge is applied, the prediction directions to which the merge mode and the AMVP mode are applied are determined (S1008). If the mode applied in the 1-REF_PIC_LIST_X direction is the merge mode, it is determined whether the reference picture set is unique (S1010). If it is unique, the reference picture index is implicitly derived (S1012). If it is not unique, the reference picture index is signaled (S1014).

[0184] As shown in Figure 10, when the value of amvpMergeModeFlag[1-REF_PIC_LIST_X] is 1, that is, when amvpMergeFlag is 1 in the coding unit and the amvpMergeModeFlag[REF_PIC_LIST_X] value is induced to false and the prediction direction of REF_PIC_LIST_X induces motion information from the AMVP prediction candidates, the reference picture index is signaled. The method in Figure 10 can be performed when the reference picture index is not signaled via the method in Figure 8, and when the reference picture set induced from the slice level is unique (when amvpMergeUniqValidRefIdx[REF_PIC_LIST_X] is not NOT_VALID), the reference picture index is implicitly induced without separate signaling.

[0185] Example 1 - 5

[0186] Examples 1-5 propose a method for constructing prediction candidates in the process of inducing motion information via the amvpMerge mode in a coding unit or a prediction unit.

[0187] As shown in Figure 11, any prediction direction for inducing motion information from the AMVP prediction candidates and any prediction direction for inducing motion information from the merge prediction candidates can be defined as refListAmvp and refListMerge respectively according to the value of amvpMergeModeFlag, and the implicitly or explicitly induced AMVP reference index can be defined as amvpRefIdx (S1102).

[0188] It is limited to the case where the reference picture indicated by amvpRefIdx is a reference picture defined so that a prediction block can be generated through amvpMerge prediction (S1104), an AMVP candidate list construction process (S1106), a process of deriving a predicted motion vector based on the AMVP candidate list (S1108), a merge candidate list construction process (S1110), a process of arranging the merge candidate list based on the bi-lateral matching error (cost) (S1112), a process of deriving a motion vector predicted based on the merge candidate list (S1114), a process of correcting the predicted motion vector (S1116), and a process of deriving an MVD from the corrected motion vector (S1118) are performed.

[0189] In Embodiments 1-5, methods for the AMVP candidate list construction process and the merge candidate list construction process are proposed.

[0190] The AMVP candidate list construction process is shown in FIGS. 12a and 12b. Referring to FIGS. 12a and 12b, the values of the variables numCand, maxStorage, and Similarity are set (S1202), and the motion information of the spatial candidates LB, RT, and LT is sequentially added to the list and the value of numCand is updated (S1204).

[0191] The identity of the motion information in the AMVP candidate list is determined and the value of numCand is updated (S1206), and it is determined whether the TMVP mode is applied, whether the number of pieces of motion information in the AMVP candidate list is less than maxStorage, and whether the size of the current block is valid (S1208). If all three conditions are satisfied, the TMVP candidate can be included in the AMVP candidate list based on the identity determination (S1210). If any one of the three conditions is not satisfied, it is determined whether the number of pieces of motion information in the AMVP candidate list is less than maxStorage and whether condition 1 is satisfied (S1212). If both of the two conditions are satisfied, the motion information of the non-adjacent candidate is added to the AMVP candidate list (S1214). If either of the two conditions is not satisfied, it is determined whether the number of pieces of motion information in the AMVP candidate list is less than maxStorage (S1216). If one condition is satisfied, the HMVP candidate can be included in the AMVP candidate list based on the identity determination (S1218).

[0192] Thereafter, it is determined whether DMVD (decoder side motion vector derivation) is applied (S1220). When DMVD is applied, after the process of sorting the AMVP candidate list based on the TM cost (S1222) and the process of correcting the candidate having the highest TM cost (S1224) are performed, the process of adding Zero MV (S1226) is performed. Here, the highest TM cost can correspond to the TM cost having the minimum value. When DMVD is not applied, the process of S1226 is immediately performed.

[0193] After that, it is determined whether the amvpMerge mode is applied (S1228). If it is applied, when the number of AMVP candidates is 1 (S1230), the second AMVP candidate is generated with the value of the first AMVP candidate, and the number of AMVP candidates is increased to 2 (S1232). If the number of candidates is not 1 (S1230), the second AMVP candidate is generated with the value of the first AMVP candidate, and the third AMVP candidate is generated with the value of the second AMVP candidate, thereby increasing the number of AMVP candidates to 3 (S1234).

[0194] In the S1212 process, condition 1 indicates whether one or more of the following two conditions are satisfied.

[0195] 1. Whether the current block is a multi hypothesis inter prediction method that generates a final predicted block by weighted averaging three or more predicted blocks

[0196] 2. Whether template matching is enabled for the current block

[0197] By the method of FIG. 12, in the process of deriving amvpMerge motion information, the motion information of the AMVP reference picture list is derived as one or two depending on the case, and this can be extended to two or three.

[0198] The AMVP candidate list construction process is shown in FIGS. 13a and 13b.

[0199] Referring to FIGS. 13a and 13b, it is determined whether a merge mode based on template matching (TM merge mode) is applied (S1302). If it is applied, a threshold is set based on the number of pixels (S1304). If it is not applied, the threshold is set to a predetermined value 1 (S1306). After the process of S1304, depending on whether the AML mode is applied (S1308), if it is applied, the maximum number of merge candidates is set to the maximum number of TM merge candidates for the TM merge mode (S1310). If it is not applied, the maximum number of merge candidates is set to the maximum number of TM merge candidates signaled via SPS / PPS / PS / PH / SH (S1312). After the process of S1306, the maximum number of merge candidates is set to the maximum number of merge candidates signaled via SPS / PPS / PH / SH (S1314).

[0200] When the amvpMerge mode is applied (S1316), the values of amvpRefList and mergeDir are derived (S1320, S1322) according to the applied mode (either the AMVP mode or one of the merge candidates) (S1318).

[0201] Also, when the spatial, temporal, non - adjacent spatial, HMVP, pairwise, affine, HMVP, and zero MV that can be candidates (S1324) satisfy the isValidAmvpMerge() condition (S1328), a threshold is referred to check the similarity between the candidate and the candidates previously added to the list. If they are not similar (S1334), the candidate is added to the candidate list (S1336). In this process, when the current block is in the amvpMerge mode (S1330), before performing the similarity check, in order to perform the similarity check using only the motion information of the merge reference list among the motion information of the candidates, the interDir of the candidate and the motion information of the AMVP candidate list are reset (S1332).

[0202] isValidAmvpMerge() is always true if the current block is not in amvpMerge mode, i.e., amvpMergeFlag is 0. It is true only when the current block is in amvpMerge mode and the reference picture indicated by the reference index of the candidate merge reference picture list is a reference picture available in amvpMerge mode, and false otherwise.

[0203] That is, isValidAmvpMerge() is true if the reference picture indicated by the reference index of the AMVP reference picture list and the reference picture indicated by the reference index of the merge candidate list of the merge candidate satisfy the value of amvpMergeValidRefPair induced at the slice level. Conversely, if the value of amvpMergeValidRefPair is false, isValidAmvpMerge() of the candidate is false.

[0204] The merge candidate list induced through the process of FIG. 13 is sorted in ascending order based on the bi-lateral Matching cost by the process shown in FIG. 14. numValidMergeCand indicates the number of candidates in the induced merge candidate list. That is, when DMVD is enabled (S1410) in the current block, the bi-lateral matching cost is calculated (S1412) for all the configured merge candidates, and the merge candidates are sorted in ascending order with reference to this value, and the candidate with the smallest error (cost) among the sorted merge candidates has the priority. In addition, if the number of induced merge candidates is not two or more (S1402), or when DMVD is disabled (S1410) in the current block, the cost of the merge candidates is set to the MAX value without calculating the bi-lateral cost (S1414, S1418), so as not to cause the effect of sorting, and the order of the induced merge candidates is maintained.

[0205] The AMVP candidate list and the merge candidate list derived through the processes described above can be used for the bidirectional predicted MV construction for the amvpMerge mode, and the process for the predicted MV construction is shown in FIG. 15.

[0206] In FIG. 15, MVPIdx indicates the AMVP candidate index derived explicitly or implicitly. numValidMergeCand indicates the number of merge candidates, and the curMv[] array and curRefIdx[] array indicate the motion information of the current block and the motion information and reference picture index for the reference picture list (REF_PIC_LIST_0, REF_PIC_LIST_1). For example, curMv[RefListMerge] indicates the motion information of the reference picture list indicated by RefListMerge, and curMv[RefListAmvp] indicates the motion information of the reference picture list represented by RefListAmvp. RefListMerge and RefListAmvp can be REF_PIC_LIST_0 or REF_PIC_LIST_1.

[0207] The MV[] array and refIdx[] array indicate the motion information of the candidate index of the AMVP candidate list or the merge candidate list, and the candidate index means the number specified in the array []. For example, mv[0] of the merge candidate list indicates the motion information of the 0th candidate configured in the merge candidate list.

[0208] Referring to FIG. 15, the reference picture list for the merge mode (merge reference picture list) is determined according to whether the prediction mode in the L0 direction is the merge mode, and the reference picture list in the direction opposite to the prediction direction to which the merge mode is applied is determined as the reference picture list for the AMVP mode (AMVP reference picture list) (S1502).

[0209] The motion information of the AMVP reference picture list is set to the motion information of the AMVP candidate indicated by the MVP candidate index among the AMVP candidates in the AMVP candidate list, and the reference picture index in the AMVP reference picture list is set (S1504). When the value of the MVP candidate index is 0 or 2 (S1506), and when the value of the MVP candidate index is 1 and (S1506) the number of merge candidates is 1 (S1508), the motion information of the merge reference picture list is set to the motion information of the 0th candidate among the merge candidates in the merge candidate list, and the reference picture index in the merge reference picture list is set to the index of the reference picture of the 0th candidate in the merge candidate list (S1510). When the value of the MVP candidate index is 1 and (S1506) the number of merge candidates is not 1 (S1508), the motion information of the merge reference picture list is set to the motion information of the 1st candidate among the merge candidates in the merge candidate list, and the reference picture index in the merge reference picture list is set to the index of the reference picture of the 1st candidate in the merge candidate list (S1510).

[0210] Example 1 - 6

[0211] Examples 1-6 propose a method for correcting (refinement) the motion information (predicted MV) induced in the process of inducing motion information using the amvpMerge mode in a coding unit or a prediction unit. The figures for Examples 1-6 are shown in FIG. 16.

[0212] Referring to FIG. 16, the merge reference picture list is determined according to whether the prediction mode in the L0 direction is the merge mode, the reference picture list in the direction opposite to the prediction direction to which the merge mode is applied is determined as the AMVP reference picture list, the POC of the reference picture indicated by the merge reference picture list and the index is determined, and the POC of the reference picture indicated by the AMVP complementary reference picture list and the index is determined (S1602).

[0213] In the case where DMVD is enabled (S1604), the motion information (predicted MV) in the amvpMerge mode is adaptively corrected based on Bi-lateral matching (S1608) or Template matching (S1610) considering the distance between the current picture and the reference picture (S1606).

[0214] Example 1 - 7

[0215] Examples 1-7 are methods of deriving motion information using the induced prediction candidates and the explicitly signaled MVD information in the process of deriving motion information using the amvpMerge mode in a coding unit or a prediction unit.

[0216] The processes for the proposed method are shown in FIGS. 6a and 6c.

[0217] Referring to FIGS. 6a to 6c, when the amvpMerge mode is not applied (S618), information indicating the prediction direction (S622), information indicating whether the Affine prediction mode is applied (S624), information indicating whether the smvd (symmetric MVD) mode is applied (S626), information on the bidirectional weight index (S628), information indicating a pair of reference pictures (S630), and information indicating the reference pictures in the reference picture list (S632) can be signaled.

[0218] If the amvpMerge mode is applied (S618), it is determined whether the prediction direction is the L1 prediction direction (S634), and the mode applied to each direction is determined between the AMVP mode and the merge mode (S636), and the reference picture index and the motion information predictor for the AMVP mode and the determined prediction direction are signaled (S638, S640).

[0219] If the affine mode is not applied (S642), it is determined whether the merge mode is applied in any direction and whether the value of the motion information predictor index in the opposite direction is less than 2 (S644). If it is true, the MVD is set to zero (S648). If it is false, the MVD in the opposite direction is signaled (S646). When the affine mode is applied (S642), the MVDs for the control points cp0, cp1, and cp2 are signaled (S650 - S654).

[0220] It is determined whether the prediction direction is the L0 prediction direction (S656), whether the smvd mode is applied (S658), and which mode is applied in each direction among the AMVP mode and the merge mode is determined (S660). The reference picture index and the motion information predictor for the AMVP mode and the determined prediction direction are signaled (S662, S664).

[0221] If the affine mode is not applied (S66), it is determined whether the merge mode is applied in any direction and whether the value of the motion information predictor index in the opposite direction is less than 2 (S668). If it is true, the MVD is set to zero (S672). If it is false, the MVD in the opposite direction is signaled (S670). When the affine mode is applied (S666), the MVDs for the control points cp0, cp1, and cp2 are signaled (S674 - S678).

[0222] Considering only the processes of S618, S636, S644, S646, S648, S660, S668, S670, and S672 among the processes represented in FIGS. 6a - 6c, if Examples 1 - 7 are further described, it is as follows.

[0223] When amvpMergeModeFlag[1] is true (1) (S644), the mvd for REF_PIC_LIST_0 is either explicitly signaled conditionally (S646) or set to zero MV without separate signaling (S648m). In this case, the condition means that the MVP index of REF_PIC_LIST_0 of the current block is greater than or equal to 2 (S644). Also, when amvpMergeModeFlag[1] is True, the MVD of REF_PIC_LIST_1 is set to zero MV (S672).

[0224] When amvpMergeModeFlag[0] is true (1) (S668), the mvd for REF_PIC_LIST_1 is either explicitly signaled conditionally (S670) or set to zero MV without separate signaling (S668). In this case, the condition means that the MVP index of REF_PIC_LIST_1 of the current block is greater than or equal to 2 (S668). Also, when amvpMergeModeFlag[0] is True, the MVD of REF_PIC_LIST_0 is set to zero MV (S648).

[0225] The final MV is derived by compensating the predicted MV for the MVD for each reference picture list induced through Examples 1-7.

[0226] Example 2

[0227] Example 2 is a modification of Examples 1-2 and proposes a method for constructing a reference picture list for the amvpMerge mode.

[0228] The figure for Example 2 is shown in FIG. 17. Since most of the processes represented in FIG. 17 operate in the same way as the processes represented in FIG. 4, the processes corresponding to the differences from FIG. 4 will be described below.

[0229] According to FIG. 4, after the reference picture list of the slice for which bLDCFlag is defined and to be encoded / decoded for the low delay slice is configured, if all the reference pictures in the reference picture list are in the forward direction, the value of bLDCFlag is set to true.

[0230] According to the process S1708 in FIG. 17, considering the information about the reference pictures in the reference picture list of the slice, if the current slice cannot configure a reference picture set that satisfies true bi-prediction, that is, if all the reference pictures are in the forward direction or the backward direction, a reference picture set available in the amvpMerge mode can be configured.

[0231] In Embodiment 2, a method for configuring a reference picture set available in the amvpMerge mode when the slice is a B slice that generates a prediction block using two or more motion information is also proposed.

[0232] An example of a method for configuring a reference picture set when the slice type is a B slice is shown in FIGS. 18a and 18b. Most of the processes represented in FIGS. 18a and 18b operate in the same way as the process represented in FIG. 5a. Therefore, below, the processes corresponding to the differences from FIG. 5a will be mainly described.

[0233] Referring to FIGS. 18a and 18b, in FIG. 5a, a process (S504) of examining whether the value of bLDCFlag is true is performed initially, but in FIGS. 18a and 18b, this process is not performed initially.

[0234] In addition, S528 to S536 in FIG. 5a are changed to the S1828 to S1844 process in FIGS. 18a and 18b. Specifically, the value of bLDCFlag is determined (S1828). If bLDCFlag = 1, a process of examining whether the proposed conditions are satisfied is performed (S1830). If bLDCFlag = 0, it is examined whether condition 1 is satisfied (S1832). Thereafter, the POC of the reference picture corresponding to refIdxLX is defined as pocLX, the POC of the reference picture corresponding to refIdxLY is defined as pocLY, and the POC of the current picture is defined as curPOC (S1834). Then, the value of bLDCFlag is determined again (S1836). If bLDCFlag = 1, it is examined whether condition 3 is satisfied (S1840). If bLDCFlag = 0, it is examined whether condition 2 is satisfied (S1838). When condition 3 is satisfied and when condition 2 is satisfied, the value of a variable (Valid) indicating whether a reference picture set available in the amvpMerge mode can be configured is set to true, and the pair of the reference picture indicated by refIdxLX and the reference picture indicated by refIdxLY is set to the reference picture set available in the amvpMerge mode (S1842). Thereafter, a determination for the next reference picture is made (S1844).

[0235] Condition 1 (condition 1, S1832) is a condition for the state of the reference picture, and can include whether the reference pictures corresponding to refIdxLXListX and refIdxInListY are encoded / decoded by reference picture resampling (RPR), and / or whether they are long-term reference pictures, and / or whether there is a weight for weighted prediction.

[0236] That is, as one example, condition 1 can be specified as a case where all or one or more of the following six conditions are satisfied.

[0237] The reference picture corresponding to refIdxInListX is not an RPR

[0238] The reference picture corresponding to refIdxInListY is not an RPR

[0239] The reference picture corresponding to refIdxInListX is not a long - term reference picture

[0240] The reference picture corresponding to refIdxInListY is not a long - term reference picture

[0241] The weight of weighted prediction for the reference picture corresponding to refIdxInListX is not signaled

[0242] The weight of weighted prediction for the reference picture corresponding to refIdxInListY is not signaled

[0243] Assuming that all of condition 1 is satisfied, it is determined whether the reference pictures corresponding to refIdxLX and refIdxLY are used as the reference picture set for the amvpMerge mode. If they cannot be used, the determination for the next reference picture is made.

[0244] Condition 2 (condition 2, S1838) indicates a condition for determining whether the reference pictures corresponding to refIdxLX and refIdxLY are suitable as reference pictures for the amvpMerge mode considering their POCs.

[0245] When the POC of the reference picture corresponding to refIdxLX is defined as pocLX, the POC of the reference picture corresponding to refIdxLY is defined as pocLY, and the POC of the current picture is defined as curPOC, if one reference picture has a POC smaller than that of the current picture and the other reference picture has a POC larger than that of the current picture, condition 2 can be satisfied. That is, condition 2 can be expressed as in the above formula 1.

[0246] Condition 3 (condition 3, S1840) indicates a condition for determining whether it is suitable as a reference picture for amvpMerge mode prediction in consideration of the POCs of the reference pictures corresponding to refIdxLX and refIdxLY when it is a low delay slice. Condition 3 is expressed as in formula 2 which means that when the POC of the reference picture corresponding to refIdxLX is defined as pocLX, the POC of the reference picture corresponding to refIdxLY is defined as pocLY, and the POC of the picture including the slice to be encoded / decoded is defined as curPOC, the two reference pictures are not the same picture. When the two pictures are not the same, it can be said that condition 3 is satisfied.

[0247]

Number

[0248] Condition 3 is expressed by formula 3 which means that the two reference pictures are reference pictures with a certain distance or more. When the two pictures have a distance difference of more than a predefined THRESHOLD, it can be said that condition 3 is satisfied.

[0249]

Number

[0250] abs() in formula 3 means the absolute value, and THRESHOLD is a predefined value that can be derived explicitly or implicitly by the decoder / encoder.

[0251] The conditions used in the process (S1830) of determining whether the proposed conditions are met can be defined as follows.

[0252] 1. How about useDMVD()

[0253] 2. State information of the reference picture

[0254] The state information of the reference picture can be specified when all or one or more of the following 6 conditions are met.

[0255] 1. The reference picture corresponding to refIdxInListX is not an RPR.

[0256] 2. The reference picture corresponding to refIdxInListY is not an RPR.

[0257] 3. The reference picture corresponding to refIdxInListX is not a long-term reference picture.

[0258] 4. The reference picture corresponding to refIdxInListY is not a long-term reference picture.

[0259] 5. The weight of weighted prediction for the reference picture corresponding to refIdxInListX is not signaled.

[0260] 6. The weight of weighted prediction for the reference picture corresponding to refIdxInListY is not signaled.

[0261] Whether the proposed conditions are met can be defined as the case where the state information of the reference picture is fully satisfied when useDMVD() is satisfied, or the case where the state information of the reference picture is partially satisfied when useDMVD() is satisfied.

[0262] Example 3

[0263] Example 3 is a modification of Examples 1-2 and Example 2, and proposes a method for constructing a reference picture list for the amvpMerge mode.

[0264] The figure for Example 3 is shown in FIG. 19. Since most of the processes represented in FIG. 19 operate in the same manner as the processes represented in FIG. 4, the processes corresponding to the differences from FIG. 4 will be described below.

[0265] According to FIG. 4, after the process of constructing a valid reference picture set for the amvpMerge mode (S416) is performed, a determination for the next slice is made. According to FIG. 19, after the process of constructing a valid reference picture set for the amvpMerge mode (S1916) is performed and the process of deriving a default reference picture for the amvpMerge mode (S1918) is performed, a determination for the next slice is made. Here, the S1918 process can correspond to the process of deriving the index of the default reference picture (default reference index) with reference to the reference picture set determined to be available in the S1916 process.

[0266] The process of deriving the default reference picture is performed through the processes represented in FIGS. 20a and 20b.

[0267] Referring to FIGS. 20a and 20b, when there is a reference picture set available in the current slice, that is, when bAmvpMergeEnabledFlag is true (1) (S2002), a process of deriving a default reference picture is performed. Various variables are defined (S2004), and it is determined whether the amvpMergeUniqValidRefIdx value of the reference picture list L0 is true (S2006). If it is true, the amvpMergeUniqValidRefIdx of the reference picture list L0 is set as the default reference picture in the L0 direction (S2008). If it is false, it is determined whether the determination has been made for all the reference pictures in the L0 direction (S2010). If there is a reference picture in the L0 direction for which the determination has not been made, it is determined whether the determination has been made for all the reference pictures in the L1 direction (S2012). If there is a reference picture in the L1 direction for which the determination has not been made, the proposed process is performed (S2014). The processes of S2010 to S2014 are repeatedly performed (S2016 to S2018), and the determination for all the reference pictures in the L0 direction and all the reference pictures in the L1 direction is completed.

[0268] After the reference picture indexes in the L1 direction and the L0 direction are initialized (S2020), it is determined whether the amvpMergeUniqValidRefIdx value of the reference picture list L1 is true (S2022). If it is true, the amvpMergeUniqValidRefIdx of the reference picture list L1 is set as the default reference picture in the L1 direction (S2024). If it is false, it is determined whether the determination has been made for all the reference pictures in the L1 direction (S2026). If there is a reference picture in the L0 direction for which the determination has not been made, it is determined whether the determination has been made for all the reference pictures in the L0 direction (S2028). If there is a reference picture in the L0 direction for which the determination has not been made, the proposed process is performed (S2030). The processes of S2026 to S2030 are repeatedly performed (S2032, S2034) until the determination for all the reference pictures in the L1 direction and all the reference pictures in the L0 direction is completed.

[0269] The proposed process means that the value of amvpMergeDefaultLX is derived when the following conditions are met. At this time, LX indicates REF_PIC_LIST_0 and REF_PIC_LIST_1.

[0270] 1. amvpMergeValidRefPair[refIdxInLX][refIdxInLY] == true

[0271] 2. abs(POC of reference picture indicated by refIdxInLX - POC of current picture) == abs(POC of reference picture indicated by refIdxInLY - POC of current picture)

[0272] 3.reference picture which has smallest value of abs(POC of reference picture indicated by refIdxInLX-POC of current picture)

[0273] 4.reference picture which has smallest value of abs(POC of reference picture indicated by refIdxInLY-POC of current picture)

[0274] The proposed process can define a reference picture that satisfies condition 2 as a default reference picture when condition 1 is met.

[0275] The proposed process can define a reference picture that satisfies condition 3 as the default reference picture in the LX direction when condition 1 is met.

[0276] The proposed process can define a reference picture that satisfies condition 4 as the default reference picture in the LX direction when condition 1 is met.

[0277] Example 3 - 4

[0278] Examples 3-4 are modifications of Examples 1-5, and propose a method for constructing prediction candidates in the process of deriving motion information via the amvpMerge mode in a coding unit or a prediction unit.

[0279] In the process of deriving a merge candidate list for amvpMerge mode prediction, motion prediction information can be derived by scaling TMVP prediction motion information with respect to a default reference picture index.

[0280] Also, in the process of deriving a merge candidate list for amvpMerge mode prediction, for the default reference picture index, the motion prediction information can be derived by scaling the HMVP and AffineHMVP prediction motion information.

[0281] Also, in the process of deriving a merge candidate list for amvpMerge mode prediction, for the default reference picture index, the motion prediction information can be derived by scaling the pairwise prediction motion information.

[0282] Also, in the process of deriving a merge candidate list for amvpMerge mode prediction, for the default reference picture index, the motion prediction information can be derived by scaling the fallback prediction motion information.

[0283] Example 5

[0284] Example 5 is a modification of Examples 1 - 5, and proposes a method for arranging a merge candidate list (merge candidates within the merge candidate list).

[0285] The figure for Example 5 is shown in FIG. 21. Since most of the processes represented in FIG. 21 operate in the same way as the processes represented in FIG. 14, the processes corresponding to the differences from FIG. 14 will be described below.

[0286] According to FIG. 14, when DMVD is enabled in the current block (S1410), only the process of calculating the bi - lateral matching cost for all the configured merge candidates (S1412) is performed. According to FIG. 21, when DMVD is enabled in the current block (S2112), depending on whether the proposed condition is satisfied (S2114), the process of calculating the template matching cost (S2116) or the process of calculating the bi - lateral matching cost (S2118) is performed.

[0287] The proposed conditions can be defined as follows.

[0288] 1. bLDCFlag == true

[0289] 2. The reference picture of the merge candidate and the reference picture of the AMVP candidate are not bidirectional.

[0290] 3. The reference picture of the merge candidate and the reference picture of the AMVP candidate are not true - bidirectional.

[0291] The proposed conditions can include one of the above three conditions. Also, the proposed conditions can include Condition 1 and Condition 2, or Condition 1 and Condition 3.

[0292] When bidirectional in Condition 2 and true - bidirectional in Condition 3 are defined with the POC of the reference picture of the merge candidate as pocMerge, the POC of the reference picture of the AMVP candidate as pocAmvp, and the POC of the current picture as curPoc, they can be defined as in the following Formulas 4 and 5.

[0293]

Equation

[0294]

Equation

[0295] Another figure for Example 5 is shown in FIG. 22. Since most of the processes represented in FIG. 22 operate in the same manner as the processes represented in FIG. 14, the processes corresponding to the differences from FIG. 14 will be described below.

[0296] Referring to FIG. 22, when bLDCFlag is false (S2210) and DMVD is enabled in the current block (S2214), the process of calculating the template matching cost (S2216) is performed. When bLDCFlag is true (S2210), that is, when all reference pictures of the current slice are backward or forward, the cost is defined as the maximum value without separate cost calculation (S2218). Such an operation substantially maintains the order of the induced merge candidate list because the sorting performance is not affected when all reference pictures of the current slice are backward or forward.

[0297] Example 6

[0298] Example 6 is a modification of Examples 1-5 and Examples 1-6, and proposes a method for constructing prediction candidates and a method for correcting motion information in the process of inducing motion information through the amvpMerge mode in a coding unit or a prediction unit.

[0299] An example for Example 6 is shown in FIG. 23. Since most of the processes represented in FIG. 23 operate in the same manner as the processes represented in FIG. 16, the processes corresponding to the differences from FIG. 16 will be described below.

[0300] According to FIG. 16, after the process of determining whether DMVD is enabled (S1604) is performed, the process of considering the distance between the current picture and the reference picture (S1606) is performed. According to FIG. 23, when DMVD is enabled (S2304), after determining whether the current picture exists between the reference picture of the merge candidate and the reference picture of the AMVP candidate based on the POC (S2306), the process of considering the distance between the current picture and the reference picture (S2308) is performed.

[0301] When the S2304 process and the S2306 process are combined, it is possible to determine whether the following bidirectional and true-bidirectional are satisfied. pocMerge indicates the POC of the reference picture of the merge candidate, pocAMVP indicates the POC of the reference picture of the amvp candidate, and curPoc indicates the POC of the current picture.

[0302] -bidirectional: (pocMerge - curPoc) × (pocAmvp - curPoc) < 0

[0303] -true-bidirectional: ((pocMerge - curPoc) × (pocAmvp - curPoc) < 0) && abs(pocMerge - curPoc) == abs(pocAmvp - curPoc)

[0304] In the case of a reference picture set that is not in a bidirectional relationship, correction is performed by the template matching method (S2310). When it satisfies bi-directional but is not in a true-bidirectional relationship, correction is performed by the template matching method (S2312).

[0305] Another example for Example 6 is shown in FIG. 24. Since most of the processes represented in FIG. 24 operate in the same way as the processes represented in FIG. 11, the processes corresponding to the differences from FIG. 11 will be described below.

[0306] According to FIG. 11, after the process (S1114) of deriving the motion vector predicted based on the merge candidate list is performed, the process (S1116) of correcting the predicted motion vector is performed. According to FIG. 24, the predicted motion vector is derived (S2414) based on the merge candidate list, and the value of bLDCFlag is determined (S2416). When bLDCFlag is false, the process (S2418) of correcting the predicted motion vector is performed, and when bLDCFlag is true, it ends.

[0307] According to Embodiment 6, when bLDCFlag is true, that is, when all reference pictures of the current slice are backward or forward, the process of correcting the motion vector can be skipped.

[0308] Example 7

[0309] Embodiment 7 proposes a method for deriving motion information for the amvpMerge mode only when it is a low delay picture.

[0310] For the prediction direction of deriving motion information in the AMVP prediction mode, among the processes of signaling the reference picture index, the AMVP candidate index, the MVD, etc., the AMVP candidate index can be implicitly derived without signaling for the AMVP candidate index.

[0311] The AMVP motion information can be derived as follows.

[0312] 1. Without signaling for the AMVP candidate index, after constructing a plurality of AMVP prediction candidates, align the AMVP prediction candidates based on the Decoder side template cost and derive the prediction candidate with the minimum cost as the AMVP motion information of the current block

[0313] 2. Construct a list with only one AMVP prediction candidate without signaling for the AMVP candidate index and navigate to the AMVP motion information of the current block.

[0314] In the process of signaling reference picture indexes, AMVP candidate indexes, MVDs, etc. for a prediction direction that induces motion information in AMVP prediction mode, the MVD can be set to zero MVD without signaling for the MVD. The setting to zero MVD can be performed when at least one of the following conditions is met.

[0315] 1. The AMVP prediction candidate list sorting process is performed based on the decoder side template cost.

[0316] 2. The motion information of amvpMerge mode based on the decoder side template cost will be refined.

[0317] In the case of a prediction direction that induces motion in the merge prediction mode, the merge prediction candidate with the 0th index in the prediction candidate list is used without additional signaling. Here, the merge prediction candidate list may be a prediction candidate list that has undergone the process of the fourth or fifth embodiment.

[0318] Encoding Method and Decoding Method

[0319] An image encoding method and a decoding method performed by an image encoding device 100 and an image decoding device 200 according to an embodiment will be described below.

[0320] 25, the image encoding device 100 and the image decoding device 200 may determine a prediction mode to be applied to each prediction direction of the current block (S2502). Here, the prediction mode may include an AMVP mode and a merge mode. For example, the prediction mode of one prediction direction among the prediction directions may be determined as the AMVP mode, and the prediction mode of the other prediction direction may be determined as the merge mode.

[0321] The image encoding device 100 and the image decoding device 200 can generate prediction blocks for each prediction direction based on at least one reference picture set, at least one AMVP candidate (AMVP prediction candidate), and at least one merge candidate (merge prediction candidate) (S2504). The image encoding device 100 and the image decoding device 200 can derive the (final) prediction block of the current block based on the prediction blocks generated in the S2504 process (S2506).

[0322] At least one reference picture set can be configured based on whether the POC of the reference picture of the current slice in which the current block is included is smaller or larger than the POC of the current picture in which the current block is included.

[0323] Also, at least one reference picture set can be configured when a predetermined first condition is satisfied. Here, the predetermined first condition can include at least one of whether DMVD is applied to the current block and the state information of the reference picture.

[0324] Also, at least one reference picture set can include a default reference picture based on at least one of the availability of a reference picture in one prediction direction and a reference picture in another prediction direction, and whether a predetermined second condition is satisfied. Here, the availability can mean whether the reference picture is a reference picture available in the amvpMerge mode. Also, the predetermined second condition can include at least one of whether the POC difference (first difference) between the current picture and a reference picture in one prediction direction is the smallest, and whether the POC difference (second difference) between the current picture and a reference picture in another prediction direction is the smallest. According to an embodiment, the predetermined second condition can further include whether the first difference and the second difference are the same as each other.

[0325] When the default reference picture is included in the reference picture set, the merge candidate can be scaled using the default reference picture, and the prediction blocks in other prediction directions (the prediction directions to which the merge mode is applied) can be derived based on the scaled merge candidate.

[0326] At least one merge candidate can be aligned based on whether the POC of the reference picture of the current slice is smaller or larger than the POC of the current picture. The alignment of the merge candidate can be performed based on either the template matching error or the bi-lateral matching error depending on whether a predetermined third condition is satisfied. The predetermined third condition can include whether the reference pictures of the merge candidate and the AMVP candidate have different prediction directions from each other. Also, the predetermined third condition can include whether the "POC difference between the reference picture of the merge candidate and the current picture" and the "POC difference between the reference picture of the AMVP candidate and the current picture" have the same absolute value.

[0327] The AMVP candidate and the merge candidate can be refined based on either the template matching error or the bi-lateral matching error depending on whether a predetermined fourth condition is satisfied. The fourth condition can include whether the reference pictures of the merge candidate and the AMVP candidate have different prediction directions from each other. Also, the predetermined fourth condition can include whether the "POC difference between the reference picture of the merge candidate and the current picture" and the "POC difference between the reference picture of the AMVP candidate and the current picture" have the same absolute value.

[0328] FIG. 26 is a diagram exemplarily showing a content streaming system to which an embodiment according to the present disclosure can be applied.

[0329] As shown in FIG. 26, a content streaming system to which an embodiment of the present disclosure is applied can generally include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.

[0330] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream, and transmits this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server can be omitted.

[0331] The bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0332] The streaming server transmits multimedia data to a user device based on a user's request via the Web server, and the Web server can serve as a medium for informing the user of what services are available. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server can play a role in controlling commands / responses between each device within the content streaming system.

[0333] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0334] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device, for example, a smartwatch, smart glass, an HMD (head mounted display), a digital TV, a desktop computer, a digital signage, and the like.

[0335] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received from each server can be processed distributively.

[0336] The scope of the present disclosure includes software or machine-executable commands (for example, an operating system, an application, firmware, a program, etc.) that enable operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium on which such software or commands are stored and can be executed on the device or the computer.

Industrial Applicability

[0337] Examples according to the present disclosure can be used to encode / decode images.

Claims

1. An image decoding method performed by an image decoding apparatus, comprising: determining a prediction mode applied to each of the prediction directions of the current block, wherein the prediction mode includes an Advanced Motion Vector Prediction (AMVP) mode and a merge mode; generating a prediction block for each of the prediction directions based on at least one reference picture set, at least one AMVP candidate, and at least one merge candidate; deriving a prediction block of the current block based on the prediction block; and wherein, among the prediction directions, the prediction mode of a certain prediction direction is determined as the AMVP mode, and the prediction mode of another prediction direction is determined as the merge mode.

2. The image decoding method according to claim 1, wherein the reference picture set is configured based on whether the Picture Order Count (POC) of the reference picture of the current slice in which the current block is included is smaller or larger than the POC of the current picture in which the current block is included.

3. The reference picture set is further configured based on satisfaction of a predetermined first condition, wherein the predetermined first condition includes whether Decoder Side Motion Vector Derivation (DMVD) is applied to the current block and at least one of the status information of the reference picture.

4. The image decoding method according to claim 3, wherein the reference picture set is further configured based on whether the reference picture in a certain prediction direction is the same as the reference picture in another prediction direction, or whether the POC difference between the reference picture in a certain prediction direction and the reference picture in another prediction direction is greater than or equal to a threshold.

5. The image decoding method according to claim 1, wherein the reference picture set includes a default reference picture based on at least one of the availability of the reference picture in a certain prediction direction and the reference picture in another prediction direction and satisfaction of a predetermined second condition.

6. The image decoding method according to claim 5, wherein the predetermined second condition includes at least one of whether the difference in POC between the current picture and the reference picture in a certain prediction direction is the smallest, and whether the difference in POC between the current picture and the reference picture in the other prediction direction is the smallest.

7. The image decoding method according to claim 5, wherein the predetermined second condition includes whether the first difference, which is the difference in POC between the current picture and the reference picture in a certain prediction direction, is the same as the second difference, which is the difference in POC between the current picture and the reference picture in the other prediction direction.

8. The merge candidate is scaled using the default reference picture, and the prediction block in the other prediction direction is derived based on the scaled merge candidate. The image decoding method according to claim 5.

9. The at least one merge candidate is aligned based on whether the POC of the reference picture of the current slice including the current block is smaller or larger than the POC of the current picture including the current block. The image decoding method according to claim 1.

10. The at least one merge candidate is aligned based on either a template matching error or a bilateral matching error based on whether a predetermined third condition is satisfied. The image decoding method according to claim 9.

11. The image decoding method according to claim 10, wherein the predetermined third condition includes whether the reference pictures of the merge candidate and the AMVP candidate have different prediction directions from each other.

12. The image decoding method according to claim 11, wherein the predetermined third condition further includes whether the difference in POC between the reference picture of the merge candidate and the current picture and the difference in POC between the reference picture of the AMVP candidate and the current picture have the same absolute value.

13. The AMVP candidate and the merge candidate are corrected based on either a template matching error or a bilateral matching error based on whether a predetermined fourth condition is satisfied, and the fourth condition includes whether the reference pictures of the merge candidate and the AMVP candidate have different prediction directions from each other. The image decoding method according to claim 1.

14. The image decoding method according to claim 13, wherein the predetermined fourth condition further includes whether the difference in POC between the reference picture of the merge candidate and the current picture and the difference in POC between the reference picture of the AMVP candidate and the current picture have the same absolute value.

15. An image encoding method performed by an image encoding apparatus, comprising: determining a prediction mode applied to each of the prediction directions of a current block, the prediction mode including an AMVP (advanced motion vector prediction) mode and a merge mode; generating a prediction block for each of the prediction directions based on at least one reference picture set, at least one AMVP candidate, and at least one merge candidate; deriving a prediction block of the current block based on the prediction block; wherein, among the prediction directions, the prediction mode of a certain prediction direction is determined as the AMVP mode, and the prediction mode of another prediction direction is determined as the merge mode.

16. A method for transmitting a bitstream generated by an image encoding method, the image encoding method comprising: determining a prediction mode applied to each of the prediction directions of a current block, the prediction mode including an AMVP (advanced motion vector prediction) mode and a merge mode; generating a prediction block for each of the prediction directions based on at least one reference picture set, at least one AMVP candidate, and at least one merge candidate; deriving a prediction block of the current block based on the prediction block; wherein, among the prediction directions, the prediction mode of a certain prediction direction is determined as the AMVP mode, and the prediction mode of another prediction direction is determined as the merge mode.

17. A computer-readable recording medium storing a bitstream generated by the image encoding method according to claim 15. ​