Image encoding / decoding method based on AMVPMERGE mode, method for transmitting bit stream, and recording medium storing bit stream

By adopting the amvpMerge mode reference picture list reordering method during image encoding/decoding, the problem of low high-resolution and high-quality image encoding/decoding efficiency is solved, and more efficient image information transmission and storage is achieved.

CN120476601APending Publication Date: 2025-08-12LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380090744.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-10
Filing Date
2023-11-06
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the encoding and decoding process of high resolution and high quality images, there is a problem of low encoding/decoding efficiency, especially when the amvpMerge mode is applied, the reordering efficiency of the reference picture list is insufficient, resulting in an increase in transmission and storage costs.

Method used

Using an image encoding/decoding method based on the amvpMerge pattern, the reference picture list is constructed and the prediction mode of the current block is reordered based on whether the prediction mode of the current block is an amvpMerge pattern, and the prediction mode of the advanced motion vector prediction (AMVP) pattern and the merge pattern are used to generate prediction samples of the current block, and avoid template matching reordering if necessary.

Benefits of technology

The efficiency of image encoding/decoding is improved, transmission and storage costs are reduced, and efficient image information processing is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476601A_ABST
    Figure CN120476601A_ABST
Patent Text Reader

Abstract

An image encoding / decoding method and apparatus are provided. An image decoding method according to the present disclosure comprises the steps of: constructing a reference picture list for a current block; rearranging the reference picture list based on whether a prediction mode of the current block is an amvpMerge mode, wherein an advanced motion vector prediction AMVP mode is applied to a first prediction direction of the current block and a merge mode is applied to a second prediction direction of the current block; and generating a prediction sample of the current block based on the reference picture list, in which a prediction mode based on the current block is an amvpMerge mode, and template matching-based rearrangement cannot be applied to the reference picture list.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for encoding / decoding an image, a method for transmitting a bitstream, and a recording medium for storing the bitstream, and more particularly, to an image encoding / decoding method and apparatus based on the amvpMerge mode, a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure, and a recording medium for storing the same. Background Art

[0002] In recent years, the demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD), has been growing in various fields. As image data becomes higher in resolution and higher in quality, the amount of information transmitted, or the bit rate, increases compared to conventional image data. This increase in the amount of information transmitted, or the bit rate, leads to increased transmission and storage costs.

[0003] Therefore, efficient image compression technology is needed to effectively transmit, store, and reproduce information of high-resolution and high-quality images. Summary of the Invention

[0004] Technical issues

[0005] The present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0006] In addition, the present disclosure provides an image encoding / decoding method and apparatus based on the amvpMerge mode.

[0007] In addition, the present disclosure is to provide an image encoding / decoding method and apparatus based on a reference picture list reordered according to whether an amvpMerge mode is applied.

[0008] Furthermore, the present disclosure is to provide a non-transitory computer-readable recording medium for storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0009] Furthermore, the present disclosure is to provide a non-transitory computer-readable recording medium for storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used for image reconstruction.

[0010] Furthermore, the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0011] The technical problems to be achieved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned can be clearly understood by those skilled in the art from the following description.

[0012] Technical Solution

[0013] According to one aspect of the present disclosure, an image decoding method includes: constructing a reference picture list for a current block; reordering the reference picture list based on whether a prediction mode of the current block is an amvpMerge mode, wherein the amvpMerge mode is a prediction mode in which an advanced motion vector prediction (AMVP) mode is applied to a first prediction direction of the current block and a merge mode is applied to a second prediction direction; and generating a prediction sample of the current block based on the reference picture list, wherein, based on the prediction mode of the current block being the amvpMerge mode, template matching-based reordering may not be applied to the reference picture list.

[0014] According to another aspect of the present disclosure, an image decoding device includes a memory and at least one processor, and the at least one processor constructs a reference picture list for a current block, reorders the reference picture list based on whether a prediction mode of the current block is an amvpMerge mode, wherein the amvpMerge mode is a prediction mode in which an advanced motion vector prediction (AMVP) mode is applied to a first prediction direction of the current block and a merge mode is applied to a second prediction direction, and generates a prediction sample of the current block based on the reference picture list, wherein, based on the prediction mode of the current block being the amvpMerge mode, template matching-based reordering may not be applied to the reference picture list.

[0015] According to another aspect of the present disclosure, an image encoding method includes: constructing a reference picture list for a current block; reordering the reference picture list based on whether a prediction mode of the current block is an amvpMerge mode, wherein the amvpMerge mode is a prediction mode in which an advanced motion vector prediction (AMVP) mode is applied to a first prediction direction of the current block and a merge mode is applied to a second prediction direction; generating a prediction sample of the current block based on the reference picture list; and encoding a reference picture index representing a reference picture used to generate the prediction sample into a bitstream, wherein, based on the prediction mode of the current block being the amvpMerge mode, template matching-based reordering may not be applied to the reference picture list.

[0016] A computer-readable recording medium according to another aspect of the present disclosure may store a bitstream generated by the image encoding method or the image encoding device of the present disclosure.

[0017] A transmission method according to another aspect of the present disclosure may transmit a bitstream generated by the image encoding method or the image encoding apparatus of the present disclosure.

[0018] The features briefly summarized above with respect to the present disclosure are merely exemplary aspects of the detailed description of the disclosure described below and do not limit the scope of the disclosure.

[0019] Beneficial effects

[0020] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0021] In addition, according to the present disclosure, an image encoding / decoding method and apparatus based on the amvpMerge mode can be provided.

[0022] In addition, according to the present disclosure, an image encoding / decoding method and apparatus based on a reference picture list reordered according to whether an amvpMerge mode is applied may be provided.

[0023] Furthermore, according to the present disclosure, a non-transitory computer-readable recording medium for storing a bitstream generated by the image encoding method or apparatus according to the present disclosure may be provided.

[0024] According to the present disclosure, a non-transitory computer-readable recording medium for storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used for image reconstruction may be provided.

[0025] Furthermore, according to the present disclosure, a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure may be provided.

[0026] Effects obtainable from the present disclosure are not limited to the above-described effects, and other effects that are not described can be clearly understood by those having ordinary skill in the art from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 A schematic diagram of a video encoding system to which embodiments of the present disclosure may be applied is shown.

[0028] Figure 2 A schematic diagram showing an image encoding device to which an embodiment of the present disclosure can be applied is shown.

[0029] Figure 3 A schematic diagram showing an image decoding device to which an embodiment of the present disclosure can be applied is shown.

[0030] Figure 4 is a diagram schematically showing an inter-frame predictor of an image encoding device.

[0031] Figure 5 is a flowchart representing a method for encoding an image based on inter-frame prediction.

[0032] Figure 6 is a diagram schematically showing an inter-frame predictor of an image decoding device.

[0033] Figure 7 is a flowchart representing a method for decoding an image based on inter-frame prediction.

[0034] Figure 8 is a flowchart showing an inter-frame prediction method.

[0035] Figure 9 It is a diagram used to describe the DMVR process.

[0036] Figure 10 A diagram for describing an encoding / decoding method based on template matching.

[0037] Figure 11 is a diagram illustrating decoder operation for amvpMerge mode according to an embodiment of the present disclosure.

[0038] Figure 12 and Figure 13 is a diagram showing a specific example of a single reference picture list.

[0039] Figure 14 is a diagram illustrating a decoder operation for amvpMerge mode according to another embodiment of the present disclosure.

[0040] Figures 15 to 17 is a diagram showing a specific example of the reference picture list reordering process in the amvpMerge mode.

[0041] Figure 18 is a diagram illustrating an MVP index parsing process according to an embodiment of the present disclosure.

[0042] Figure 19 is a diagram illustrating a decoder operation according to an embodiment of the present disclosure.

[0043] Figure 20 is a diagram illustrating an mvp index parsing method in amvpMerge mode according to an embodiment of the present disclosure.

[0044] Figure 21 is a diagram illustrating a method for reordering a single reference picture list based on whether RPR is performed.

[0045] Figure 22 is a diagram illustrating an mvp index parsing method in amvpMerge mode according to another embodiment of the present disclosure.

[0046] Figure 23 is a diagram illustrating an mvp index parsing method in amvpMerge mode according to another embodiment of the present disclosure.

[0047] Figure 24 is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure.

[0048] Figure 25 is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure.

[0049] Figure 26 An exemplary diagram showing a content streaming system to which embodiments of the present disclosure may be applied is shown. DETAILED DESCRIPTION

[0050] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. However, the present disclosure can be implemented in various forms and is not limited to the embodiments described herein.

[0051] When describing the embodiments of the present disclosure, when it is considered that a detailed description of a well-known configuration or function will obscure the main points of the present disclosure, its detailed description is omitted. Additionally, parts not related to the description of the present disclosure are omitted from the drawings, and similar reference numerals are assigned to similar components.

[0052] In the present disclosure, when a component is described as being "connected," "coupled," or "linked" to another component, this may include not only a direct connection but also an indirect connection with another component interposed therebetween. Additionally, when a component is described as "including" or "having" another component, unless explicitly stated otherwise, this does not exclude other components but may further include additional components.

[0053] In the present disclosure, the terms first, second, etc. are used only to distinguish one component from another component and do not limit the order or importance of the components unless otherwise explicitly stated. Therefore, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.

[0054] In this disclosure, distinguishable components are described for the purpose of clearly illustrating their respective characteristics and do not necessarily imply that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed across multiple hardware or software units. Therefore, such integrated or distributed implementations are also included within the scope of this disclosure, even if they are not explicitly described.

[0055] In the present disclosure, the components described in the various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also included in the scope of the present disclosure. Additionally, an embodiment including additional components in addition to the components described in the various embodiments is also included in the scope of the present disclosure.

[0056] The present disclosure relates to encoding and decoding of images, and unless otherwise defined in the present disclosure, terms used herein may have ordinary meanings commonly used in the technical field to which the present disclosure belongs.

[0057] In this disclosure, a "picture" generally refers to a unit representing a single image at a specific point in time. A slice / tile is a coding unit that constitutes a portion of a picture, and a picture may be composed of one or more slices / tiles. Additionally, a slice / tile may include one or more coding tree units (CTUs).

[0058] In the present disclosure, "pixel" or "picture element" may refer to the smallest unit constituting a picture (or image). Additionally, the term "sample" may be used as a corresponding term for a pixel. A sample may generally represent a pixel or a pixel value, and may indicate only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.

[0059] In the present disclosure, a "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific area of a picture or information related to the area. Depending on the context, the term "unit" may be used interchangeably with "sample array," "block," "area," and the like. Typically, an M×N block may include a set (or array) of samples (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.

[0060] In the present disclosure, the term "current block" may refer to one of a "current coding block," a "current coding unit," a "coding target block," a "decoding target block," or a "processing target block." When prediction is performed, the "current block" may refer to a "current prediction block" or a "prediction target block." When transform (inverse transform) / quantization (dequantization) is performed, the "current block" may refer to a "current transform block" or a "transform target block." When filtering is performed, the "current block" may refer to a "filtering target block."

[0061] In the present disclosure, unless explicitly stated as a chroma block, the term "current block" may refer to a block including both a luma component block and a chroma component block, or may refer to a "luma block of the current block." The luma component block of the current block may be explicitly expressed with terms such as "luma block" or "current luma block," which clearly indicates that it is a luma component block. Additionally, the chroma component block of the current block may be explicitly expressed with terms such as "chroma block" or "current chroma block," which clearly indicates that it is a chroma component block.

[0062] In the present disclosure, " / " and "," may refer to "and / or". For example, "A / B" and "A, B" may refer to "A and / or B". Additionally, "A / B / C" and "A, B, C" may refer to "at least one of A, B and / or C".

[0063] In the present disclosure, "or" may refer to "and / or". For example, "A or B" may mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" may also mean "additionally or alternatively".

[0064] Overview of Video Coding Systems

[0065] Figure 1 A schematic diagram of a video encoding system to which embodiments of the present disclosure may be applied is shown.

[0066] The video encoding system according to an embodiment may include an encoder device 10 and a decoder device 20. The encoder device 10 may transmit encoded video and / or image information or data to the decoder device 20 in the form of a file or stream via a digital storage medium or a network.

[0067] The encoder device 10 according to the embodiment may include a video source generator 11, an encoder 12, and a transmitter 13. The decoder device 20 according to the embodiment may include a receiver 21, a decoder 22, and a renderer 23. The encoder 12 may be referred to as a video / image encoder, and the decoder 22 may be referred to as a video / image decoder. The transmitter 13 may be included in the encoder 12. The receiver 21 may be included in the decoder 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0068] The video source generator 11 can obtain a video / image by capturing, synthesizing, or generating a video / image. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet, or a smartphone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc., and in this case, the video / image capture process may be replaced by a process that generates relevant data.

[0069] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoder 12 may output encoded data (encoded video / image information) in the form of a bitstream.

[0070] Transmitter 13 can obtain the encoded video / image information or data output in the form of a bitstream and transmit it in the form of a file or stream to receiver 21 of decoder device 20 or another external device via a digital storage medium or network. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 may include components for generating media files using a predetermined file format and components for transmission via a broadcast / communication network. Transmitter 13 may be provided as a transmission device separate from encoding device 10. In this case, the transmission device may include at least one processor for obtaining the encoded video / image information or data in the form of a bitstream and a transmitter for delivering the data in the form of a file or stream. Receiver 21 may extract / receive the bitstream from the storage medium or network and transmit it to decoder 22.

[0071] The decoder 22 may decode the video / image by performing a series of processes corresponding to the operations of the encoder 12 , such as dequantization, inverse transformation, prediction, and the like.

[0072] The renderer 23 may render the decoded video / image. The rendered video / image may be displayed via a display unit.

[0073] Overview of Image Coding Devices

[0074] Figure 2 A schematic diagram showing an image encoding device to which an embodiment of the present disclosure can be applied is shown.

[0075] like Figure 2 As described above, the image encoding apparatus 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as a "predictor." The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may further include a subtractor 115.

[0076] All or at least some of the components constituting the image encoding apparatus 100 may be implemented as a single hardware component (ie, an encoder or a processor) depending on the embodiment. Additionally, the memory 170 may include a decoded picture buffer (DPB) and may be implemented by a digital storage medium.

[0077] The image splitter 110 may split the input image (or picture, frame) input to the image encoding device 100 into at least one processing unit. For example, a processing unit may be referred to as a coding unit (CU). A coding unit may be obtained by recursively splitting a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree, binary tree, or ternary tree (QT / BT / TT) structure. For example, a coding unit may be divided into coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. To split a coding unit, a quadtree structure may be applied first, followed by a binary tree structure and / or a ternary tree structure. The encoding process according to the present disclosure may be performed based on a final coding unit that is not further split. A maximum coding unit may be used directly as the final coding unit, or a coding unit of greater depth obtained by splitting the maximum coding unit may be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and / or reconstruction, which will be described later. As another example, the processing unit used in the encoding process may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may each be divided or partitioned from the final coding unit. A prediction unit may be a unit for sample prediction, and a transform unit may be a unit for deriving a transform coefficient and / or a residual signal from the transform coefficient.

[0078] The predictor (inter-frame predictor 180 or intra-frame predictor 185) can perform prediction on the target block (current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block or coding unit (CU). The predictor can generate various information related to the prediction of the current block and send it to the entropy encoder 190. The prediction-related information can be encoded by the entropy encoder 190 and can be output in the form of a bitstream.

[0079] The intra-frame predictor 185 can predict the current block by referencing samples within the current picture. The referenced samples may be located in an adjacent area of the current block, or may be located at a farther position depending on the intra-frame prediction mode and / or intra-frame prediction method. The intra-frame prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional mode may include, for example, a DC mode and a planar mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the granularity of the prediction direction. However, this is an example, and depending on the configuration, a greater or lesser number of directional prediction modes may be used. The intra-frame predictor 185 may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0080] The inter-frame predictor 180 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. Motion information may include a motion vector and a reference picture index. Motion information may also include information regarding the inter-frame prediction direction (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a collocated reference block or collocated coding unit (colCU), and the reference picture including the temporally neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 180 may construct a motion information candidate list based on the neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 180 can use the motion information of the neighboring blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be sent. In motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may refer to the difference between the motion vector of the current block and the motion vector predictor.

[0081] The predictor can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the predictor can apply intra-frame prediction or inter-frame prediction for the prediction of the current block, and can also apply both intra-frame and inter-frame prediction simultaneously. A prediction method that applies both intra-frame and inter-frame prediction for the prediction of the current block is referred to as combined intra-frame inter-frame prediction (CIIP). Additionally, the predictor can perform intra-block copying (IBC) for the prediction of the current block. For example, intra-block copying can be used in applications such as picture content coding (SCC) in gaming content image / video coding. IBC is a method that predicts the current block using a pre-reconstructed reference block within the current picture that is located at a predetermined distance from the current block. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction within the current picture, but because it derives the reference block within the current picture, it can operate similarly to inter-frame prediction. In other words, IBC can use at least one of the inter-frame prediction methods described in this disclosure.

[0082] The prediction signal generated by the predictor can be used to generate a reconstructed signal or a residual signal. The subtractor 115 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the predictor from the input image signal (original block, original sample array). The generated residual signal can be sent to the transformer 120.

[0083] The transformer 120 may generate transform coefficients by applying a transform method to the residual signal. For example, the transform method may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. The transform process may be applied to pixel blocks of the same square size or non-square, variable-sized blocks.

[0084] The quantizer 130 may quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 130 may rearrange the block-shaped quantized transform coefficients into a one-dimensional vector based on a coefficient scanning order and may generate information about the quantized transform coefficients based on the one-dimensional vector of the quantized transform coefficients.

[0085] The entropy encoder 190 can perform various encoding methods, such as Exponential Golomb, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can not only encode the quantized transform coefficients, but also encode the information required for video / image reconstruction (i.e., the values of syntax elements) together with the quantized transform coefficients or separately. The encoded information (i.e., the encoded video / image information) can be transmitted or stored in the form of a bitstream in a network abstraction layer (NAL) unit. The video / image information may also include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The signaling information, transmitted information, and / or syntax elements described in this disclosure may be encoded through the above-mentioned encoding process and included in the bitstream.

[0086] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting a signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be provided as an internal / external element of the image encoding apparatus 100, or the transmitter may be configured as a component of the entropy encoder 190.

[0087] The quantized transform coefficients output from the quantizer 130 may be used to generate a residual signal. For example, the residual signal (residual block or residual sample) may be reconstructed by dequantizing and inverse-transforming the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.

[0088] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-frame predictor 180 or the intra-frame predictor 185. When the target block has no residual (such as when skip mode is applied), the prediction block can be used as a reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current picture, and as described later, after filtering, it can also be used for inter-frame prediction of the next picture.

[0089] The filter 160 may apply filtering to the reconstructed signal to enhance the subjective / objective quality. For example, the filter 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and the modified reconstructed picture may be stored in the memory 170, specifically in the DPB of the memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 may generate various filter-related information as described later in the description of each filtering method and send it to the entropy encoder 190. The filter-related information may be encoded by the entropy encoder 190 and output in the form of a bitstream.

[0090] The modified reconstructed picture transmitted to the memory 170 may be used as a reference picture in the inter-frame predictor 180. In this case, when applying inter-frame prediction, the image encoding apparatus 100 may avoid prediction mismatch between the image encoding apparatus 100 and the image decoding apparatus, and may improve encoding efficiency.

[0091] The DPB in the memory 170 can store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 180. The memory 170 can store motion information of blocks in the current picture for which motion information has been derived (or encoded) and / or motion information of blocks in the reconstructed image. The stored motion information can be sent to the inter-frame predictor 180 to be used as motion information for spatially adjacent blocks or temporally adjacent blocks. The memory 170 can store reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 185.

[0092] Overview of Image Decoding Equipment

[0093] Figure 3 A schematic diagram showing an image decoding device to which an embodiment of the present disclosure can be applied is shown.

[0094] like Figure 3 As shown, the image decoding apparatus 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame predictor 265. The inter-frame predictor 260 and the intra-frame predictor 265 may be collectively referred to as a "predictor." The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.

[0095] All or at least some of the plurality of components constituting the image decoding apparatus 200 may be implemented as a single hardware component (ie, a decoder or a processor) depending on an embodiment. Additionally, the memory 170 may include a DPB and may be implemented by a digital storage medium.

[0096] The image decoding apparatus 200 receiving a bit stream including video / image information may perform the same operation as that performed by Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image encoding device 100 in the image decoding device 200. For example, the image decoding device 200 can use the processing unit applied in the image encoding device to perform decoding. Therefore, the processing unit used for decoding can be, for example, a coding unit. The coding unit can be a coding tree unit, or can be obtained by dividing the maximum coding unit. In addition, the reconstructed image signal decoded and output by the image decoding device 200 can be played back by a playback device (not shown).

[0097] The image decoding device 200 may receive Figure 2The signal is output by the image encoding device in the form of a bitstream. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to extract information required for image reconstruction (or picture reconstruction) (i.e., video / image information). The video / image information may also include information about various parameter sets such as the adaptive parameter set (APS), the picture parameter set (PPS), the sequence parameter set (SPS), or the video parameter set (VPS). Additionally, the video / image information may also include general constraint information. The image decoding device may additionally use the information about the parameter set and / or the general constraint information to decode the image. The signaling information, reception information, and / or syntax elements described in the present disclosure can be obtained from the bitstream by decoding through a decoding process. For example, the entropy decoder 210 can decode the information in the bitstream based on a coding method such as Exponential Golomb, CAVLC, or CABAC, and can output syntax element values required for image reconstruction and quantized values of transform coefficients related to the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to a syntax element in a bitstream, can use information about the decoded target syntax element, decoded information about the decoded target block and adjacent blocks, or information about previously decoded symbols / bins to determine a context model, can predict the probability of a bin appearing based on the determined context model, and can perform arithmetic decoding of the bin to generate a symbol corresponding to each syntax element. In this case, the CABAC entropy decoding method can use the decoded symbol / bin information to update the context model for the next symbol / bin after determining the context model. Among the decoded information from the entropy decoder 210, prediction-related information can be provided to the predictor (inter-frame predictor 260 and intra-frame predictor 265), and the residual value entropy-decoded by the entropy decoder 210 (in other words, quantized transform coefficients and related parameter information) can be input to the dequantizer 220. Additionally, among the decoded information from the entropy decoder 210, filtering-related information can be provided to the filter 240. In addition, a receiver (not shown) that receives a signal output from the image encoding apparatus may be additionally configured as an internal / external element of the image decoding apparatus 200 , or the receiver may be configured as a component of the entropy decoder 210 .

[0098] In addition, the image decoding device according to the present disclosure may also be referred to as a video / image / picture decoding device. The image decoding device may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210, and the sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or an intra-frame predictor 265.

[0099] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scanning order used in the image encoding apparatus. The dequantizer 220 may dequantize the quantized transform coefficients using a quantization parameter (i.e., quantization step size information) and obtain the transform coefficients.

[0100] The inverse transformer 230 may perform inverse transform on the transform coefficients to obtain a residual signal (a residual block or a residual sample array).

[0101] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the prediction-related information output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction method).

[0102] The predictor can generate a prediction signal based on various prediction methods (techniques) to be described later, which is the same as the description of the predictor in the image encoding device 100 .

[0103] The intra predictor 265 may predict the current block by referring to samples within the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265 in the same manner.

[0104] The inter-frame predictor 260 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. Motion information may include a motion vector and a reference picture index. Motion information may also include information regarding the inter-frame prediction direction (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 260 may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes (methods), and prediction-related information may include information indicating the inter-frame prediction mode (method) applied to the current block.

[0105] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 260 and / or the intra-frame predictor 265). When the target block has no residual (such as when skip mode is applied), the prediction block can be used as the reconstructed block. The description of the adder 155 can also be applied to the adder 235 in the same manner. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current picture, and as described later, after filtering, it can also be used for inter-frame prediction of the next picture.

[0106] The filter 240 may apply filtering to the reconstructed signal to enhance the subjective / objective quality. For example, the filter 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and the modified reconstructed picture may be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0107] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-frame predictor 260. The memory 250 can store motion information of blocks in the current picture for which motion information has been derived (or decoded) and / or motion information of blocks in the reconstructed image. The stored motion information can be sent to the inter-frame predictor 260 for use as motion information of spatially or temporally neighboring blocks. The memory 250 can store reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 265.

[0108] In this specification, the embodiments described for the filter 160, the inter-frame predictor 180, and the intra-frame predictor 185 of the image encoding device 100 may be applied to the filter 240, the inter-frame predictor 260, and the intra-frame predictor 265 of the image decoding device 200 in the same or corresponding manner.

[0109] Inter-frame prediction

[0110] The predictor of the image encoding device 100 and the image decoding device 200 can perform inter-frame prediction on a block-by-block basis to derive prediction samples. Inter-frame prediction can be prediction derived based on data elements (e.g., sample values or motion information) from pictures other than the current picture. When inter-frame prediction is applied to the current block, the prediction block (prediction sample array) of the current block can be derived based on a reference block (reference sample array) specified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information on the inter-frame prediction type (L0 prediction, L1 prediction, bidirectional prediction, etc.). When inter-frame prediction is applied, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in a reference picture. The reference picture comprising the reference block and the reference picture comprising the temporal neighboring blocks can be the same or different. Temporally neighboring blocks may be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including temporally neighboring blocks may be referred to as a collocated picture (colPic). For example, a motion information candidate list may be constructed based on neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter-frame prediction may be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block may be the same as that of the selected neighboring block. For skip mode, unlike merge mode, a residual signal may not be transmitted. For motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block may be derived by using the sum of the motion vector predictor and the motion vector difference.

[0111] Motion information may include L0 motion information and / or L1 motion information depending on the inter-frame prediction type (L0 prediction, L1 prediction, bidirectional prediction, etc.). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on the L0 motion vector may be referred to as L0 prediction, prediction based on the L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 and L1 motion vectors may be referred to as bi-prediction. Here, the L0 motion vector may represent a motion vector associated with reference picture list L0 (L0), and the L1 motion vector may represent a motion vector associated with reference picture list L1 (L1). Reference picture list L0 may include pictures preceding the current picture in output order as reference pictures, and reference picture list L1 may include pictures following the current picture in output order as reference pictures. The preceding picture may be referred to as a forward (reference) picture, and the subsequent picture may be referred to as a backward (reference) picture. Reference picture list L0 may also include pictures following the current picture in output order as reference pictures. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed later. Reference picture list L1 may also include pictures preceding the current picture in output order as reference pictures. In this case, the subsequent picture may be indexed first in the reference picture list L1, and the previous picture may be indexed later. Here, the output order may correspond to the picture order count (POC) order.

[0112] Figure 4 is a diagram schematically showing the inter-frame predictor 180 of the image encoding device 100, and Figure 5 is a flowchart representing a method for encoding an image based on inter-frame prediction.

[0113] The image encoding device 100 may perform inter-frame prediction on the current block S510. The image encoding device 100 may derive an inter-frame prediction mode and motion information of the current block, and generate prediction samples of the current block. Here, the processes of determining the inter-frame prediction mode, deriving motion information, and generating prediction samples may be performed simultaneously, or any one process may be performed before the other process. For example, the inter-frame predictor 180 of the image encoding device 100 may include a prediction mode determiner 181, a motion information deriver 182, and a prediction sample deriver 183, and may determine the prediction mode of the current block in the prediction mode determiner 181, derive motion information of the current block in the motion information deriver 182, and derive prediction samples of the current block in the prediction sample deriver 183. For example, the inter-frame predictor 180 of the image encoding device 100 may search for a block similar to the current block in a specific area (search range) of a reference picture through motion estimation, and derive a reference block having the smallest difference from the current block or less than or equal to a certain standard. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the position difference between the reference block and the current block. The image encoding apparatus 100 can determine a mode to be applied to the current block from among various prediction modes. The image encoding apparatus 100 can compare the RD costs of various prediction modes and determine the optimal prediction mode for the current block.

[0114] For example, when skip mode or merge mode is applied to the current block, the image encoding apparatus 100 may construct a merge candidate list as described below and derive a reference block whose difference from the current block is minimal or less than or equal to a certain standard among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding device. The motion information of the current block may be derived using the motion information of the selected merge candidate.

[0115] As another example, when the (A)MVP mode is applied to the current block, the image encoding device 100 may construct the (A)MVP candidate list described below and use the motion vector of a motion vector predictor (MVP) candidate selected from among the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP for the current block. In this case, for example, the motion vector indicating the reference block derived through the above-described motion estimation may be used as the motion vector of the current block, and the MVP candidate with the smallest motion vector difference from the motion vector of the current block among the MVP candidates may be selected as the MVP candidate. A motion vector difference (MVD) may be derived as the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information regarding the MVD may be signaled to the image decoding device 200. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index may be constructed as reference picture index information and separately signaled to the image decoding device 200.

[0116] The image encoding apparatus 100 may induce residual samples based on the prediction samples S520. The image encoding apparatus 100 may induce residual samples by comparing original samples of the current block with the prediction samples.

[0117] The image encoding device 100 may encode image information including prediction information and residual information S530. The image encoding device 100 may output the encoded image information in the form of a bitstream. Prediction information is information related to the prediction process and may include information about prediction mode information (e.g., a skip flag, a merge flag, or a mode index, etc.) and motion information. The information about the motion information may include candidate selection information (e.g., a merge index, an mvp flag, or an mvp index), which is information used to derive a motion vector. In addition, the information about the motion information may include information about the above-mentioned MVD and / or reference picture index information. In addition, the information about the motion information may include information indicating whether L0 prediction, L1 prediction, or bidirectional prediction is applied. The residual information is information about the residual samples. The residual information may include information about the quantized transform coefficients of the residual samples.

[0118] The output bit stream may be stored in a (digital) storage medium and transmitted to a decoding device, or may be transmitted to the image decoding apparatus 200 via a network.

[0119] Furthermore, as described above, the image encoding apparatus 100 may generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to derive the same result from the image encoding apparatus 100 as the prediction result performed by the image decoding apparatus 200, which can improve encoding efficiency. Therefore, the image encoding apparatus 100 may store the reconstructed picture (or reconstructed sample, reconstructed block) in a memory and use it as a reference picture for inter-frame prediction. As described above, a loop filtering process, etc. may be further applied to the reconstructed picture.

[0120] Figure 6 is a diagram schematically showing the inter-frame predictor 260 of the image decoding apparatus 200, and Figure 7 is a flowchart representing a method for decoding an image based on inter-frame prediction.

[0121] The image decoding apparatus 200 may perform an operation corresponding to the operation performed in the image encoding apparatus 100. The image decoding apparatus 200 may perform prediction on the current block based on the received prediction information and derive a prediction sample.

[0122] Specifically, the image decoding apparatus 200 may determine a prediction mode of the current block based on the received prediction information S710. The image decoding apparatus 200 may determine which inter prediction mode to apply to the current block based on prediction mode information within the prediction information.

[0123] For example, whether merge mode or (A)MVP mode is applied to the current block may be determined based on the merge flag. Alternatively, one of various inter-frame prediction mode candidates may be selected based on the mode index. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter-frame prediction modes described below.

[0124] The image decoding apparatus 200 may derive motion information of the current block based on the determined inter-frame prediction mode (S720). For example, when skip mode or merge mode is applied to the current block, the image decoding apparatus 200 may construct a merge candidate list described below and select one of the merge candidates included in the merge candidate list. The selection may be performed based on the above-mentioned selection information (merge index). The motion information of the current block may be derived by using the motion information of the selected merge candidate. The motion information of the selected merge candidate may be used as the motion information of the current block.

[0125] As another example, when the (A)MVP mode is applied to the current block, the image decoding apparatus 200 may construct the (A)MVP candidate list described below and use the motion vector of the motion vector predictor (MVP) candidate selected from among the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the above-mentioned selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on the information about the MVD, and the motion vector of the current block may be derived based on the MVP and MVD of the current block. In addition, the reference picture index of the current block may be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list of the current block may be derived as the reference picture used for inter-frame prediction of the current block.

[0126] In addition, as described below, the motion information of the current block can be derived without constructing a candidate list, and in this case, the motion information of the current block can be derived according to the process initiated in the prediction mode described below. In this case, the candidate list construction as described above can be omitted.

[0127] The image decoding apparatus 200 may generate prediction samples of the current block based on the motion information of the current block (S730). In this case, a reference picture may be derived based on a reference picture index of the current block, and the prediction samples of the current block may be derived by using samples of the reference block indicated by the motion vector of the current block on the reference picture. In this case, as described below, in some cases, a prediction sample filtering process may be further performed on all or part of the prediction samples of the current block.

[0128] For example, the inter-frame predictor 260 of the image decoding device 200 may include a prediction mode determiner 261, a motion information deriver 262 and a prediction sample deriver 263, and may determine the prediction mode of the current block based on the received prediction mode information in the prediction mode determiner 181, derive the motion information (motion vector and / or reference picture index, etc.) of the current block based on the received information about the motion information in the motion information deriver 182, and derive the prediction sample of the current block in the prediction sample deriver 183.

[0129] The image decoding apparatus 200 may generate residual samples of the current block based on the received residual information (S740). The image decoding apparatus 200 may generate reconstructed samples of the current block based on the predicted samples and the residual samples, and generate a reconstructed picture based thereon (S750). Hereinafter, as described above, an in-loop filtering process, etc. may be further applied to the reconstructed picture.

[0130] Reference Figure 8 As described above, the inter-frame prediction process may include an inter-frame prediction mode determination step S810, a motion information derivation step S820 according to the determined prediction mode, and a prediction execution (prediction sample generation) step S830 based on the derived motion information. The inter-frame prediction process may be performed in the image encoding apparatus 100 and the image decoding apparatus 200 as described above.

[0131] Determine inter prediction mode

[0132] A variety of inter-frame prediction modes can be used for the prediction of the current block in the picture. For example, various modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode, etc. can be used. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bidirectional prediction with CU-level weights (BCW), bidirectional optical flow (BDOF), etc. can be used as additional modes or alternative modes. Affine mode can also be called affine motion prediction mode. MVP mode can also be called advanced motion vector prediction (AMVP) mode. In this document, some modes and / or motion information candidates derived from some modes can be included as one of the motion information-related candidates of other modes. For example, an HMVP candidate can be added as a merge candidate for merge / skip mode, or can be added as an MVP candidate for MVP mode.

[0133] Prediction mode information indicating the inter-frame prediction mode of the current block can be signaled from the image encoding device 100 to the image decoding device 200. The prediction mode information can be included in the bitstream and received by the image decoding device 200. The prediction mode information can include index information indicating one of multiple candidate modes. Alternatively, the inter-frame prediction mode can be indicated by hierarchical signaling of flag information. In this case, the prediction mode information can include at least one flag. For example, whether skip mode is applied can be indicated by signaling a skip flag, and when skip mode is not applied, whether merge mode is applied can be indicated by signaling a merge flag, and when merge mode is not applied, MVP mode can be indicated or a flag for additional distinction can be further signaled. Affine mode can be signaled as an independent mode, or can be signaled as a related mode such as merge mode or MVP mode. For example, affine mode can include affine merge mode and affine MVP mode.

[0134] DMVR (Decoder-side Motion Vector Optimization)

[0135] When DMVR is applied, the decoder can derive optimized motion information through cost comparison based on a template generated by using motion information of adjacent blocks in merge / skip mode. According to an embodiment of the present invention, the accuracy of motion prediction can be increased and compression performance can be improved without using additional signaling information.

[0136] Hereinafter, for convenience of description, the description will be based on the decoder side, but the DMVR of the present disclosure can be performed in the same manner as on the encoder side. In addition, DMVR can be regarded as an example of a process for deriving motion information (in which motion vectors are optimized), or as an example of a process for generating prediction samples (in which prediction samples are generated based on a reference block indicated by an optimized MV pair).

[0137] In the bidirectional prediction process, the optimized MV can be searched based on the initial MVs in the reference picture list L0 and the reference picture list L1. According to the bilateral matching (BM) method, the distortion between two candidate blocks in the reference picture list L0 and the list L1 can be calculated. Figure 9 , the SAD between red blocks (indicated by MVdiff) can be calculated based on each MV candidate within the search range centered on the initial MV. In this case, the MV candidate with the lowest SAD can be the optimized MV and can be used to generate a bidirectional prediction signal.

[0138] In an embodiment, the decoder may invoke the DMVR process to improve the accuracy of the initial motion compensated prediction (i.e., motion compensated prediction via conventional merge / skip modes). For example, the decoder may perform the DMVR process when the prediction mode of the current block is merge mode or skip mode and bidirectional bi-prediction (where bidirectional reference pictures are in opposite directions based on the current picture in display order) is applied to the current block.

[0139] Specifically, for example, DMVR can be applied to CUs encoded with the following modes and features:

[0140] CU-level merge mode with bi-directional prediction MV

[0141] One reference picture precedes the current picture and another reference picture follows the current picture

[0142] The distances from the two reference pictures to the current picture (i.e., POC difference) are the same

[0143] CU has 64 luma samples

[0144] CU height and CU width are both greater than or equal to 8 luma samples

[0145] BCW weight index indicates equal weight

[0146] WP is not enabled for the current block

[0147] The optimized MV derived by the DMVR process can not only be used to generate inter-frame prediction samples, but also to predict the temporal motion vector for future picture coding. The original MV can not only be used in the deblocking process, but also to predict the spatial motion vector for future CU coding.

[0148] AMVP with merge (AMVP-merge)

[0149] In AMVP-merge mode, the bidirectional predictor can be composed of an AMVP predictor in one direction and a merge predictor in the other direction. This mode can be enabled for the coding block when the selected merge predictor and AMVP predictor meet the predetermined DMVR condition. Here, the DMVR condition can be based on the situation where there is at least one past reference picture and a future reference picture for the current picture and the (time) distance between each reference picture and the current picture is the same. In this case, as a starting point, bilateral matching MV optimization can be applied to the merge MV candidate and AMVP MVP. Conversely, when the template matching function is enabled, template matching MV optimization can be applied to the merge predictor or AMVP predictor with a higher template matching cost.

[0150] The AMVP portion of the pattern may be signaled as regular unidirectional AMVP (ie, reference index and MVD) and may have a derived mvp index when signaling the mvp index for cases where template matching is used or disabled.

[0151] For the AMVP direction LX (where X can be 0 or 1), the merge portion in the other direction (1-LX) can be implicitly derived by minimizing the bilateral matching cost between the AMVP predictor and the merge predictor (i.e., the AMVP and merge motion vector pair). For all merge candidates in the merge candidate list with a motion vector in the other direction (1-LX), the bilateral matching cost can be calculated by using the merge candidate MV and the AMVP MV. In addition, the merge candidate with the lowest cost can be selected. Starting from the selected merge candidate MV and AMVP MV, the bilateral matching optimization can be applied to the coding block.

[0152] The third pass of multi-pass DMVR (which is an 8x8 sub-PU BDOF optimization of multi-pass DMVR) may be enabled for blocks encoded in AMVP-merge mode.

[0153] The above-mentioned AMVP-merge mode can be indicated by a flag, and when this mode is enabled, LX in the AMVP direction can be additionally indicated by a flag.

[0154] When bilateral matching (BM) AMVP-merge mode is used for the current block and template matching is enabled, MVD may not be signaled. An additional pair of AMVP-merge MVP is introduced. The merge candidate list may be sorted in ascending order of BM cost. To indicate the merge candidate to be used in the sorted merge candidate list, an index (0 or 1) may be signaled. When there is only one candidate in the merge candidate list, the AMVP MVP and merge MVP pair may be populated without bilateral matching MV optimization.

[0155] Block-level reference picture list reordering

[0156] A block-level reference picture reordering method based on template matching can be used. For unidirectional AMVP mode, the reference pictures in List 0 and List 1 can be mixed to generate a joint list. For each hypothesis of a reference picture in the joint list, template matching can be performed to calculate the cost. The joint list can be reordered based on ascending template matching cost. In addition, the index of the reference picture selected in the reordered joint list can be signaled in the bitstream. For bidirectional prediction AMVP mode, a list of pairs of reference pictures from List 0 and List 1 can be generated and similarly reordered based on template matching cost. In addition, the index of the selected pair can be signaled.

[0157] Reference picture resampling

[0158] Reference picture resampling is inherited from VVC. Compared to the filter length of VVC (e.g., 8 taps for luma affine coded blocks, 6 taps for luma non-affine coded blocks, and 4 taps for each chroma block), the corresponding RPR filters in ECM can be increased to 12, 10, and 6 taps.

[0159] Local Illumination Compensation (LIC)

[0160] LIC is an inter-frame prediction technique that models the local illumination variation between the current block and its predicted block as a function of the current block template and the reference block template. In this case, the parameters of the function can be expressed as a scale α and an offset β constituting the linear equation α*p[x]+β (where p[x] is the reference sample indicated by the MV at position x of the reference picture). When surround motion compensation is enabled, the MV must be cropped to account for the surround offset. Since the parameters α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for the parameters, other than signaling the LIC flag indicating the use of LIC for AMVP mode.

[0161] LIC is used for unidirectionally predicted inter-frame coding units and may have the following characteristics:

[0162] - Intra-frame neighboring samples can be used in the process of deriving LIC parameters.

[0163] - Disable LIC for blocks smaller than 32 luma samples.

[0164] - For non-subblock and affine modes, LIC parameters are derived based on the template block samples corresponding to the current CU, rather than the partial template block samples corresponding to the first top-left 16x16 unit.

[0165] Furthermore, samples of the reference block template may be generated by using motion compensation and block MV together without rounding to integer pixel precision.

[0166] Template Matching(TM)

[0167] Figure 10 This is a diagram for describing an encoding / decoding method based on template matching.

[0168] Template matching (TM) is a method for deriving motion vectors performed on the decoder side. It is a method that refines the motion information of the current block (e.g., the current coding unit, the current CU) by searching for a template (hereinafter referred to as a "reference template") that is most similar to a template adjacent to the current block (e.g., the current coding unit, the current CU) in a reference picture. The current template may be the upper and / or left neighboring blocks of the current block, or a portion of these neighboring blocks. In addition, the reference template may be determined to have the same size as the current template.

[0169] like Figure 10 As shown, once an initial motion vector for the current block is derived, a search for a better motion vector can be performed in the neighboring region of the initial motion vector. For example, the search can be performed within a search range of [-8, +8] pixels based on the initial motion vector. Furthermore, the search step size used for the search can be determined based on the AMVR mode of the current block. Furthermore, template matching can be performed continuously with the bilateral matching process in merge mode.

[0170] When the prediction mode of the current block is AMVP mode, a motion vector predictor candidate (MVP candidate) may be determined based on the template matching error. For example, a motion vector predictor candidate (MVP candidate) that minimizes the error between the current template and the reference template may be selected. Template matching for optimizing the motion vector may then be performed on the selected motion vector predictor candidate. In this case, template matching for optimizing the motion vector may not be performed on unselected motion vector predictor candidates.

[0171] More specifically, the optimization of the selected motion vector predictor candidate can be started from full pixel (integer pixel) accuracy within the [-8, +8] pixel search range using an iterative diamond search. Alternatively, for the 4-pixel AMVR mode, it can start from 4-pixel accuracy. Thereafter, a search for half pixel and / or quarter pixel accuracy can be performed according to the AMVR mode. Depending on the search process, the motion vector predictor candidate can maintain the same motion vector accuracy as indicated by the AMVR mode even after the template matching process. During the iterative search process, the search process terminates when the difference between the previous minimum cost and the current minimum cost is less than an arbitrary threshold. The threshold can be the same as the area of the block (i.e., the number of samples in the block). Table 1 shows an example of a search pattern according to the AMVR mode and the merge mode accompanying AMVR.

[0172] [Table 1]

[0173]

[0174] When the prediction mode of the current block is merge mode, a similar search method can be applied to the merge candidates indicated by the merge index. As shown in Table 1, template matching can be performed up to 1 / 8 pixel accuracy or can skip half-pixel accuracy or less, depending on whether an alternative interpolation filter is used based on the merge motion information. In this case, the alternative interpolation filter can be the filter used when AMVR is in half-pixel mode. In addition, when template matching is available, depending on whether bilateral matching (BM) is available, template matching can be run as a standalone process or as an additional motion vector optimization process between block-based bilateral matching and sub-block-based bilateral matching. Whether template matching is available and / or whether bilateral matching is available can be determined based on an availability condition check. In the above, the accuracy of the motion vector can refer to the accuracy of the motion vector difference (MVD).

[0175] Problems in traditional technologies

[0176] For the above amvpMerge mode, when the amvpMerge flag is sent and the value of the corresponding flag is 1, bidirectional prediction is performed without sending the inter_pred_idc value in the following table.

[0177] [Table 2]

[0178]

[0179] For the amvpMerge mode, prediction motion information based on the AMVP mode is derived by referring to the reference pictures in the reference picture list of one direction, and prediction motion information based on the merge mode is derived by referring to the reference pictures in the reference picture list of the opposite direction. To this end, when the reference picture index for block-level reference picture list reordering is signaled, an AMVP candidate referencing the reference picture of the corresponding list is derived based on whether the reference picture indicated by the reference picture index belongs to the L0 list or the L1 list, and prediction motion information is derived using the AMVP method.

[0180] In order to perform this process, block-level reference picture list reordering must have been performed in the parsing step, but pipeline delay issues may arise since it may be known to which reference picture list the reference picture indicated by the reference picture index belongs.

[0181] To address this pipeline delay issue, a method can be considered that performs block-level reference picture list reordering during the decoding process, rather than during the parsing step. However, in this case, it is unknown to which list, L0 or L1, the reference picture indicated by the reference picture index belongs. Furthermore, since it is unknown whether the corresponding reference picture is a resampled reference picture (RPR), the following problem may arise: parsing may not be performed again when the index of the AMVP candidate list is sent in a subsequent step. This problem occurs because, when sending the AMVP candidate index for amvpMerge, adaptive signaling is performed to indicate two candidates when template matching-based motion compensation of the amvpMerge prediction candidate is possible on the decoding side, and three candidates when motion compensation is not possible on the decoding side. Furthermore, when the referenced reference picture is an RPR, the template matching-based motion compensation process may not be performed. As a result, when it is unknown whether the currently referenced reference picture is an RPR, it is unknown whether to parse two or three candidates during the parsing step, and therefore the decoding side may not perform the parsing again.

[0182] Furthermore, when block-level reference picture list reordering is performed in the decoding process rather than the parsing step, the reference picture list used for prediction in the AMVP mode of amvpMerge during the parsing step can be defined as L0 or L1 unidirectional prediction. In this case, if it is arbitrarily defined as L1 unidirectional prediction, the Symmetric Motion Vector Difference (SMVD) flag is transmitted even though the SMVD scheme may not be applied in amvpMerge mode. This leads to the following problem: unnecessary bits are transmitted even though the aforementioned Local Illumination Compensation (LIC) may not be used in amvpMerge mode. Furthermore, if it is arbitrarily defined as L0 unidirectional prediction to address this issue, the transmission of the SMVD flag can be prevented, but the problem of transmitting the LIC flag still cannot be resolved.

[0183] To address the above-mentioned issues, according to an embodiment of the present disclosure, a single reference picture list can be configured for amvpMerge mode, and the list can be skipped or reordered without pipeline delay during the parsing step. In addition, in amvpMerge mode, the maximum number of candidates can be adaptively determined based on whether template matching-based motion compensation is allowed at a higher level, and the cMax value in the truncated binarization process can be determined based on the determined maximum number of candidates. In addition, in amvpMerge mode, the LIC flag can be not parsed regardless of the prediction direction (i.e., the value of inter_pred_idc). In addition, in amvpMerge mode, the maximum number of candidates can be adaptively determined based on whether the reference picture is an RPR, and the binarized MVP index can be parsed using the determined maximum number of candidates. Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. The embodiments described below can be implemented individually or in combination of at least two.

[0184] Implementation Method 1

[0185] According to the first embodiment of the present disclosure, during the motion information prediction process in amvpMerge mode, the reference picture list referenced by the amvpMerge prediction candidate can be composed of a single list based on the L0 list and the L1 list. The reordering process for the single list can be skipped or limited to simple reordering. As a result, the reference picture information indicated by the reference picture index can be known without pipeline delay during the parsing step. In addition, whether reference picture resampling (RPR) is performed or the direction of the reference picture list (L0 or L1) can also be known.

[0186] Figure 11 This figure illustrates decoder operations for amvpMerge mode according to an embodiment of the present disclosure. Decoder operations for amvpMerge mode can be divided into a parsing process and a decoding process. During the parsing process, a reference picture list can be configured based on whether amvpMerge mode is applied. Furthermore, during the decoding process, the reference picture list can be reordered based on whether amvpMerge mode is applied.

[0187] Reference Figure 11 During the parsing process, the image decoding apparatus may parse the amvpMerge flag from the bitstream S1110.

[0188] When the amvpMerge flag is 0 or false (No in S1120 ), the amvpMerge mode may not be applied, and the image decoding apparatus may configure a single reference picture list for the AMVP mode S1130 .

[0189] On the contrary, when the amvpMerge flag is 1 or true ("Yes" in S1120), the amvpMerge mode can be applied, and the image decoding apparatus can configure a single reference picture list for the amvpMerge mode S1140. In this case, a single reference picture list can be configured by using the L0 list and the L1 list.

[0190] In addition, the image decoding apparatus may obtain a reference picture index S1150 indicating a reference picture to be used for motion information prediction within the reference picture list configured in S1130 or S1140 described above.

[0191] Next, during the decoding process, when the amvpMerge flag is 0 or false (No in S1160 ), the image decoding apparatus may reorder the single reference picture list for the AMVP mode S1170 . In an embodiment, the reordering process may be performed based on all template matching techniques.

[0192] Conversely, when the amvpMerge flag is 1 or true ("Yes" in S1160), reordering of a single reference picture list for amvpMerge mode can be skipped. By not reordering the order of pictures in the reference picture list (in units of CUs, PUs, or arbitrary units) in amvpMerge mode, the reference picture information indicated by the reference picture index obtained in the parsing step can be interpreted without pipeline delay. In other words, the reference picture information referenced by the reference picture index transmitted in the bitstream can be referenced during the encoding / decoding process of the candidate index transmitted at the next position of the reference picture index and the TM-based coding tool control flag.

[0193] In addition, since template matching-based reordering is not performed during decoding, index transmission efficiency can be determined based on the picture order of the reference picture list configured in the parsing step. Taking this into account, in amvpMerge mode, a single reference picture list is configured based on the L0 list and the L1 list. A specific example of a single reference picture list is as follows Figure 12 and Figure 13 As shown. Figure 12 and Figure 13 In , each number (eg, 0, 8, 32, 16, etc.) represents a picture order count (POC) value of the corresponding reference picture.

[0194] First, refer to Figure 12In the example of FIG. 1 , reference picture 0 of the first index in the L0 list is inserted into the first index position of the single list, and reference picture 4 of the first index in the L1 list is inserted into the second index position of the single list. Next, reference picture 8 of the second index in the L0 list is inserted into the third index position of the single list, and reference picture 32 of the second index in the L1 list is inserted into the fourth index position of the single list. Since reference picture 32 has already been inserted from the L1 list into the single list, reference picture 32 of the third index in the L0 list is discarded. In addition, reference picture 16 of the fourth index in the L0 list is inserted into the fifth index position of the single list, and reference picture 2 of the third index in the L1 list is inserted into the last index position of the single list, thereby completing the single list configuration.

[0195] Refer to below Figure 13 In the example of FIG, reference picture 0 of the first index in the L0 list is inserted into the first index position of the single list, and reference picture 4 of the first index in the L1 list is inserted into the second index position of the single list. Next, reference picture 8 of the second index in the L0 list is inserted into the third index position of the single list, and reference picture 32 of the second index in the L1 list is inserted into the fourth index position of the single list. Next, reference picture 2 of the third index in the L1 list is inserted into the fifth index position of the single list. Since reference picture 32 has already been inserted from the L1 list into the single list, reference picture 32 of the third index in the L0 list is discarded. In addition, reference picture 2 of the third index in the L1 list is inserted into the fifth index position of the single list, and reference picture 16 of the fourth index in the L0 list is inserted into the last index position of the single list, thereby completing the single list configuration.

[0196] In addition, as in Figure 11 In the method, the process of reordering the reference picture list in amvpMerge mode can be changed to be performed during the parsing process instead of being skipped completely. The specific details are as follows Figure 14 shown.

[0197] Figure 14 is a diagram showing the decoder operation for amvpMerge mode according to another embodiment of the present disclosure. Figure 11 In this way, the decoder operation for amvpMerge mode can be divided into a parsing process and a decoding process.

[0198] Reference Figure 14 During the parsing process, the image decoding apparatus may parse the amvpMerge flag from the bitstream S1410.

[0199] When the amvpMerge flag is 0 or false (No in S1420 ), the amvpMerge mode may not be applied, and the image decoding apparatus may configure a single reference picture list for the AMVP mode S1430 .

[0200] On the contrary, when the amvpMerge flag is 1 or true ("Yes" in S1420), the amvpMerge mode may be applied, and the image decoding device may configure a single reference picture list for the amvpMerge mode S1440. In addition, the image decoding device may perform a reordering process S1450 on the configured single reference picture list. In this case, the reordering process is not performed based on template matching, but may be performed only by using pre-decoded reference picture information such as picture order count (POC), temporal ID (TID), quantization parameter (QP), etc. In this regard, Figure 14 The decoder operation in Figure 11 different.

[0201] The image decoding apparatus may obtain a reference picture index S1460 indicating a reference picture to be used for motion information prediction within the reference picture list configured through S1430 or S1440 and S1450 .

[0202] Next, during the decoding process, when the amvpMerge flag is 0 or false ("No" in S1470), the image decoding apparatus may perform a reordering process S1480 for a single reference picture list for the AMVP mode. In this case, unlike the reordering process in S1450, the reordering process may be performed based on template matching.

[0203] In contrast, when the amvpMerge flag is 1 or true ("Yes" in S1470), the template matching-based reordering process for a single reference picture list of the amvpMerge mode may be skipped.

[0204] In addition, a specific example of the process for reordering the reference picture list of the amvpMerge mode performed in the parsing step is as follows: Figures 15 to 17 As shown. Figures 15 to 17 In , each number (e.g., 0, 4, 8, etc.) represents the picture order count (POC) value of the corresponding reference picture. Figures 15 to 17 In the example, it is assumed that the POC value of the current image is 1.

[0205] As in Figure 15In the example of

[0065] , a single list for amvpMerge mode may be reordered based on the POC value of each picture. For example, the single list may be reordered in the order of "reference picture 0 → reference picture 2 → reference picture 4 → reference picture 8 → reference picture 16 → reference picture 32" (i.e., in ascending order of POC).

[0206] Alternatively, as in Figure 16 In this example, a single list for amvpMerge mode can be reordered based on the QP value of each picture. For example, the single list can be reordered in the order of "reference picture 0 → reference picture 32 → reference picture 16 → reference picture 8 → reference picture 4 → reference picture 2" (i.e., in ascending order of QP).

[0207] Alternatively, as in Figure 17 In this example, a single list for amvpMerge mode can be reordered based on the temporal ID value of each picture. For example, the single list can be reordered in the order of "reference picture 0 → reference picture 32 → reference picture 16 → reference picture 8 → reference picture 4 → reference picture 2" (i.e., in ascending order of TID).

[0208] As described above, according to Embodiment 1 of the present disclosure, a single reference picture list for amvpMerge mode can be configured, and reordering of the list can be skipped or performed without pipeline delay in the parsing step.

[0209] Implementation Method 2

[0210] According to embodiment 2 of the present disclosure, in amvpMerge mode, the AMVP maximum candidate index value (or the maximum number of candidates) can be adaptively determined based on whether template matching-based motion compensation is allowed, as signaled in the high-level syntax. When template matching-based motion compensation is allowed in amvpMerge mode, the maximum number of allowed candidates can be determined to be the same regardless of whether specific conditions of the reference picture actually referenced (e.g., whether it is an RPR) are met.

[0211] Figure 18 is a diagram illustrating an MVP index parsing process according to an embodiment of the present disclosure.

[0212] Reference Figure 18 , the image decoding apparatus may parse the amvpMerge flag from the bitstream S1810 and obtain a reference picture index S1820.

[0213] When the amvpMerge flag is 1 or true ("Yes" in S1830), the amvpMerge mode is applied, so the image decoding device can determine whether to allow motion compensation based on template matching (condition *) S1840. Whether to allow can be determined at the sequence, picture, or slice level through high-level syntax (e.g., SPS, PPS, PH, SH, etc.). For example, when motion compensation based on template matching is not allowed at the sequence level ("Yes" in S1840, condition * = true), the maximum candidate index can be determined as a predefined first value (case 1, S1850), and the mvp index can be signaled based on the determined maximum candidate index. On the contrary, when motion compensation based on template matching is allowed at the sequence level ("No" in S1840, condition * = false), the maximum candidate index can be determined as a predefined second value (case 2, S1855), and the mvp index can be signaled based on the determined maximum candidate index. For example, when it is determined in the SPS that AMVP based on template matching is not performed, regardless of whether the actually referenced reference picture is reference picture resampling (RPR), binaryization of the mvp index can be applied to indicate up to three candidate indexes. Alternatively, when it is determined in the SPS that AMVP based on template matching is not performed, regardless of whether the actually referenced reference picture is RPR, binaryization of the mvp index can be applied to indicate up to two candidate indexes. In addition, the image decoding device can parse amvpMerge_mvp_index[][] based on the determined maximum candidate index value S1860. In the example, amvpMerge_mvp_index[][] can follow a truncated binaryization process. In this case, the input of the process can be a syntax element with a synVal value and a binaryization request for cMax, and the output of the process can be the binaryized value of the syntax element. The bin string of the syntax element can be represented by the following formula.

[0214] [Equation 1]

[0215] n = cMax + 1

[0216] k = Floor(Log2(n))

[0217] u = (1 << (k + 1)) - n

[0218] When synVal is less than the variable u in Equation 1 above, the bin string can be derived by performing a fixed-length (FL) binaryization process on synVal with cMax value equal to (1 << k) - 1. On the contrary, when synVal is greater than or equal to the variable u in Equation 1 above, the bin string can be derived by performing a fixed-length (FL) binaryization process on (synVal + u) with cMax value equal to (1 << (k + 1)) - 1.

[0219] When the amvpMerge flag is 0 or false (No in S1830), the amvpMerge mode is not applied, and the image decoding device may determine the maximum candidate index of the conventional AMVP mode as a predefined third value S1870. Then, the image decoding device may parse mvp_index_LX[][] based on the determined maximum candidate index value S1880.

[0220] As described above, according to Embodiment 2 of the present disclosure, in amvpMerge mode, the maximum number of candidates can be adaptively determined based on whether template matching-based motion compensation is allowed at a higher level. In addition, the cMax value in the truncated binarization process can be determined based on the determined maximum number of candidates.

[0221] Implementation 3

[0222] According to embodiment 3 of the present disclosure, when the amvpMerge mode is applied, parsing of the local illumination compensation (LIC) flag may be skipped regardless of the prediction direction (ie, the value of inter_pred_idc).

[0223] Figure 19 is a diagram illustrating a decoder operation according to an embodiment of the present disclosure.

[0224] Reference Figure 19 , the image decoding device can parse the amvpMerge flag S1910.

[0225] If the amvpMerge flag is 0 or false ("No" in S1920), the amvpMerge mode is not applied. Therefore, the image decoding apparatus may determine the signaling condition of the LIC flag in S1930. In an embodiment, the signaling condition of the LIC flag may be whether inter_pred_idc is defined as bidirectional prediction. For example, if inter_pred_idc is defined as bidirectional prediction, parsing of the LIC flag may be skipped.

[0226] As a result of the determination, if the signaling condition of the LIC flag is satisfied (Yes in S1930), the image decoding apparatus may parse the LIC flag S1940. Conversely, if the signaling condition of the LIC flag is not satisfied (No in S1930), the image decoding apparatus may skip parsing the LIC flag S1940.

[0227] In addition, when the amvpMerge flag is 1 or true ("Yes" in S1920), the amvpMerge mode is applied, and thus the image decoding apparatus can skip the two processes S1930 and S1940 of parsing the LIC flag described above.

[0228] As described above, according to Embodiment 3 of the present disclosure, in amvpMerge mode, regardless of the prediction direction (i.e., the value of inter_pred_idc), the LIC flag may not be parsed. Therefore, when inter_pred_idc is not defined as bidirectional prediction in amvpMerge mode, the following problem can be solved: because amvpMerge is identified as unidirectional prediction during the LIC parsing process, even though LIC may not be applied, compression efficiency is reduced by transmitting the LIC flag.

[0229] Implementation 4

[0230] According to embodiment 4 of the present disclosure, in amvpMerge mode, the reference picture list may be composed of a single reference picture list. The single reference picture list may be reordered based on a predetermined RPR condition, and the maximum number of candidates for the amvpMerge mode may be adaptively determined based on whether RPR is performed. In other words, when the reference picture indicated by the reference picture index obtained in the parsing process of the mvp index is resampled, the maximum number of allowed candidates may be adaptively predefined to apply different binarization processes. According to an embodiment, when the predetermined maximum number of candidates is 1, the mvp index value may be derived (or inferred) to be 0 without a separate parsing process.

[0231] Figure 20 2 is a diagram illustrating an mvp index parsing method in amvpMerge mode according to an embodiment of the present disclosure.

[0232] Reference Figure 20 The image decoding apparatus may parse the syntax element AllowUnitedRefPicList (S2010) from the bitstream to indicate whether a single reference picture list is allowed. In an embodiment, AllowUnitedRefPicList may be obtained at the CU level, in which case a reference picture list may be configured for each CU. In another embodiment, AllowUnitedRefPicList may be obtained at a higher layer (e.g., SPS, PPS, PH, SH), in which case a reference picture list may be configured identically for all CUs that reference the corresponding higher layer.

[0233] When AllowUnitedRefPicList is 1 or true ("Yes" in S2020), the image decoding apparatus may configure a single reference picture list S2030 and reorder the configured single reference picture list S2040. In this case, the reordering may be performed based on whether RPR is performed, size comparison of template costs, etc. A specific example of a method for reordering a single reference picture list based on whether RPR is performed is as follows: Figure 21shown.

[0234] Reference Figure 21 The image decoding apparatus may add the L0 reference picture of index idxY to the single reference picture list uniRefList by performing a loop as many times as the number of L0 reference pictures (S2110). To this end, the loop control variable Y may be initialized to 0 (S2111). In addition, when the variable Y is less than the number of L0 reference pictures ("Yes" in S2112), the image decoding apparatus may add the L0 reference picture of index idxY to the single reference picture list uniRefList (S2113).

[0235] When the loop terminates, the variable Y is initialized to 0 again (S2121). In addition, the image decoding device may add the L1 reference picture of index idxY to the single reference picture list uniRefList (S2120) by looping by the number of L1 reference pictures. Specifically, when the variable Y is less than the number of L1 reference pictures ("Yes" in S2122), the image decoding device may determine whether the L1 reference picture of index idxY does not exist in the single reference picture list uniRefList (S2123). As a result of the determination, when the corresponding reference picture does not exist in the single reference picture list uniRefList ("Yes" in S2123), the image decoding device may add the corresponding reference picture to the single reference picture list uniRefList (S2124) and repeat the above-mentioned processes S2122 to S2124. On the contrary, when the corresponding reference picture already exists in the single reference picture list uniRefList ("No" in S2123), the image decoding device may increase the variable Y by 1 and repeat the above processes S2122 to S2124 without adding the corresponding reference picture to the single reference picture list uniRefList.

[0236] The image decoding apparatus may reorder the single reference picture list uniRefList configured through the above-described process based on whether RPR is performed (S2130). Specifically, the image decoding apparatus may determine whether each reference picture (LX-idxY) in the single reference picture list has been resampled (S2131). As a result of the determination, when the corresponding reference picture has been resampled ("Yes" in S2131), the image decoding apparatus may set the template cost of the corresponding reference picture to the maximum (S2132). Conversely, when the corresponding reference picture has not been resampled ("No" in S2131), the image decoding apparatus may set the template cost of the corresponding reference picture to 0 (S2133). In addition, the image decoding apparatus may sort the single reference picture list uniRefList in ascending order of template cost (S2134).

[0237] according to Figure 21In this reordering method, the reference picture with the smaller template cost may have a reference picture index value with fewer bits. In other words, the reference picture with the smaller template cost is located at the front in the single reference picture list.

[0238] Refer again Figure 20 , when AllowUnitedRefPicList is 0 or false (“No” in S2020 ), the above-mentioned single reference picture list configuration S2030 and reordering S2040 may be skipped.

[0239] The image decoding device may parse the bitstream for amvpMergeFlag indicating whether the amvpMerge mode is applied (S2050). In embodiments, the amvpMergeFlag may be obtained at the CU or PU level. Furthermore, the image decoding device may parse the bitstream for an amvpMerge reference index (i.e., refIdxAmvpMerge) indicating a reference picture to be used in the amvpMerge mode (S2060).

[0240] When amvpMergeFlag is 1 or true ("Yes" in S2070), the image decoding device may parse the mvp indexes (S2080, S2090, and S2095) based on whether RPR is performed. Specifically, when the reference picture indicated by refIdxAmvpMerge is resampled (i.e., reference picture (refIdxAmvpMerge) = RPR) ("Yes" in S2080), the image decoding device may parse the mvp index (S2090) binarized using the maximum number of candidates M from the bitstream. Conversely, when the reference picture indicated by refIdxAmvpMerge is not resampled (i.e., reference picture (refIdxAmvpMerge) != RPR) ("No" in S2080), the image decoding device may parse the mvp index (S2095) binarized using the maximum number of candidates N from the bitstream. In this manner, in the amvpMerge mode, the maximum number of candidates may be adaptively determined as M or N based on whether the reference picture is resampled, and the mvp index may be resolved by using a binarization process based on the maximum number of candidates.

[0241] Furthermore, when amvpMergeFlag is 0 or false (No in S2070 ), the above-described RPR check S2080 and the mvp index parsing based on whether RPR is performed S2090 and S2095 may be skipped.

[0242] Figure 22 is a diagram illustrating an mvp index parsing method in amvpMerge mode according to another embodiment of the present disclosure. Figure 22The single reference picture list configuration S2210 to S2240 and amvpMergeFlag and refIdxAmvpMerge parsing process S2250 and S2260 are as described above. Figure 20 Hereinafter, the MVP index parsing process will be described in detail.

[0243] When amvpMergeFlag is 1 or true (Yes in S2270 ), the image decoding apparatus may parse the mvp index based on whether the AMVP candidates are reordered based on template matching and whether RPR is performed.

[0244] Specifically, when template matching-based reordering is performed on the AMVP candidates ("Yes" in S2280), the image decoding device may determine whether to perform RPR (S2282). As a result of the determination, when the reference picture indicated by refIdxAmvpMerge is resampled (i.e., reference picture (refIdxAmvpMerge) = RPR) ("Yes" in S2282), the image decoding device may parse the bitstream and obtain the mvp index binarized using the maximum number of candidates M (S2290). On the contrary, when the reference picture indicated by refIdxAmvpMerge is not resampled (i.e., reference picture (refIdxAmvpMerge) ! = RPR) ("No" in S2282), the image decoding device may parse the bitstream and obtain the mvp index binarized using the maximum number of candidates N (S2292).

[0245] When template matching-based reordering is not performed on the AMVP candidates (No in S2280 ), the image decoding apparatus may parse the mvp index binarized with the maximum number Q of candidates from the bitstream S2294 .

[0246] In this way, in amvpMerge mode, the maximum number of candidates can be adaptively determined to be any one of M, N, or Q based on whether AMVP candidates are reordered based on template matching and whether RPR is performed. In addition, a binarization process based on the maximum number of candidates can be used to resolve the mvp index.

[0247] When amvpMergeFlag is 0 or false (No in S2270), all of the above steps of checking whether to perform reordering based on template matching S2280, checking whether to perform RPR S2282, and parsing the mvp index based on whether to perform RPR S2290 to S2294 can be skipped.

[0248] Figure 23 is a diagram illustrating an mvp index parsing method in amvpMerge mode according to another embodiment of the present disclosure. Figure 23The single reference picture list configuration S2310 to S2340 and amvpMergeFlag and refIdxAmvpMerge parsing process S2350 and S2360 are as described above. Figure 20 Hereinafter, the MVP index parsing process will be described in detail.

[0249] When amvpMergeFlag is 1 or true (Yes in S2370 ), the image decoding apparatus may parse the mvp index based on whether template matching-based reordering is performed and whether RPR is performed.

[0250] Specifically, when template matching-based reordering is not performed on the AMVP candidates (No in S2380), the image decoding device may determine whether to perform RPR (S2382). As a result of the determination, when the reference picture indicated by refIdxAmvpMerge is resampled (i.e., reference picture (refIdxAmvpMerge) = RPR) (Yes in S2382), the image decoding device may parse the bitstream and obtain the mvp index binarized using the maximum number of candidates M (S2390). Conversely, when the reference picture indicated by refIdxAmvpMerge is not resampled (i.e., reference picture (refIdxAmvpMerge) ! = RPR) (No in S2382), the image decoding device may parse the bitstream and obtain the mvp index binarized using the maximum number of candidates N (S2392).

[0251] When template matching-based reordering is performed on the AMVP candidates (Yes in S2380), the image decoding device may determine whether to perform RPR (S2384). As a result of the determination, when the reference picture indicated by refIdxAmvpMerge is resampled (i.e., reference picture (refIdxAmvpMerge) = RPR) (Yes in S2384), the image decoding device may parse the bitstream and obtain the binarized mvp index using the maximum number of candidates Q (S2394). On the contrary, when the reference picture indicated by refIdxAmvpMerge is not resampled (i.e., reference picture (refIdxAmvpMerge) ! = RPR) (No in S2384), the image decoding device may parse the bitstream and obtain the binarized mvp index using the maximum number of candidates P (S2396).

[0252] In this way, in amvpMerge mode, the maximum number of candidates can be adaptively determined to be any one of M, N, Q, or P based on whether AMVP candidates are reordered based on template matching and whether RPR is performed. In addition, a binarization process based on the maximum number of candidates can be used to resolve the MVP index.

[0253] When amvpMergeFlag is 0 or false (No in S2370), all of the above steps of checking whether to perform reordering based on template matching S2380, checking whether to perform RPR S2382, and parsing the mvp index based on whether to perform RPR S2390 to S2396 can be skipped.

[0254] In the above Figures 20 to 23 In the example of , variables M, N, Q, and P are predefined values, and when the corresponding values are 1, the maximum number of allowed candidates refers to 1. When the maximum number of candidates is 1, the MVP index is not signaled, and the MVP index value can be derived (or inferred) as 0. Conversely, when the maximum number of candidates is greater than 1, the MVP index can be signaled. For example, when the maximum number of candidates is 2, a 1-bit MVP index can be signaled, and when the maximum number of candidates is 3, a 2-bit MVP index can be signaled.

[0255] As described above, according to embodiment 4 of the present disclosure, in amvpMerge mode, the maximum number of candidates may be adaptively determined based on whether the reference picture is an RPR, and a binarization process for an mvp index may be applied based on the determined maximum number of candidates.

[0256] In the following, reference is made to Figure 24 and Figure 25 , an image encoding / decoding method according to an embodiment of the present disclosure will be described in detail.

[0257] Figure 24 is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure. Figure 24 The image decoding method in Figure 3 The image decoding device is executed in.

[0258] Reference Figure 24 When inter-frame prediction is applied to the current block, the image decoding device may configure a reference picture list for the current block (S2410). The image decoding device may reorder the reference picture list based on whether the prediction mode of the current block is the amvpMerge mode (S2420). Here, the amvpMerge mode refers to a prediction mode in which the Advanced Motion Vector Prediction (AMVP) mode is applied to the first prediction direction of the current block and the merge mode is applied to the second prediction direction. In an embodiment, based on the prediction mode of the current block being the amvpMerge mode, template matching-based reordering may not be applied to the reference picture list. In addition, the image decoding device may generate prediction samples of the current block based on the reference picture list (S2430).

[0259] In an embodiment, the reference picture list may be configured as a single list based on a first reference picture list for a first prediction direction and a second reference picture list for a second prediction direction.

[0260] In an embodiment, the prediction mode based on the current block is amvpMerge mode, and the reference picture list can be reordered based on at least one of the picture order count (POC) difference between the current picture and the reference picture, the temporal ID or the quantization parameter of the reference picture.

[0261] In an embodiment, based on the prediction mode of the current block being the amvpMerge mode, the maximum number of motion vector predictor (MVP) candidates for the first prediction direction may be determined differently based on whether template matching-based motion compensation is allowed.

[0262] In an embodiment, based on the prediction mode of the current block being the amvpMerge mode, a local illumination compensation (LIC) flag indicating whether LIC is applied to the current block may be obtained from a bitstream only when bidirectional prediction is applied to the current block.

[0263] In an embodiment, based on the prediction mode of the current block being amvpMerge mode, the maximum number of AMVP candidates for the first prediction direction may be determined differently depending on whether the reference picture of the current block has been resampled. In this case, reordering may be performed based on whether each reference picture in the reference picture list has been resampled. Furthermore, reordering may be performed by setting the template cost of the resampled reference pictures in the reference picture list to a predetermined maximum value.

[0264] In an embodiment, based on the prediction mode of the current block being the amvpMerge mode, the maximum number of AMVP candidates for the first prediction direction may be determined differently based on whether template matching-based reordering is applied to the AMVP candidates.

[0265] Figure 25 is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure. Figure 25 The image coding method in Figure 2 The image encoding device in is executed.

[0266] Reference Figure 25When inter-frame prediction is applied to the current block, the image encoding device may configure a reference picture list for the current block S2510. The image encoding device may reorder the reference picture list based on whether the prediction mode of the current block is the amvpMerge mode S2520. Here, the amvpMerge mode refers to a prediction mode in which the Advanced Motion Vector Prediction (AMVP) mode is applied to the first prediction direction of the current block and the merge mode is applied to the second prediction direction. In an embodiment, based on the prediction mode of the current block being the amvpMerge mode, template matching-based reordering may not be applied to the reference picture list. The image encoding device may generate prediction samples of the current block based on the reference picture list S2530. Then, the image encoding device may encode a reference picture index indicating the reference picture used to generate the prediction samples into the bitstream S2540.

[0267] In an embodiment, the reference picture list may be configured as a single list based on a first reference picture list for a first prediction direction and a second reference picture list for a second prediction direction.

[0268] In an embodiment, the prediction mode based on the current block is amvpMerge mode, and the reference picture list can be reordered based on at least one of the picture order count (POC) difference between the current picture and the reference picture, the temporal ID or the quantization parameter of the reference picture.

[0269] In addition, the bit stream generated by the above-mentioned image encoding method can be stored in a computer-readable recording medium or transmitted to an image decoding device through a network.

[0270] Although the exemplary methods of the present disclosure are shown as a series of operations for clarity of description, this is not intended to limit the order in which the steps are performed, and each step may be performed simultaneously or in a different order if desired. In order to implement the method according to the present disclosure, another step may be additionally included in the exemplary steps, or the remaining steps may be included in addition to some steps, or another additional step may be included in addition to some steps.

[0271] In the present disclosure, an image encoding device or image decoding device that performs a predetermined operation (step) may perform an operation (step) for checking a condition or situation for performing the corresponding operation (step). For example, when it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or image decoding device may perform an operation for checking whether the predetermined condition is satisfied and then perform the predetermined operation.

[0272] The various embodiments of the present disclosure do not list all possible combinations but are intended to describe representative aspects of the present disclosure, and matters described in the various embodiments can be applied independently or in combination of at least two.

[0273] In addition, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. For implementation by hardware, they may be implemented by one or more ASICs (application-specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), general-purpose processors, controllers, microcontrollers, microprocessors, and the like.

[0274] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied may be included in multimedia broadcast transmission and reception devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video communication devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video on demand (VoD) service providers, OTT video (over-the-top video) devices, Internet streaming service providers, three-dimensional (3D) video devices, video phone video devices, medical video devices, and the like, and may be used to process video signals or data signals. For example, OTT video (over-the-top video) devices may include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), and the like.

[0275] Figure 26 An exemplary schematic diagram showing a content streaming system to which embodiments of the present disclosure may be applied is shown.

[0276] like Figure 26 As shown, a content streaming system to which the embodiments of the present disclosure are applied may generally include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.

[0277] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data, generates a bitstream, and sends it to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server can be omitted.

[0278] A bitstream may be generated by applying the video encoding method and / or the image encoding apparatus according to the embodiments of the present disclosure, and the streaming server may temporarily store the bitstream during a process of transmitting or receiving the bitstream.

[0279] The streaming server can transmit multimedia data to a user device via a web server based on user requests, and the web server can act as an intermediary to notify users of available services. When a user requests a desired service from the web server, the web server can transmit the request to the streaming server, and the streaming server can transmit the multimedia data to the user. In this case, the content streaming system can include a separate control server, and in this case, the control server can be used to control the command / response exchanges between devices within the content streaming system.

[0280] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, in order to provide a seamless streaming service, the streaming server may store the bitstream for a period of time.

[0281] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet PCs, ultrabooks, wearable devices (i.e., smart watches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signage.

[0282] Each server within the content streaming system may operate as a distributed server, in which case data received by each server may be processed in a distributed manner.

[0283] The scope of the present disclosure includes software or machine-executable instructions (i.e., operating systems, applications, firmware, programs, etc.) that enable methods according to various embodiments to be performed on a device or computer, as well as non-transitory computer-readable media in which such software or instructions are stored and can be executed on a device or computer.

[0284] Industrial Applicability

[0285] The embodiments of the present disclosure may be used for encoding / decoding of images.

Claims

1. An image decoding method performed by an image decoding device, the method comprising the following steps: Build a reference picture list for the current block; reordering the reference picture list based on whether a prediction mode of the current block is an amvpMerge mode, wherein the amvpMerge mode is a prediction mode in which an advanced motion vector prediction (AMVP) mode is applied to a first prediction direction of the current block and a merge mode is applied to a second prediction direction; and Generate prediction samples of the current block based on the reference picture list, The prediction mode based on the current block is the amvpMerge mode, and template matching-based reordering is not applied to the reference picture list.

2. The method according to claim 1, wherein The reference picture list consists of a single list of a first reference picture list based on the first prediction direction and a second reference picture list based on the second prediction direction.

3. The method according to claim 1, wherein Based on the prediction mode of the current block being the amvpMerge mode, the reference picture list is reordered based on at least one of a picture order count (POC) difference between the current picture and the reference picture, a temporal identifier (TID) of the reference picture, or a quantization parameter, wherein the TID is a TID.

4. The method according to claim 1, wherein Based on whether the prediction mode of the current block is the amvpMerge mode, a maximum number of motion vector predictor MVP candidates for the first prediction direction is determined differently based on whether template matching-based motion compensation is allowed.

5. The method according to claim 1, wherein Based on the prediction mode of the current block being the amvpMerge mode, only when bidirectional prediction is applied to the current block, an LIC flag indicating whether local illumination compensation (LIC) is applied to the current block is obtained from a bitstream.

6. The method according to claim 1, wherein Based on whether the prediction mode of the current block is the amvpMerge mode, a maximum number of AMVP candidates for the first prediction direction is determined differently based on whether a reference picture of the current block is resampled.

7. The method according to claim 1, wherein The reordering is performed based on whether each reference picture in the reference picture list is resampled.

8. The method according to claim 7, wherein: The reordering is performed by setting the template costs of the resampled reference pictures in the reference picture list to a predetermined maximum value.

9. The method according to claim 1, wherein: Based on the prediction mode of the current block being the amvpMerge mode, the maximum number of AMVP candidates is determined differently based on whether the template matching-based reordering is applied to the AMVP candidates for the first prediction direction.

10. An image decoding device, comprising a memory and at least one processor, wherein: The at least one processor: Build a reference picture list for the current block; reordering the reference picture list based on whether a prediction mode of the current block is an amvpMerge mode, wherein the amvpMerge mode is a prediction mode in which an advanced motion vector prediction (AMVP) mode is applied to a first prediction direction of the current block and a merge mode is applied to a second prediction direction; and Generate prediction samples of the current block based on the reference picture list, The prediction mode based on the current block is the amvpMerge mode, and template matching-based reordering is not applied to the reference picture list.

11. An image encoding method performed by an image encoding device, the method comprising the following steps: Build a reference picture list for the current block; reordering the reference picture list based on whether a prediction mode of the current block is an amvpMerge mode, wherein the amvpMerge mode is a prediction mode in which an advanced motion vector prediction (AMVP) mode is applied to a first prediction direction of the current block and a merge mode is applied to a second prediction direction; generating a prediction sample of the current block based on the reference picture list; and encoding into a bitstream a reference picture index representing a reference picture used to generate the prediction sample, The prediction mode based on the current block is the amvpMerge mode, and template matching-based reordering is not applied to the reference picture list.

12. The method according to claim 11, wherein The reference picture list consists of a single list of a first reference picture list based on the first prediction direction and a second reference picture list based on the second prediction direction.

13. The method according to claim 11, wherein Based on the prediction mode of the current block being the amvpMerge mode, the reference picture list is reordered based on at least one of a picture order count (POC) difference between the current picture and the reference picture, a temporal identifier (TID) of the reference picture, or a quantization parameter, wherein the TID is a TID. 14 . A non-transitory computer-readable recording medium storing a bit stream generated by the image encoding method according to claim 11 .

15. A method for transmitting a bit stream generated by an image encoding method, the image encoding method comprising the steps of: Build a reference picture list for the current block; reordering the reference picture list based on whether a prediction mode of the current block is an amvpMerge mode, wherein the amvpMerge mode is a prediction mode in which an advanced motion vector prediction (AMVP) mode is applied to a first prediction direction of the current block and a merge mode is applied to a second prediction direction; generating a prediction sample of the current block based on the reference picture list; and encoding into a bitstream a reference picture index representing a reference picture used to generate the prediction sample, The prediction mode based on the current block is the amvpMerge mode, and template matching-based reordering is not applied to the reference picture list.