Image encoding / decoding method using template matching, method for transmitting bitstream, and recording medium storing bitstream

JP2025513819A5Pending Publication Date: 2026-05-13LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2023-04-11
Publication Date
2026-05-13

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An image decoding method according to the present disclosure is an image decoding method performed by an image decoding device, the image decoding method including the steps of: determining information about template matching for a current block on which a template matching-based technique is performed, calculating a template matching cost based on the determined information about template matching, and performing the template matching-based technique based on the calculated template matching cost.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an image encoding / decoding method using template matching (TM), a method for transmitting a bitstream, and a recording medium storing the bitstream, and more particularly, to a method for performing inter prediction using template matching. [Background technology]

[0002] Recently, the demand for high-resolution, high-quality images, for example, HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various fields. As the resolution and quality of image data increases, the amount of information or bits transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits transmitted leads to an increase in transmission costs and storage costs.

[0003] This requires a highly efficient image compression technique for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide a method for performing inter prediction using template matching.

[0006] Another object of the present disclosure is to provide a method for signaling (transmitting or receiving) additional information required when using a coding tool based on template matching (such as the size of the template, a cost function for calculating the cost of template matching, and application of weights to the cost of template matching).

[0007] Another object of the present disclosure is to provide various methods for calculating more accurate template matching costs based on various conditions.

[0008] Another object of the present disclosure is to provide a non-transitory computer-readable recording medium for storing a bitstream generated by the image encoding method according to the present disclosure.

[0009] Another object of the present disclosure is to provide a non-transitory computer-readable recording medium for storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to reconstruct an image.

[0010] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method according to the present disclosure.

[0011] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not described above will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the following description. [Means for solving the problem]

[0012] An image decoding method according to one aspect of the present disclosure is an image decoding method performed by an image decoding device, and may include a step of determining information regarding template matching for a current block on which a template matching-based technique is performed, a step of calculating a template matching cost based on the determined template matching information, and a step of performing the template matching-based technique based on the calculated template matching cost.

[0013] In the image decoding method according to the present disclosure, the information on template matching may include information on the size of the template for calculating the template matching cost, or information on a difference base function for calculating the template matching cost.

[0014] In the image decoding method according to the present disclosure, information on the size of the template or information on the difference base function may be included in and signaled at a higher level of the current block.

[0015] In the image decoding method according to the present disclosure, the information on the size of the template or the information on the difference base function may vary depending on the template matching based technology.

[0016] In the image decoding method according to the present disclosure, a template for calculating the template matching cost may be adaptively determined based on coding information of the current block or a size of the current block.

[0017] In the image decoding method according to the present disclosure, the template for calculating the template matching cost may be adaptively determined based on a comparison between the size of the current block and a predetermined threshold.

[0018] In the image decoding method according to the present disclosure, the template for calculating the template matching cost may be adaptively determined based on the position or type of the merging candidate that is the subject of the template matching cost calculation.

[0019] In the image decoding method according to the present disclosure, the template for calculating the template matching cost may be adaptively determined based on the magnitude or directionality of a motion vector that is a target of the template matching cost calculation.

[0020] In the image decoding method according to the present disclosure, the template matching cost is adjusted to a weighted sum of the template matching cost and the spatial similarity cost, and the spatial similarity cost can be derived based on a difference between a reference template adjacent to the reference block and a pixel value of a pixel in the reference block adjacent to the reference template.

[0021] In the image decoding method according to the present disclosure, the weights of the weighted sum may be signaled by being included in a higher level of the current block, or may be derived based on the size of the current block or the coding information of the current block.

[0022] In the image decoding method according to the present disclosure, the template matching cost may be adjusted to a value obtained by applying a weight to the template matching cost.

[0023] In the image decoding method according to the present disclosure, the weights can be adaptively determined based on the positions or types of merging candidates that are the subject of template matching cost calculation.

[0024] In the image decoding method according to the present disclosure, the weight can be adaptively determined based on the magnitude or directionality of a motion vector that is the subject of template matching cost calculation.

[0025] An image encoding method according to another aspect of the present disclosure may perform operations corresponding to the image decoding method according to one aspect of the present disclosure.

[0026] An image encoding method according to the present disclosure is an image encoding method performed by an image encoding device, and may include a step of determining information related to template matching for a current block on which a template matching-based technique is performed, a step of calculating a template matching cost based on the determined template matching information, and a step of performing the template matching-based technique based on the calculated template matching cost.

[0027] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or apparatus of the present disclosure.

[0028] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image coding method or apparatus of the present disclosure.

[0029] The features described above in the brief summary of the present disclosure are merely exemplary embodiments of the detailed description of the present disclosure that follows and are not intended to limit the scope of the present disclosure. Effect of the Invention

[0030] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0031] Furthermore, according to the present disclosure, a method for performing inter prediction using template matching can be provided.

[0032] In addition, according to the present disclosure, when using a coding tool based on template matching, a method for signaling (transmitting or receiving) additional information required for the coding tool (such as the size of the template, a cost function for calculating the cost of template matching, and application of weights to the cost of template matching) is provided, thereby enabling efficient template matching-based coding.

[0033] The present disclosure also provides various methods for calculating more accurate template matching costs based on various conditions.

[0034] According to the present disclosure, a non-transitory computer-readable recording medium can be provided that stores a bitstream generated by the image encoding method according to the present disclosure.

[0035] In addition, according to the present disclosure, a non-transitory computer-readable recording medium can be provided that stores a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.

[0036] Furthermore, according to the present disclosure, a method for transmitting a bitstream generated by an image coding method can be provided.

[0037] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief description of the drawings]

[0038] [Figure 1] FIG. 1 is a schematic diagram illustrating a video coding system to which an embodiment of the present disclosure can be applied. [Diagram 2] 1 is a diagram illustrating an image encoding device to which an embodiment of the present disclosure can be applied; [Diagram 3] 1 is a diagram illustrating an image decoding device to which an embodiment of the present disclosure can be applied; [Figure 4] FIG. 2 is a diagram illustrating an inter prediction unit of an image encoding device. [Diagram 5] 1 is a flowchart illustrating a method for encoding an image based on inter prediction. [Figure 6] FIG. 2 is a diagram illustrating an inter prediction unit of an image decoding device. [Figure 7] 1 is a flowchart illustrating a method for decoding an image based on inter prediction. [Figure 8] 1 is a flowchart illustrating an inter prediction method. [Figure 9] 1 is a diagram illustrating a template matching-based encoding / decoding method according to the present disclosure. [Figure 10] 1 is a diagram illustrating an example of a template of a current block and a reference sample of the template in a reference picture. [Figure 11]FIG. 13 is a diagram for explaining a method for identifying a template using motion information of a sub-block. [Figure 12] FIG. 2 is a diagram for explaining an image decoding method according to the present disclosure. [Figure 13] FIG. 1 is a diagram for explaining an image encoding method according to the present disclosure. [Figure 14] 11A to 11C are diagrams illustrating various examples of selectable templates according to an embodiment of the present disclosure. [Figure 15] 11A to 11C are diagrams illustrating various examples of selectable templates according to another embodiment of the present disclosure. [Figure 16] 10 is a diagram illustrating an example of a template of a reference block and boundary pixels of the reference block for calculating a spatial similarity cost; [Figure 17] FIG. 2 is a diagram for explaining an image decoding / encoding method according to the present disclosure. [Figure 18] FIG. 1 is a diagram illustrating an exemplary content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0039] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present disclosure will be described in detail below with reference to the accompanying drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0040] In describing the embodiments of the present disclosure, if it is determined that a specific description of a known configuration or function may make the gist of the present disclosure unclear, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure are omitted, and similar parts are denoted by similar reference numerals.

[0041] In the present disclosure, when a certain component is "connected," "coupled," or "connected" to another component, this includes not only a direct connection relationship, but also an indirect connection relationship in which another component exists between them. Furthermore, when a certain component is described as "including" or "having" another component, this does not mean that the other component is excluded, but that the other component can be further included, unless otherwise specified.

[0042] In this disclosure, terms such as "first" and "second" are used only for the purpose of distinguishing one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be called a second component in another embodiment, and similarly, a second component in one embodiment may be called a first component in another embodiment.

[0043] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component, and do not necessarily mean that the components are separate. In other words, multiple components may be integrated and configured as a single hardware or software unit, or one component may be distributed and configured as multiple hardware or software units. Thus, even if not otherwise stated, such integrated or distributed embodiments are also included in the scope of the present disclosure.

[0044] In the present disclosure, the components described in the various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also included in the scope of the present disclosure. In addition, an embodiment including other components in addition to the components described in the various embodiments is also included in the scope of the present disclosure.

[0045] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have ordinary meanings in the technical field to which the present disclosure belongs, unless they are newly defined in this disclosure.

[0046] In this disclosure, a "picture" generally means a unit indicating any one image in a particular time period, a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. Also, a slice / tile may include one or more coding tree units (CTUs).

[0047] In this disclosure, a "pixel" or a "pel" may refer to the smallest unit that constitutes one picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally indicate a pixel or a pixel value, may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.

[0048] In this disclosure, a "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. A unit may be used interchangeably with terms such as "sample array," "block," or "area," depending on the case. In a general case, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.

[0049] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." If prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." If transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." If filtering is performed, a "current block" may refer to a "block to be filtered."

[0050] Additionally, in this disclosure, unless explicitly stated as a chroma block, a "current block" may refer to a block including both a luma component block and a chroma component block, or a "luma block of the current block." The luma component block of the current block may be explicitly expressed as including an explicit description of the luma component block, such as a "luma block" or a "current luma block." Additionally, the chroma component block of the current block may be explicitly expressed as including an explicit description of the chroma component block, such as a "chroma block" or a "current chroma block."

[0051] In the present disclosure, " / " and "," can be interpreted as "and / or." For example, "A / B" and "A, B" can be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."

[0052] In this disclosure, "or" can be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."

[0053] Video Coding System Overview

[0054] FIG. 1 is a schematic diagram of a video coding system to which embodiments of the present disclosure can be applied.

[0055] A video coding system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0056] The encoding device 10 according to an embodiment may include a video source generating unit 11, an encoding unit 12, and a transmitting unit 13. The decoding device 20 according to an embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitting unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.

[0057] The video source generating unit 11 can obtain videos / images through a process of capturing, synthesizing, or generating videos / images. The video source generating unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured videos / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate videos / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.

[0058] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of steps such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in a bitstream format.

[0059] The transmitting unit 13 may obtain the encoded video / image information or data output in a bitstream format and may transmit the same to the receiving unit 21 of the decoding device 20 or other external objects in a file or streaming format via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit 13 may include elements for generating a media file via a predetermined file format and may include elements for transmitting via a broadcasting / communication network. The transmitting unit 13 may be provided as a transmitting device separate from the encoding device 12, in which case the transmitting device may include at least one processor for obtaining the encoded video / image information or data output in a bitstream format and a transmitting unit for transmitting the same in a file or stream format. The receiving unit 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0060] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.

[0061] The rendering unit 23 can render the decoded video / images. The rendered video / images can be displayed via a display unit.

[0062] Overview of the image encoding device

[0063] FIG. 2 is a diagram illustrating an image encoding device to which an embodiment of the present disclosure can be applied.

[0064] As shown in Fig. 2, the image coding device 100 may include an image division unit 110, a subtraction unit 115, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy coding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transformation unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0065] All or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor) depending on the embodiment. Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.

[0066] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / Binary-tree / Ternal-tree) structure. For example, one coding unit may be divided into a plurality of coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For dividing the coding units, a quad-tree structure may be applied first, and a binary-tree structure and / or a ternary-tree structure may be applied later. A coding procedure according to the present disclosure may be performed based on a final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, and a lower depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0067] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a block to be processed (current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied in units of a current block or a CU. The prediction unit may generate various information related to prediction of the current block and transmit it to the entropy encoding unit 190. The information related to prediction may be encoded by the entropy encoding unit 190 and output in a bitstream format.

[0068] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from the current block according to an intra prediction mode and / or an intra prediction technique. The intra prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional mode may include, for example, a DC mode and a Planar mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the fineness of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the setting. The intra prediction unit 185 may also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.

[0069] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of a block, a sub-block, or a sample based on the correlation of motion information between a neighboring block and a current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring block may include a spatial neighboring block present in the current picture and a temporal neighboring block present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter prediction unit 180 may generate information indicating which candidate is used to derive a motion vector and / or a reference picture index of the current block by forming a motion information candidate list based on neighboring blocks. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode and a merge mode, the inter prediction unit 180 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of the current block can be signaled by using the motion vector of a neighboring block as a motion vector predictor and encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0070] The prediction unit may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for prediction of the current block, and may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction for prediction of the current block may be called combined inter and intra prediction (CIIP). The prediction unit may also perform intra block copy (IBC) for prediction of the current block. Intra block copy can be used for content image / video coding such as games, for example, as in screen content coding (SCC). IBC is a method of predicting a current block using an already restored reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC is basically performed in the current picture, but may be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.

[0071] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, prediction sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.

[0072] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size, or may be applied to non-square, variable-sized blocks.

[0073] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may code the quantized signal (information on the quantized transform coefficients) and output the coded signal in a bitstream format. The information on the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information on the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.

[0074] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information required for video / image restoration (e.g., values ​​of syntax elements, etc.) together or separately in addition to the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a network abstraction layer (NAL) unit unit in a bitstream format. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may further include general constraint information. The signaling information, transmitted information and / or syntax elements referred to in this disclosure may be encoded through the above-mentioned encoding procedures and included in the bitstream.

[0075] The bitstream may be transmitted via a network or may be stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storing unit (not shown) for storing the signal may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.

[0076] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the quantized transform coefficients are subjected to inverse quantization and inverse transformation via the inverse quantization unit 140 and the inverse transformation unit 150, so that a residual signal (residual block or residual sample) can be restored.

[0077] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to a prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, a predicted block may be used as a reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering as described below.

[0078] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, and the like. The filtering unit 160 may generate various information related to filtering, as will be described later in the description of each filtering method, and transmit the information to the entropy coding unit 190. The information related to filtering may be coded by the entropy coding unit 190 and output in a bitstream format.

[0079] The modified reconstructed picture transmitted to the memory 170 may be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 may avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and may also improve encoding efficiency.

[0080] The DPB in the memory 170 may store modified reconstructed pictures to be used as reference pictures in the inter prediction unit 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter prediction unit 180 to be used as motion information of spatial surrounding blocks or motion information of temporal surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra prediction unit 185.

[0081] Overview of the image decoding device

[0082] FIG. 3 is a diagram illustrating an image decoding device to which an embodiment of the present disclosure can be applied.

[0083] 3, the image decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.

[0084] All or at least some of the components constituting the image decoding device 200 may be realized by one hardware component (e.g., a decoder or a processor) depending on the embodiment. Also, the memory 170 may include a DPB and may be realized by a digital storage medium.

[0085] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by executing a process corresponding to the process performed by the image encoding device 100 of Fig. 2. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Thus, the processing unit for decoding can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).

[0086] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in the form of a bitstream. The received signal may be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 may derive information (e.g., video / image information) required for image restoration (or picture restoration) by parsing the bitstream. The video / image information may further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The image decoding apparatus may further use information on the parameter set and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to the residual. More specifically, the CABAC entropy decoding method may receive bins corresponding to each syntax element from the bitstream, determine a context model using syntax element information to be decoded and decoded information of neighboring blocks and a block to be decoded, or information of a symbol / bin decoded in a previous step, predict the occurrence probability of the bin based on the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method may update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to a prediction unit (inter prediction unit 260 and intra prediction unit 265), and residual values ​​entropy-decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0087] Meanwhile, the image decoding device according to the present disclosure may be called a video / image / picture decoding device. The image decoding device may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.

[0088] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients to output transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on a coefficient scan order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0089] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0090] The prediction unit may perform prediction on a current block and generate a predicted block including a prediction sample for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).

[0091] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as has been described in the explanation of the prediction unit of the image encoding device 100.

[0092] The intra prediction unit 265 may predict the current block by referring to samples in the current picture. The description of the intra prediction unit 185 may be similarly applied to the intra prediction unit 265.

[0093] The inter prediction unit 260 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index for the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating a mode (technique) of inter prediction for the current block.

[0094] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 may also be applied to the adder 235. The adder 235 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, and may also be used for inter prediction of the next picture via filtering, as described below.

[0095] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0096] The (modified) reconstructed picture stored in the DPB of the memory 250 may be used as a reference picture in the inter prediction unit 260. The memory 250 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter prediction unit 260 to be used as motion information of a spatial surrounding block or motion information of a temporal surrounding block. The memory 250 may store a reconstructed sample of a reconstructed block in the current picture and transmit it to the intra prediction unit 265.

[0097] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.

[0098] Inter Prediction

[0099] The prediction unit of the image encoding device 100 and the image decoding device 200 may perform inter prediction on a block basis to derive a prediction sample. The inter prediction may be a prediction derived in a manner dependent on data elements (e.g., sample values ​​or motion information, etc.) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector on a reference picture indicated by a reference picture index. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block may be predicted on a block, sub-block, or sample basis based on correlation between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal peripheral block may be the same or different from each other. The temporal peripheral block may be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal peripheral block may be called a collocated picture (colPic). For example, a motion information candidate list may be constructed based on the peripheral blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, a motion vector of a selected neighboring block may be used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.

[0100] The motion information may include L0 motion information and / or L1 motion information depending on an inter prediction type (such as L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be referred to as an L0 prediction, a prediction based on an L1 motion vector may be referred to as an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be referred to as a bi-prediction (Bi) prediction. Here, the L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and the L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include a previous picture in output order from the current picture as a reference picture, and the reference picture list L1 may include a subsequent picture in output order from the current picture. The previous picture may be referred to as a forward (reference) picture, and the subsequent picture may be referred to as a backward (reference) picture. The reference picture list L0 may further include a subsequent picture in output order from the current picture as a reference picture. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed next. The reference picture list L1 may further include a previous picture in output order from the current picture as a reference picture. In this case, the subsequent picture may be indexed first in the reference picture list 1, and the previous picture may be indexed next. Here, the output order may correspond to a picture order count (POC) order.

[0101] FIG. 4 is a diagram illustrating an inter prediction unit (180) of the image decoding device 100, and FIG. 5 is a flowchart showing a method for encoding an image based on inter prediction.

[0102] The image encoding device 100 may perform inter prediction on a current block (S510). The image encoding device 100 may derive an inter prediction mode and motion information of the current block, and generate a prediction sample of the current block. Here, the steps of determining the inter prediction mode, deriving the motion information, and generating the prediction sample may be performed simultaneously, or one of the steps may be performed prior to the other steps. For example, the inter prediction unit 180 of the image encoding device 100 may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183, and the prediction mode determination unit 181 may determine a prediction mode for the current block, the motion information derivation unit 182 may derive motion information of the current block, and the prediction sample derivation unit 183 may derive a prediction sample of the current block. For example, the inter prediction unit 180 of the image encoding device 100 may search for a block similar to the current block within a certain region (search region) of a reference picture through motion estimation, and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion. Based on this, a reference picture index indicating a reference picture in which the reference block is located may be derived, and a motion vector may be derived based on a position difference between the reference block and the current block. The image encoding device 100 may determine a mode to be applied to the current block among various prediction modes. The image encoding device 100 may compare RD costs for the various prediction modes and determine an optimal prediction mode for the current block.

[0103] For example, when a skip mode or a merge mode is applied to the current block, the image encoding device 100 may construct a merge candidate list (to be described later) and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate related to the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to a decoding device. Motion information of the current block may be derived using motion information of the selected merge candidate.

[0104] As another example, when the (A)MVP mode is applied to the current block, the image encoding apparatus 100 may construct an (A)MVP candidate list (to be described later) and use a motion vector of a selected MVP candidate among the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector indicating a reference block derived by the above-mentioned motion estimation may be used as the motion vector of the current block, and an MVP candidate having a motion vector with the smallest difference from the motion vector of the current block among the MVP candidates may become the selected MVP candidate. A motion vector difference (MVD), which is a difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information on the MVD may be signaled to the image decoding apparatus 200. In addition, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the image decoding apparatus 200.

[0105] The image encoding device 100 may derive a residual sample based on the predicted sample (S520). The image encoding device 100 may derive the residual sample by comparing an original sample of the current block with the predicted sample.

[0106] The image encoding device 100 may encode image information including prediction information and residual information (S530). The image encoding device 100 may output the encoded image information in a bitstream format. The prediction information may include prediction mode information (e.g., skip flag, merge flag or mode index, etc.) and information on motion information as information related to the prediction procedure. The information on the motion information may include candidate selection information (e.g., merge index, mvp flag or mvp index) which is information for deriving a motion vector. In addition, the information on the motion information may include the above-mentioned information on MVD and / or reference picture index information. In addition, the information on the motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information on the residual sample. The residual information may include information on a quantized transform coefficient for the residual sample.

[0107] The output bitstream may be stored in a (digital) storage medium and then transmitted to the image decoding device 200, or may be transmitted to the image decoding device 200 via a network.

[0108] Meanwhile, as described above, the image coding apparatus 100 may generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the same prediction result as that performed by the image decoding apparatus 200 is derived by the image coding apparatus 100, thereby improving coding efficiency. Therefore, the image coding apparatus 100 may store the reconstructed picture (or reconstructed sample, reconstructed block) in a memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure or the like may be further applied to the reconstructed picture.

[0109] FIG. 6 is a diagram illustrating the inter prediction unit 260 of the image decoding device 200, and FIG. 7 is a flowchart showing a method for decoding an image based on inter prediction.

[0110] The image decoding apparatus 200 may perform operations corresponding to those performed by the image encoding apparatus 100. The image decoding apparatus 200 may perform prediction on a current block based on received prediction information and derive a prediction sample.

[0111] Specifically, the image decoding apparatus 200 may determine a prediction mode for the current block based on the received prediction information (S710). The image decoding apparatus 200 may determine which inter prediction mode is applied to the current block based on prediction mode information in the prediction information.

[0112] For example, it may be determined whether the merge mode is applied to the current block or the (A)MVP mode is determined based on the merge flag. Alternatively, one of various inter prediction mode candidates may be selected based on the mode index. The inter prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter prediction modes described below.

[0113] The image decoding apparatus 200 may derive motion information of the current block based on the determined inter prediction mode (S720). For example, when a skip mode or a merge mode is applied to the current block, the image decoding apparatus 200 may form a merge candidate list (to be described later) and select one merge candidate from among the merge candidates included in the merge candidate list. The selection may be performed based on the above-mentioned selection information (merge index). Motion information of the selected merge candidate may be used to derive motion information of the current block. The motion information of the selected merge candidate may be used as motion information of the current block.

[0114] As another example, when the (A)MVP mode is applied to the current block, the image decoding apparatus 200 may construct an (A)MVP candidate list described below, and may use a motion vector of a selected MVP candidate among the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the above-mentioned selection information (mvp flag or mvp index). In this case, the MVD of the current block may be derived based on information on the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. Also, the reference picture index of the current block may be derived based on the reference picture index information. A picture pointed to by the reference picture index in a reference picture list for the current block may be derived as a reference picture referenced for inter prediction of the current block.

[0115] Meanwhile, the motion information of the current block may be derived without constructing a candidate list, as described below, in which case the motion information of the current block may be derived according to a procedure disclosed in a prediction mode, as described below. In this case, the candidate list construction as described above may be omitted.

[0116] The image decoding apparatus 200 may generate a prediction sample for the current block based on the motion information of the current block (S730). In this case, the reference picture may be derived based on a reference picture index of the current block, and the prediction sample of the current block may be derived using a sample of a reference block to which the motion vector of the current block points on the reference picture. In this case, as described below, a prediction sample filtering procedure may be further performed on all or a part of the prediction sample of the current block, if necessary.

[0117] For example, the inter prediction unit 260 of the image decoding device 200 may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263, and may determine a prediction mode for the current block based on prediction mode information received by the prediction mode determination unit 181, derive motion information (such as a motion vector and / or a reference picture index) of the current block based on information regarding the motion information received by the motion information derivation unit 182, and derive a prediction sample of the current block by the prediction sample derivation unit 183.

[0118] The image decoding apparatus 200 may generate a residual sample for the current block based on the received residual information (S740). The image decoding apparatus 200 may generate a reconstructed sample for the current block based on the predicted sample and the residual sample, and may generate a reconstructed picture based on the reconstructed sample (S750). Thereafter, an in-loop filtering procedure may be further applied to the reconstructed picture, as described above.

[0119] 8, as described above, the inter prediction procedure may include an inter prediction mode determination step (S810), a motion information derivation step (S820) according to the determined prediction mode, and a prediction execution step (a prediction sample generation step) based on the derived motion information (S830). The inter prediction procedure may be performed in the image encoding device 100 and the image decoding device 200 as described above.

[0120] Inter prediction mode decision

[0121] Various inter prediction modes can be used to predict a current block in a picture. For example, various modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, and merge with MVD (MVVD) mode can be used. Decoder side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, Bi-prediction with CU-level weight (BCW), Bi-directional optical flow (BDOF), etc. can be used in addition to or instead of the affine mode. The affine mode may also be referred to as affine motion prediction mode. The MVP mode may also be referred to as advanced motion vector prediction mode (AMVP). In this document, motion information candidates derived by some modes and / or some modes may be included as one of the motion information related candidates of other modes. For example, an HMVP candidate may be added as a merge candidate of the merge / skip mode, or may be added as an MVP candidate of the MVP mode.

[0122] Prediction mode information indicating an inter prediction mode of a current block may be signaled from the image encoding apparatus 100 to the image decoding apparatus 200. The prediction mode information may be included in a bitstream and received by the image decoding apparatus 200. The prediction mode information may include index information indicating one of a number of candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether or not a skip mode is applied, and if the skip mode is not applied, a merge flag may be signaled to indicate whether or not a merge mode is applied, and if the merge mode is not applied, an MVP mode may be indicated to be applied, or a flag for additional division may be further signaled. The affine mode may be signaled in an independent mode or in a mode dependent on the merge mode or MVP mode. For example, the affine mode may include an affine merge mode and an affine MVP mode.

[0123] Deriving motion information

[0124] Inter prediction may be performed using motion information of the current block. The image encoding device 100 may derive optimal motion information for the current block through a motion estimation procedure. For example, the image encoding device 100 may search for a similar reference block having high correlation with an original block in an original picture for the current block in a fractional pixel unit within a predetermined search range in the reference picture, thereby deriving motion information. The similarity of the block may be derived based on a difference in phase-based sample values. For example, the similarity of the block may be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, the motion information may be derived based on the reference block having the smallest SAD in the search range. The derived motion information may be signaled to the image decoding device 200 according to various methods based on the inter prediction mode.

[0125] Generating prediction samples

[0126] A predicted block for the current block may be derived based on the motion information derived according to the prediction mode. The predicted block may include a prediction sample (prediction sample array) of the current block. If the motion vector of the current block points to a fractional sample unit, an interpolation procedure may be performed, whereby a prediction sample of the current block may be derived based on a reference sample of a fractional sample unit in a reference picture. If affine inter prediction is applied to the current block, a prediction sample may be generated based on a sample / sub-block unit MV. If bi-prediction is applied, a prediction sample derived through a weighted sum or weighted average (by phase) of a prediction sample derived based on L0 prediction (i.e., prediction using a reference picture in a reference picture list L0 and MVL0) and a prediction sample derived based on L1 prediction (i.e., prediction using a reference picture in a reference picture list L1 and MVL1) may be used as a prediction sample of the current block. When bi-prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions from the current picture (i.e., bi-prediction but also bidirectional prediction), this can be called true bi-prediction.

[0127] As described above, reconstructed samples and reconstructed pictures can be generated based on the derived predicted samples, and then procedures such as in-loop filtering can be performed.

[0128] Template matching(TM)

[0129] FIG. 9 is a diagram for explaining a template matching based encoding / decoding method according to the present disclosure.

[0130] Template matching (TM) is a method of deriving a motion vector performed at the decoder end, which can refine motion information of a current block (e.g., current coding unit, current CU) by finding a template (hereinafter, referred to as a "reference template") in a reference picture that is most similar to a template (hereinafter, referred to as a "current template") adjacent to the current block. The current template may be a top neighboring block and / or a left neighboring block of the current block, or a part of these neighboring blocks. In addition, the reference template can be determined to be the same size as the current template.

[0131] As shown in Fig. 9, if an initial motion vector of a current block is derived, a search for a better motion vector can be performed in a surrounding area of ​​the initial motion vector. For example, the range of the surrounding area in which the search is performed may be within a [-8, +8]-pel search area centered on the initial motion vector. Also, the size of a search step for performing the search can be determined based on the AMVR mode of the current block. Also, template matching may be performed consecutively with a bilateral matching process in a merge mode.

[0132] When the prediction mode of the current block is the AMVP mode, a motion vector predictor candidate (MVP candidate) can be determined based on a template matching error. For example, a motion vector predictor candidate (MVP candidate) that minimizes an error between a current template and a reference template can be selected. Then, template matching for improving a motion vector can be performed on the selected motion vector predictor candidate. In this case, template matching for improving a motion vector may not be performed on a motion vector predictor candidate that is not selected.

[0133] More specifically, the refinement of the selected motion vector predictor candidate may start with full-pel accuracy in a [-8, +8]-pel search range using an iterative diamond search. Or, in the case of 4-pel AMVR mode, it may start with 4-pel accuracy. Then, depending on the AMVR mode, a search may be continued with half-pel and / or quarter-pel accuracy. According to the search process, the motion vector predictor candidate may maintain the same motion vector accuracy as indicated by the AMVR mode even after the template matching process. In the iterative search process, if the difference between the previous minimum cost and the current minimum cost is smaller than an arbitrary threshold, the search process ends. The threshold may be equal to the area of ​​the block, i.e., the number of samples in the block. Table 1 shows an example of a search pattern according to the AMVR mode and the merge mode with AMVR.

[0134] [Table 1]

[0135] If the prediction mode of the current block is a merge mode, a similar search method can be applied to the merge candidates indicated by the merge index. As shown in Table 1, the template matching can be performed up to 1 / 8-pel accuracy, or can be skipped below half-pel accuracy, which can be determined depending on whether an alternative interpolation filter is used based on the merge motion information. In this case, the alternative interpolation filter may be a filter used when the AMVR is in half-pel mode. Also, if the template matching is available, depending on whether bilateral matching (BM) is available, the template matching can be operated as an independent process, or can be operated as an additional motion vector improvement process between block-based bilateral matching and sub-block-based bilateral matching. Whether the template matching is available and / or whether the bilateral matching is available can be determined by checking availability conditions. The accuracy of the motion vector in the above can mean the accuracy of a motion vector difference (MVD).

[0136] Adaptive reordering of merge candidates with template matching (ARMC-TM)

[0137] Merge candidates can be adaptively reordered by template matching. The reordering method can be applied to general merge mode, template matching (TM) merge mode, and affine merge mode (except for SbTMVP candidates). In the case of TM merge mode, the reordering of merge candidates can be performed before the motion vector refinement process described above.

[0138] After the merge candidate list is constructed, the merge candidates may be divided into one or more subgroups. The size of the subgroups for the general merge mode and the TM merge mode may be 5, and the size of the subgroups for the affine merge mode may be 3. The merge candidates in each subgroup may be reordered in ascending order by cost values ​​based on template matching. For simplicity, the merge candidates in the last subgroup but not the first subgroup may not be reordered.

[0139] The template matching cost of a merging candidate can be measured by the sum of absolute differences (SAD) between the template samples of the current block and the corresponding reference samples. In this case, the template can include a set of reconstructed samples adjacent to the current block. The template reference samples can be located according to the motion information of the merging candidate.

[0140] FIG. 10 is a diagram illustrating an example of a template of a current block and a reference sample of a template in a reference picture. When a merge candidate uses bidirectional prediction, a reference sample of the template of the merge candidate may be generated as shown in FIG. 10. Specifically, after a reference block in a list0 reference picture is identified based on a list0 (L0) motion vector of a merge candidate of a current block, a reference template (reference template 0, RT0) in the list0 reference picture may be identified. Similarly, after a reference block in a list1 reference picture is identified based on a list1 (L1) motion vector of a merge candidate of a current block, a reference template (reference template 1, RT1) in the list1 reference picture may be identified.

[0141] In the case of a subblock-based merging candidate with a subblock size of Wsub×Hsub, the top template may include one or more subtemplates of size Wsub×1, and the left template may include one or more subtemplates of size 1×Hsub.

[0142] FIG. 11 is a diagram for explaining a method of specifying a template using motion information of sub-blocks.

[0143] In the example shown in FIG. 11, a reference sample of each sub-template may be derived using motion information of sub-blocks included in the first column and the first row of the current block. More specifically, a reference sub-block (A_ref, B_ref, C_ref, D_ref, E_ref, F_ref, G_ref) corresponding to a reference picture may be identified using motion vectors of sub-blocks (A, B, C, D, E, F, G) included in the first column and the first row of the current block in the current picture. For example, a position of a corresponding reference sub-block may be identified from a collocated block in the reference picture based on the motion vector of each sub-block. Then, a reference template may be configured from a restoration area adjacent to each reference sub-block. As shown in FIG. 11, when the size of a sub-block is Wsev×Hsub, the top reference template may include one or more sub-reference templates of size Wsub×1, and the left reference template may include one or more sub-reference templates of size 1×Hsub.

[0144] In the present disclosure, template matching may be a process of searching for a reference template that has the highest similarity to a current template. According to the present disclosure, a template matching cost may be calculated to measure the similarity, and a cost function such as SAD may be used for this purpose. A high template matching cost may mean a high template matching error, and thus a low similarity between templates. Conversely, a low template matching cost may mean a low template matching error, and thus a high similarity between templates.

[0145] In the present disclosure, a cost function for calculating a template matching cost may be a function that utilizes the difference between a sample value in a current template and a corresponding sample value in a reference template. Thus, the cost function may be referred to as a "difference (error)-based function" or a "difference (error)-based equation" between corresponding samples in two templates. Also, the template matching cost calculated by the cost function may be referred to as a "difference (error)-based function value" or a "difference (error)-based value" between corresponding samples in two templates.

[0146] Various embodiments according to the present disclosure will now be described with reference to the drawings.

[0147] Various embodiments of the present disclosure relate to the above-mentioned template matching-based encoding / decoding operation, such as a template matching-based motion vector correction (improvement) process or a template matching-based merge candidate list rearrangement process, etc. The present disclosure proposes the following configurations (Configuration 1 to Configuration 5), which may be applied individually or in combination with one or more other configurations.

[0148] Configuration 1) The size of a template required for calculating the cost of template matching can be transmitted using a higher level syntax such as a sequence parameter set (SPS) or a picture parameter set (PPS).

[0149] Configuration 2) A cost function for calculating the cost of template matching can be transmitted using a higher level syntax such as a sequence parameter set (SPS) or a picture parameter set (PPS).

[0150] Configuration 3) The position or size of the template can be variably selected based on the coding information of the current block or the size information of the current block.

[0151] Configuration 4) The cost of template matching can be calculated by a weighted sum of the template matching error and the template spatial similarity cost.

[0152] Configuration 5) The cost of template matching can be weighted according to the magnitude of the motion vector, the position of the motion vector, or the type of the motion vector.

[0153] Example 1

[0154] According to this embodiment, the size of a template required for calculating the cost of template matching can be variably set, and information on the size of the template can be signaled. For example, the information on the size of the template can be transmitted from the image encoding device to the image decoding device using a higher level syntax such as a sequence parameter set (SPS) or a picture parameter set (PPS).

[0155] The availability of the method of rearranging merging candidates based on template matching can be signaled at a higher level, such as the SPS, as shown in Table 2.

[0156] [Table 2]

[0157] In Table 2, sps_aml_enabled_flag may be information indicating whether aml is available in a coded layer video sequence (CLVS) that refers to the SPS. The aml may mean an adaptive merge list, which may refer to a technique for adaptively rearranging a merge list based on a template matching cost. According to this embodiment, if the aml is available, information regarding the size of a template for calculating a template matching cost may be additionally signaled, for example, as shown in Table 3.

[0158] [Table 3]

[0159] In Table 3, sps_aml_tm_size_idx may be information (index) on the size of a template used in a merge candidate list rearrangement technique based on template matching. The size of the template can be derived from the index and the table in Table 4.

[0160] [Table 4]

[0161] Table 4 may be predefined in the image encoding device and the image decoding device. The derivation of the template size by the above method is merely an example, and for example, the sps_aml_tm_size_idx value may be input into a predetermined formula to derive the template size as the resultant value. For example, the predetermined formula may include left-shifting.

[0162] The ranges of syntax values ​​and / or template sizes are merely examples and are not limited to the above examples, and various other combinations are possible.

[0163] Furthermore, the information regarding the size of the template may be transmitted at any other higher level, such as a picture parameter set or a picture header, other than the above-mentioned SPS.

[0164] Example 2

[0165] According to this embodiment, the size of a template required for calculating the cost of template matching can be variably set, and information on the size of the template can be signaled. For example, the information on the size of the template can be transmitted from the image encoding device to the image decoding device using a higher level syntax such as a sequence parameter set (SPS) or a picture parameter set (PPS).

[0166] The availability of the motion vector correction (improvement) method based on template matching can be transmitted at a higher level such as SPS, for example, as shown in Table 5.

[0167] [Table 5]

[0168] In Table 5, sps_dmvd_enabled_flag may be information indicating whether dmvd is available in a coded layer video sequence (CLVS) referring to the SPS. The dmvd may mean decoder-side motion vector derivation, which may mean a technique for correcting a motion vector in a decoder based on a template matching cost. According to this embodiment, if the dmvd is available, information on the size of a template for calculating a template matching cost may be additionally signaled, for example, as shown in Table 6.

[0169] [Table 6]

[0170] In Table 6, sps_dmvd_tm_size_idx may be information (index) on the size of a template used in a motion vector correction technique based on template matching. The size of the template may be derived according to the index and Table 7.

[0171] [Table 7]

[0172] Table 7 may be predefined in the image encoding device and the image decoding device. The derivation of the template size by the above method is merely an example, and for example, the sps_dmvd_tm_size_idx value may be input into a predetermined formula to derive the template size as the resultant value. For example, the predetermined formula may include left-shifting.

[0173] The ranges of syntax values ​​and / or template sizes are merely examples and are not limited to the above examples, and various other combinations are possible.

[0174] Furthermore, the information regarding the template size can be transmitted at any other higher level, such as a picture parameter set or a picture header, in addition to the above-mentioned SPS.

[0175] Example 3

[0176] According to this embodiment, information on the template size can be commonly signaled for the techniques based on template matching cost. For example, the information on the template size can be transmitted from the image encoding device to the image decoding device using a higher level syntax such as a sequence parameter set (SPS) or a picture parameter set (PPS).

[0177] According to this embodiment, information on the template size for calculating a common template matching cost for the template matching cost based techniques can be signaled as shown in Table 8, for example.

[0178] [Table 8]

[0179] In Table 8, sps_tm_size_idx may be information (index) on the size of a template commonly used in techniques based on template matching. The size of the template can be derived according to the index and Table 9.

[0180] [Table 9]

[0181] Table 9 may be predefined in the image encoding device and the image decoding device. The derivation of the template size by the above method is merely an example, and for example, the sps_tm_size_idx value may be input into a predetermined formula, and the template size may be derived as the resultant value. For example, the predetermined formula may include left-shifting.

[0182] The ranges of syntax values ​​and / or template sizes are merely examples and are not limited to the above examples, and various other combinations are possible.

[0183] Furthermore, the information regarding the template size can be transmitted at any other higher level, such as a picture parameter set or a picture header, in addition to the above-mentioned SPS.

[0184] This embodiment relates to signaling of information on template size commonly applied to all technologies based on template matching. However, as a modified example, when different template sizes are used for each technology, information on the template size for each technology can be separately signaled by additionally transmitting a separate syntax as in the first or second embodiment. When a separate syntax is additionally transmitted as described above, the template size for the corresponding technology can be induced by the size information separately signaled for the corresponding technology instead of the information on the common template size.

[0185] As shown in Tables 4, 7, and 9, the size of the template may be composed of only squared values. In this case, the size of the template may be derived from information on the size of the template in the form of a log value of squared or a shift value of squared without table mapping. However, the present invention is not limited thereto, and in the first to third embodiments, the size of the template may have a value other than squared values ​​of squared values. The size of the template may be different depending on the encoding / decoding tool. For example, the size of the template used for template matching-based motion vector compensation may be different from the size of the template used for template matching-based merge candidate rearrangement. The size of the template may be defined to be different depending on the size or shape of the block. For example, the size of the template for a block larger than a critical size may be different from the size of the template for a block smaller than a critical size. The size of the template when the width-to-height ratio is larger than or smaller than a threshold value may be different from the size of the template when the width-to-height ratio is not larger than or smaller than a threshold value.

[0186] Example 4

[0187] According to this embodiment, a cost function used in calculating the cost of template matching can be variably set, and information on the cost function can be signaled. For example, the information on the cost function can be transmitted from the image encoding device to the image decoding device using a higher level syntax such as a sequence parameter set (SPS) or a picture parameter set (PPS).

[0188] The availability of the method of rearranging merging candidates based on template matching may be transmitted at a higher level such as SPS as shown in Table 2. Table 2 has been described in the first embodiment, so a duplicate description will be omitted.

[0189] According to this embodiment, if the AML is available, information regarding a cost function for calculating the template matching cost can be additionally signaled, for example as shown in Table 10.

[0190] [Table 10]

[0191] In Table 10, sps_aml_cost_function_idx may be information (index) about a cost function used in a merge candidate list rearrangement technique based on template matching. A cost function can be derived according to the index and Table 11.

[0192] [Table 11]

[0193] The above-mentioned Table 11 may be predefined in the image encoding device and the image decoding device. The above-mentioned syntax value ranges and / or cost functions are merely examples, and are not limited to the above-mentioned examples, and various other combinations are possible.

[0194] Also, the information about the cost function can be transmitted at any other higher level, such as in the picture parameter set or the picture header, besides the above mentioned SPS.

[0195] Example 5

[0196] According to this embodiment, a cost function used in calculating the cost of template matching can be variably set, and information on the cost function can be signaled. For example, the information on the cost function can be transmitted from the image encoding device to the image decoding device using a higher level syntax such as a sequence parameter set (SPS) or a picture parameter set (PPS).

[0197] Whether or not the motion vector correction (improvement) method based on template matching can be used can be transmitted at a higher level such as SPS as shown in Table 5. Table 5 has been described in the second embodiment, so a duplicate description will be omitted.

[0198] According to this embodiment, if the dmvd is available, information regarding a cost function for calculating the template matching cost can be additionally signaled, for example as shown in Table 12.

[0199] [Table 12]

[0200] In Table 12, sps_dmvd_cost_function_idx may be information (index) about a cost function used in a motion vector correction technique based on template matching. A cost function can be derived according to the index and Table 13.

[0201] [Table 13]

[0202] The above-mentioned Table 13 may be predefined in the image encoding device and the image decoding device. The above-mentioned syntax value ranges and / or cost functions are merely examples and are not limited to the above-mentioned examples, and various other combinations are possible.

[0203] Also, the information about the cost function can be transmitted at any other higher level, such as in the picture parameter set or the picture header, besides the above mentioned SPS.

[0204] Example 6

[0205] According to this embodiment, information on a cost function can be commonly signaled for techniques based on template matching cost. For example, the information on the cost function can be transmitted from an image encoding device to an image decoding device using a higher level syntax such as a sequence parameter set (SPS) or a picture parameter set (PPS).

[0206] According to this embodiment, information about a cost function for calculating the template matching cost, which is common to the techniques based on the template matching cost, can be signaled, for example, as shown in Table 14.

[0207] [Table 14]

[0208] In Table 14, sps_cost_function_idx may be information (index) about a cost function commonly used in techniques based on template matching. A cost function can be derived according to the index and the table in Table 15.

[0209] [Table 15]

[0210] The above-mentioned Table 15 may be predefined in the image encoding device and the image decoding device. The above-mentioned syntax value ranges and / or cost functions are merely examples and are not limited to the above-mentioned examples, and various other combinations are possible.

[0211] Also, the information about the cost function can be transmitted at any other higher level, such as in the picture parameter set or the picture header, besides the above mentioned SPS.

[0212] This embodiment relates to signaling information on a cost function commonly applied to all template matching based technologies. However, as a modified example, when different cost functions are used for each technology, information on the cost function for each technology can be separately signaled by additionally transmitting a separate syntax as in the fourth or fifth embodiment. When a separate syntax is additionally transmitted as described above, the cost function for the technology can be induced by the cost function information separately signaled for the technology instead of the information on the common cost function.

[0213] The cost function commonly described in the fourth to sixth embodiments will be described below.

[0214] When the template of the current block is curT and the template of the reference block specified by the motion vector is refT, the cost function can be defined as follows: where n can indicate the size or number of pixels of the template.

[0215] The sum of absolute difference (SAD) may indicate the sum of absolute values ​​of the difference values ​​between the template of the current block and the template of the reference block, and may be calculated according to Equation 1.

[0216] [Formula 1]

number

[0217] The sum of square error (SSE) may indicate the sum of squares of the difference between the template of the current block and the template of the reference block, and may be calculated by Equation 2.

[0218] [Formula 2]

number

[0219] The mean removal SAD (MR-SAD) may indicate the sum of absolute values ​​of difference values ​​obtained by subtracting the mean value from the template of the current block and the template of the reference block, and may be calculated by Equation 3.

[0220] [Formula 3]

number

[0221] During the ceremony, JPEG2025513819000020.jpg523 may indicate the template average of the current block and the template average of the reference block respectively.

[0222] HAD (Hadamard absolute difference) can be calculated using the existing Hadamard transform, and a detailed derivation process will be omitted.

[0223] Fig. 12 is a diagram for explaining an image decoding method for carrying out the first to sixth embodiments of the present disclosure. The image decoding method of Fig. 12 can be performed by an image decoding device. Depending on the embodiment, some of the steps included in Fig. 12 may be omitted, new steps may be added, or the order of execution may be changed. In Fig. 12, the template matching-based technology may include the above-mentioned rearrangement of merging candidates based on template matching, motion vector correction based on template matching, etc.

[0224] As shown in Fig. 12, in step S1210, information indicating whether the template matching based technology is available can be obtained. The template matching based technology availability information may include, for example, the above-mentioned sps_aml_enabled_flag or sps_dmvd_enabled_flag. It may be determined whether the template matching based technology is available based on the template matching based technology availability information obtained in step S1210 (S1220). Alternatively, if the template matching based technology is always available, steps S1210 and S1220 may be omitted.

[0225] If the template matching based technology is not available, the procedure of Fig. 12 related to the template matching based technology can end. If the template matching based technology is available, information about template matching can be obtained (S1230). The information about template matching can include, for example, information about the size of the template required for the template matching cost calculation described above, or information about the cost function used in the template matching cost calculation.

[0226] Thereafter, for a current block to which the template matching based technique is applied, a template matching cost may be calculated based on the information acquired in step S1230 (S1240). For example, a template matching cost may be derived for the current block using a cost function determined based on information on the template size and / or cost function derived based on information on the template size. Thereafter, a prediction for the current block may be made by performing the template matching based technique based on the derived template matching cost (S1250).

[0227] Fig. 13 is a diagram for explaining an image coding method for carrying out the first to sixth embodiments of the present disclosure. The image coding method of Fig. 13 can be performed by an image coding device. Depending on the embodiment, some of the steps included in Fig. 13 may be omitted, new steps may be added, or the order of execution may be changed. In Fig. 13, the template matching-based technology may include the above-mentioned rearrangement of merging candidates based on template matching, motion vector correction based on template matching, etc.

[0228] As shown in FIG. 13, in step S1310, it may be determined whether a template matching based technology is available. Based on the determination in step S1310, template matching based technology availability information may be encoded. The template matching based technology availability information may include, for example, the above-mentioned sps_aml_enabled_flag or sps_dmvd_enabled_flag. In step S1320, it is determined whether the template matching based technology is available. If the template matching based technology is not available, the template matching based technology availability information is encoded and the procedure of FIG. 13 may end. If the template matching based technology is always available, steps S1310 and S1320 may be omitted.

[0229] If template matching-based technology is available, information related to template matching can be determined (S1330). The information related to template matching can include, for example, information about the size of the template required for calculating the cost of template matching as described above, or information about a cost function used to calculate the cost of template matching.

[0230] Thereafter, for a current block to which the template matching based technique is applied, a template matching cost may be calculated based on the information determined in step S1330 (S1340). For example, a template matching cost may be derived for the current block using a cost function determined based on information on the template size and / or cost function derived based on information on the template size. Then, prediction for the current block may be performed by performing the template matching based technique based on the derived template matching cost (S1350). Also, template matching based technique availability information and information on template matching may be encoded (S1360).

[0231] According to the embodiment described with reference to the first to sixth embodiments of the present disclosure and Fig. 12 and Fig. 13, the template matching based technology can be adaptively applied. Therefore, it is possible to improve the coding efficiency by applying the template matching based technology, and thereby improve the image compression efficiency.

[0232] Example 7

[0233] According to the present embodiment, the position and / or size of the template required for calculating the cost of template matching can be variably set. For example, the position and / or size of the template can be variable based on the coding information of the current block or the size information of the current block.

[0234] When using the neighboring pixel information of the current block and the neighboring pixel information of the reference block to calculate the cost of template matching, one or a combination of two or more of the following examples (Examples 1 to 6) can be used. In the following, the "target block" can be the "current block" and / or the "reference block".

[0235] Example 1) All the surrounding pixels of the block are used, but it is possible to use more surrounding pixels than are in the block.

[0236] Example 2) All the surrounding pixels of the block are used, but the number of surrounding pixels can be the same as the number of pixels in the block.

[0237] Example 3) It is possible to use only a portion of one or more of the surrounding pixels of the block.

[0238] Example 4) Only a portion of one or more of the neighboring pixels of the block may be used, but in a subsampled form.

[0239] Example 5) Only a portion of one or more of the surrounding pixels of the block is used, but a portion of the surrounding pixels at any position can be used.

[0240] Example 6) A specific pixel that represents the surrounding pixels of the block can be selected and used. For example, two of the available surrounding pixels (the surrounding pixel with the maximum sample value and the surrounding pixel with the minimum sample value) can be used. In this case, pixels other than the selected specific pixel among the surrounding pixels may not be used.

[0241] The selection of the neighboring pixel information may be performed based on the width and / or height of the block. For example, the maximum available pixel size may be arbitrarily determined, and neighboring pixels may be used up to a certain width or height, and only some of the pixels may be selected and used when the width or height is greater than the maximum available pixel size.

[0242] As another example, a subsampling ratio can be determined so that some pixels are adaptively selected depending on the width or height, for example, if the subsampling ratio is 2 and the width or height is 4, only two surrounding pixels can be selected for use, and if the width or height is 8, only four surrounding pixels can be selected for use.

[0243] As another example, the selection of the neighboring pixel information may be performed based on an arbitrary threshold. For example, the values ​​of the neighboring pixels may be compared with a threshold, and only neighboring pixels having values ​​greater than the threshold may be selected for use. Alternatively, the values ​​of the neighboring pixels may be compared with a threshold, and only neighboring pixels having values ​​smaller than the threshold may be selected for use.

[0244] It will be apparent to those skilled in the art that the above methods may be used in combination.

[0245] The subsampling ratio and the arbitrary threshold may be predefined and used in the image encoding device and the image decoding device, or the subsampling ratio and the arbitrary threshold may be included in a specific header (SPS, PPS, Picture Header, Slice Header, etc.) and transmitted (signaled) from the image encoding device to the image decoding device.

[0246] 14 is a diagram for explaining various examples of templates that can be selected according to an embodiment of the present disclosure. The example shown in FIG. 14 may correspond to any of the above-mentioned examples 1 to 6.

[0247] In Fig. 14, the shaded portions are examples of templates required for calculating the cost of template matching. As shown in Fig. 14, the position and / or size of the template can be variably set.

[0248] According to another example of the present disclosure, when selecting the neighboring pixel information, one or a combination of two or more of the following examples (Example 7 to Example 12) may be applied. Hereinafter, the "target block" may be a "current block" and / or a "reference block".

[0249] Example 7) The neighboring pixels can be selected from one reference line adjacent to the block.

[0250] Example 8) The neighboring pixels can be selected from one reference line that is not adjacent to the block.

[0251] Example 9) The surrounding pixels may be selected from an area consisting of two or more reference lines adjacent to the block.

[0252] Example 10) The neighboring pixels may be selected from an area consisting of one or more reference lines adjacent to the left side of the block.

[0253] Example 11) The neighboring pixels can be selected from a region formed by one or more reference lines adjacent to the top edge of the block.

[0254] Example 12) The surrounding pixels may be selected from a region formed by combining two or more of the methods of Examples a) to e).

[0255] 15 is a diagram for explaining various examples of templates that can be selected according to another embodiment of the present disclosure. The example shown in FIG. 15 may correspond to any of the above-mentioned Examples 7 to 11.

[0256] In Fig. 15, the shaded portions are examples of templates required for calculating the cost of template matching. As shown in Fig. 15, the position and / or size of the template can be variably set.

[0257] In the method for selecting the area to which the neighboring pixel information belongs, described with reference to FIG. 15, the width and height of a block can be used as criteria. According to the present disclosure, for example, the minimum number of pixels required to derive the template matching cost can be arbitrarily determined. In this case, if a block has a specific width or height that can secure the minimum number of pixels even in the area of ​​only one reference line, the neighboring pixels for the block can be selected from the area of ​​one reference line that is adjacent or not adjacent to the block. However, if not, the neighboring pixels can be selected from an area composed of two or more reference lines.

[0258] According to another example of the present disclosure, an arbitrary threshold can be used as a method for selecting the surrounding pixel information. For example, the values ​​of the surrounding pixels can be compared with a threshold, and only the surrounding pixels whose values ​​are greater than the threshold can be selectively used. Alternatively, the values ​​of the surrounding pixels can be compared with a threshold, and only the surrounding pixels whose values ​​are smaller than the threshold can be selectively used. In this case, if the minimum number of pixels required for template matching can be secured from one reference line region, which may be adjacent or not adjacent, the surrounding pixels can be selected only from the region formed by one reference line. Otherwise, the surrounding pixels required for calculating the template matching cost can be secured from the region formed by multiple reference lines.

[0259] The information for selecting the region to which the neighboring pixels belong may be predefined in the image encoding device and the image decoding device. Alternatively, the information for selecting the region to which the neighboring pixels belong may be included in a specific header (SPS, PPS, Picture Header, Slice Header, etc.) and transmitted from the image encoding device to the image decoding device. Alternatively, the region to which the neighboring pixels belong may be derived by a method predefined in the image encoding device and the image decoding device based on the information signaled in a specific header and the size (width and / or height) of the block.

[0260] According to another example of the present disclosure, when rearranging merging candidates based on template matching, the position of the template for calculating the template matching cost may be determined based on the position of the merging candidate. For example, if the merging candidate is a left block of the current block, the template matching cost may be calculated using only the left template. Also, if the merging candidate is an upper block of the current block, the template matching cost may be calculated using only the upper template.

[0261] According to another example of the present disclosure, when rearranging merge candidates based on template matching, the position and / or size of the template for calculating a template matching cost may be adaptively determined according to the type of merge candidate, where the type of merge candidate may mean a spatial motion vector candidate, a temporal motion vector candidate, a non-adjacent spatial motion vector candidate, a history-based motion vector candidate, etc.

[0262] In the above example, when selecting a template to rearrange merging candidates based on template matching, the number of pixels used in the template cost may differ for each merging candidate. According to the present disclosure, in order to make a fair comparison of template costs between merging candidates, the template cost may be normalized by the number of pixels used in the template cost.

[0263] According to another embodiment of the present disclosure, the template for calculating the template matching cost may be adaptively determined based on the direction of the motion vector. More specifically, the template may be selected based on the difference in magnitude between the motion vector in the x-axis direction and the motion vector in the y-axis direction. Alternatively, the size of the template may be adjusted based on the magnitude of the absolute value of the motion vector.

[0264] According to the present embodiment, it is possible to adaptively / variably select the position and / or size of the template, which may allow for more accurate calculation of the template matching cost, thereby improving the efficiency of image coding based on the template matching cost.

[0265] Various examples described in this embodiment may be combined for use. For example, the method of selecting neighboring pixels according to (c) of Fig. 14 and the method of selecting neighboring pixels according to (b) of Fig. 15 may be combined, in which case one or more neighboring pixels may be selected from a reference line that is not adjacent to the block.

[0266] Alternatively, at least two of the surrounding pixel selection methods described in this embodiment can be used, and the template matching costs calculated based on each selection method can be weighted together to calculate a final cost, and a template matching cost-based technique can be performed based on the final cost.

[0267] Alternatively, any of the methods of selecting neighboring pixels described in this embodiment may be selected based on the size and / or shape of the current block. For example, if the width-to-height ratio of the current block is greater than or less than a threshold, neighboring pixels may be selected from the left region of FIG. 15(d) or the top region of FIG. 15(e).

[0268] Example 8

[0269] According to this embodiment, the cost of template matching can be derived (adjusted) by a weighted sum of the cost of the template matching and the spatial similarity cost of the template.

[0270] For example, if the similarity between the template of the reference block and the pixels of the reference block is low or has other statistical characteristics, rearrangement of merge candidates based on template matching may not help improve coding efficiency or may even result in degradation of coding performance. According to the present disclosure, in order to improve the accuracy of the template matching cost, the final template cost can be calculated by additionally using a spatial similarity cost to the existing template cost.

[0271] The spatial similarity cost can be calculated based on the difference between the template of the reference block and the pixel values ​​in the reference block adjacent to the template. The spatial similarity of the reference template can also be expressed as a reliability for the reference template.

[0272] FIG. 16 is a diagram illustrating an example of a template of a reference block and boundary pixels of the reference block for calculating a spatial similarity cost.

[0273] When the reference template of the reference block identified by the motion vector is refT and the boundary pixel in the reference block adjacent to the reference template is refB, the spatial similarity cost can be calculated, for example, as shown in Equation 4.

[0274] [Formula 4]

number

[0275] In the formula, n can indicate the size (or number of pixels) of the template.

[0276] Equation 4 is a cost calculation formula based on SAD similar to Equation 1. However, the method of calculating the spatial similarity cost is not limited to Equation 4, and it is obvious to those skilled in the art that various methods based on the difference between refT and refB (e.g., Equations 1 to 3) can be used.

[0277] The final cost of template matching can be derived by a weighted sum of the template cost and the spatial similarity cost, for example, as follows:

[0278] Final cost = a*template_cost + b*spatial_similarity_cost

[0279] In the above, the template cost may be derived by any of the template cost functions described in the various embodiments above, and the spatial similarity cost may be derived according to the method described in this embodiment. In the above, the weight a of the template cost and the weight b of the spatial similarity cost may be explicitly signaled by being included in a predetermined header (e.g., SPS, PPS, Picture Header, Slice Header, etc.). Alternatively, the weight may be derived to a value predefined in the image encoding device and the image decoding device based on the size of the current block and / or the coding information of the current block. Alternatively, the weight may be derived by a combination of the size of the current block and / or the coding information of the current block and information signaled via the bitstream.

[0280] According to this embodiment, by adjusting the cost of template matching taking into account the spatial similarity of the reference template, it is possible to calculate a more accurate template matching cost, thereby making it possible to improve the efficiency of image coding based on the template matching cost.

[0281] Example 9

[0282] According to this embodiment, the cost function of template matching can have different weights based on the magnitude of the motion vector, the position of the motion vector, and / or the type of the motion vector.

[0283] According to the seventh embodiment, the template direction or the template size can be adaptively selected according to the type of the merge candidate. Similarly, according to the present embodiment, the weight of the template matching cost can be adjusted according to the type of the merge candidate. More specifically, different weights can be applied to the template matching cost according to the type of the merge candidate, such as a spatial motion vector candidate, a temporal motion vector candidate, a non-adjacent spatial motion vector candidate, a history-based motion vector candidate, etc., as follows:

[0284] Final cost=a merge_type *Template_cost

[0285] In the above, merge_type denotes a weight according to the type of the corresponding merge candidate, and the template cost can be derived by one of the template cost functions described in the various embodiments above.

[0286] For example, if the weight of a spatial motion vector candidate is set lower than the weight of a non-adjacent spatial motion vector candidate, the final template cost for the spatial motion vector candidate can be relatively reduced. Thus, the spatial motion vector candidate can be adjusted to be selected more than the non-adjacent spatial motion vector candidate. That is, the compression efficiency can be improved by adjusting the selection rate of a particular merge candidate by applying a weight value according to the type of merge candidate according to the statistical characteristics of the image. The weight according to the type of merge candidate can be explicitly signaled by being included in a predetermined header (SPS, PPS, Picture Header, Slice Header, etc.), or can be induced to a value predefined in the image encoding device and the image decoding device based on the size of the current block and / or the coding information of the current block. Alternatively, the weight can be induced by a combination of the size of the current block and / or the coding information of the current block and information signaled via a bitstream.

[0287] Also, according to the seventh embodiment, when performing template matching-based motion vector compensation, the direction and / or size of the template can be selected according to the magnitude of the motion vector or the difference between the x-direction component and the y-direction component of the motion vector. Similarly, according to other embodiments of the present disclosure, the weight can be adjusted based on the magnitude of the motion vector. More specifically, different weights can be applied according to the magnitude of the absolute value of the motion vector or the difference between the x-direction component and the y-direction component of the motion vector, as follows:

[0288] Final cost=a abs_mv *Template_cost

[0289] In the above, abs_mv means a weight according to the magnitude of the absolute value of a motion vector or a difference between the x-direction component and the y-direction component of a motion vector, and the template cost can be derived by one of the template cost functions described in the various embodiments above.

[0290] For example, if the magnitude of the absolute value of the motion vector is small, setting the weight low has the effect of reducing the final template matching cost. That is, by applying different weights according to the magnitude of the motion vector according to the statistical characteristics of the image, it is possible to adjust so that the motion vector indicating the reference block at a position more similar to the position of the current block is selected. The weight according to the magnitude of the absolute value of the motion vector or the difference between the x-direction component and the y-direction component of the motion vector may be included in a predetermined header (SPS, PPS, Picture Header, Slice Header, etc.) and explicitly signaled, or may be induced to a value predefined in the image encoding device and the image decoding device based on the size of the current block and / or the coding information of the current block. Alternatively, the weight may be induced by a combination of the size of the current block and / or the coding information of the current block and information signaled through the bitstream.

[0291] This embodiment can be combined with Example 8 of the present disclosure to derive the final cost, for example, as follows:

[0292] Final cost=a merge_type * template_cost + b * spatial_similarity_cost

[0293] In the above merge_type means the weight according to the type of the merge candidate, which is the weight according to the absolute value of the motion vector or the difference between the x-direction component and the y-direction component of the motion vector depending on the technology. abs_mv In addition, the b represents a weight for the spatial similarity cost.

[0294] In the above, the template cost can be derived by one of the template cost functions described in the various embodiments above, and the spatial similarity cost can be derived by the method described in the eighth embodiment of the present disclosure. abs_mvThe weights a and b may be explicitly signaled by being included in a predetermined header (e.g., SPS, PPS, Picture Header, Slice Header, etc.), or the weights may be derived to predefined values ​​in the image encoding device and the image decoding device based on the size of the current block and / or the coding information of the current block, or the weights may be derived by a combination of the size of the current block and / or the coding information of the current block and information signaled via the bitstream.

[0295] Two or more of the various embodiments described in the present disclosure may be combined. For example, a combination of the third embodiment and the seventh embodiment may lead to an embodiment in which different template sizes are used depending on the position of the template region or the size / shape of the block. Or, a combination of the sixth embodiment and the seventh embodiment may lead to an embodiment in which different cost functions are applied to the same (or different) template region and then weighted sum is calculated to calculate the final cost. Or, a combination of the third embodiment and the eighth embodiment may lead to an embodiment in which the size of the top template or the left template is variably determined based on the reliability cost. Or, a combination of the seventh embodiment and the eighth embodiment may lead to an embodiment in which the size and / or position of the template region is determined based on the reliability cost. Furthermore, examples of combinations of two or more embodiments are not limited to the above examples, except when combinations between the embodiments are not possible.

[0296] According to the present disclosure, it becomes possible to calculate more sophisticated template matching costs depending on the type of merging candidate, the magnitude of the motion vector, etc., which is expected to improve the efficiency of encoding based on template matching.

[0297] FIG. 17 is a diagram for explaining an image decoding method according to at least one of the embodiments of the present disclosure.

[0298] The image decoding method of FIG. 17 can be performed in an image decoding device.

[0299] When the template matching cost-based technique is applied to the current block, the image decoding apparatus may determine various information for calculating the template matching cost (S1710). The various information in step S1710 may include not only the template size and cost function described above, but also the template position, threshold, method of selecting neighboring pixel information, region to which the neighboring pixels belong, various weights, and final cost function described in the seventh embodiment of the present disclosure, and may include at least one of all information for calculating the template matching cost according to at least one of the various embodiments of the present disclosure. The determination of the various information in step S1710 may be performed based on various conditions such as information signaled through a predetermined header in the bitstream, information already defined in the image encoding apparatus and the image decoding apparatus, size (width and / or height) of the block, coding information of the block, type of merge candidate, and magnitude of the motion vector, as described in the embodiments of the present disclosure.

[0300] Once the information for calculating the template matching cost is determined in step S1710, the template matching cost may be calculated based on the determined information (S1720). According to one embodiment, the template matching cost may be adjusted by a weighted sum with the spatial similarity cost. According to another embodiment, the template matching cost may be weighted based on the type of merge candidate, the magnitude of the motion vector, etc. According to another embodiment, the template matching cost may be weighted based on the type of merge candidate, the magnitude of the motion vector, etc., and may be adjusted by a weighted sum with the spatial similarity cost.

[0301] When the template matching cost is finally calculated in step S1720, a template matching based technique may be performed based on the calculated template matching cost (S1730), as described with reference to FIGS. 12 and 13.

[0302] Also, the method of FIG. 17 may be performed as a part of an image encoding method in an image encoding device. That is, the description of FIG. 17 regarding the image decoding method can be applied to the image encoding method in the same manner. Therefore, the image encoding device determines information on template matching based on the above-mentioned various conditions (S1710), calculates a template matching cost based on the determined information (S1720), and then performs a template matching-based technology based on the calculated template matching cost (S1730). However, the image encoding device does not need to use information signaled through a predetermined header in a bitstream to determine information on template matching. Instead, since the image encoding device is a device that generates a bitstream, information that needs to be transmitted to an image decoding device among the above-mentioned various conditions for determining information on template matching can be encoded in a predetermined header of the bitstream.

[0303] According to the present disclosure, the cost calculation of template matching can be adaptively made more accurate, and therefore, it is expected to have an effect of improving the coding efficiency of template matching-based techniques.

[0304] FIG. 18 is a diagram illustrating an example of a content streaming system to which an embodiment of the present disclosure can be applied.

[0305] As shown in FIG. 18, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.

[0306] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, video camera, etc. directly generates a bitstream, the encoding server can be omitted.

[0307] The bitstream may be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0308] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit the multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server can control commands / responses between devices in the content streaming system.

[0309] The streaming server may receive the content from a media storage and / or an encoding server. For example, when receiving the content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time in order to provide a smooth streaming service.

[0310] Examples of the user devices include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices such as smartwatches, smart glass, head mounted displays (HMDs), digital TVs, desktop computers, and digital signage.

[0311] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0312] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of the various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]

[0313] The embodiments of the present disclosure can be used to encode / decode images.

Claims

1. An image decoding method performed by an image decoding device, The steps include determining information regarding template matching for the current block on which the template matching infrastructure technology is being applied, A step of deriving a template matching value based on the information determined regarding template matching, An image decoding method comprising the step of performing the template matching base technology based on the template matching value.

2. The image decoding method according to claim 1, wherein the information relating to template matching includes information relating to the size of a template for the template matching value, or information relating to a difference-based function for the template matching value.

3. The image decoding method according to claim 2, wherein the information relating to the size of the template, or the information relating to the difference-based function, is included in the higher level of the current block and signaled.

4. The image decoding method according to claim 2, wherein the information relating to the size of the template or the information relating to the difference-based function varies depending on the template matching base technology.

5. The image decoding method according to claim 1, wherein the template for the template matching value is adaptively determined based on the encoding information of the current block or the size of the current block.

6. The image decoding method according to claim 5, wherein the template for the template matching value is adaptively determined based on a comparison of the size of the current block with a predetermined threshold.

7. The image decoding method according to claim 5, wherein the template for the template matching value is adaptively determined based on the position or type of the merge candidate that is the subject of the template matching.

8. The image decoding method according to claim 5, wherein the template for weights is adaptively determined based on the magnitude or direction of the motion vector that is the subject of the template matching.

9. The template matching value is derived by a weighted sum of the template matching value and the spatial similarity value, The image decoding method according to claim 1, wherein the spatial similarity value is derived based on the difference between a reference template adjacent to a reference block and the pixel value of a pixel in the reference block adjacent to the reference template.

10. The image decoding method according to claim 9, wherein the weights of the weighted sum are included in the higher level of the current block and signaled, or are derived based on the size of the current block or the encoding information of the current block.

11. The image decoding method according to claim 1, wherein the template matching value is adjusted to a value obtained by applying a weight to the template matching value.

12. The image decoding method according to claim 11, wherein the weight is adaptively determined based on the position or type of the merge candidate that is the target of the template matching value.

13. The image decoding method according to claim 11, wherein the weight is adaptively determined based on the magnitude or direction of the motion vector that is the target of the template matching value.

14. An image encoding method performed by an image encoding device, The steps include determining information regarding template matching for the current block on which the template matching infrastructure technology is being applied, A step of deriving a template matching value based on the information determined regarding template matching, An image encoding method comprising the step of performing the template matching base technology based on the template matching value.

15. In a method for transmitting a bitstream generated by an image encoding method, the image encoding method is: The steps include determining information regarding template matching for the current block on which the template matching infrastructure technology is being applied, A step of deriving a template matching value based on the information determined regarding template matching, A method comprising the step of performing the template matching base technology based on the template matching values.

16. A non-temporary computer-readable recording medium for storing a bitstream generated by the image encoding method described in claim 14.