Intra prediction mode-based image encoding / decoding method and apparatus using multiple reference lines (mrl), and recording medium for storing bitstream
By using the multi-reference line (MRL) intra prediction mode in image encoding/decoding, weighted prediction blocks are generated, and the problem of low high-resolution and high-quality image encoding/decoding efficiency is solved, and more efficient image processing and storage is achieved.
Patent Information
- Application Number
- CN202380074291.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-20
- Filing Date
- 2023-09-20
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively improve the encoding/decoding efficiency of high resolution and high-quality images, resulting in increased transmission and storage costs.
Using a multi-reference line (MRL)-based intra prediction mode, the encoding/decoding efficiency is improved by determining the weighting sum of multiple reference samples.
Improves image encoding/decoding efficiency, reduces transmission and storage costs, and supports efficient processing of high-resolution and high-quality images.
Smart Images

Figure CN120019644A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method, an apparatus, and a recording medium for storing a bit stream, and more particularly, to an image encoding / decoding method and an apparatus based on an intra-frame prediction mode using a multiple reference line (MRL), and a recording medium for storing a bit stream generated by the image encoding method / apparatus of the present disclosure. Background Art
[0002] Recently, the demand for high-resolution and high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images has increased in various fields. As image data becomes high-resolution and high-quality, the amount of information or bit rate transmitted increases relative to conventional image data. The increase in the amount of information or bit rate transmitted leads to an increase in transmission cost and storage cost.
[0003] Therefore, an efficient image compression technology is needed to effectively transmit, store and reproduce information of high-resolution and high-quality images. Summary of the invention
[0004] Technical issues
[0005] The present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0006] The present disclosure is to provide an image encoding / decoding method and apparatus for performing an intra-frame prediction mode.
[0007] The present disclosure is to provide an image encoding / decoding method and apparatus for performing an intra prediction mode using a multiple reference line (MRL).
[0008] The present disclosure is to provide an image encoding / decoding method and apparatus that fuses multiple prediction blocks generated using MRL.
[0009] The present disclosure is to provide an image encoding / decoding method and apparatus for performing a planar mode using MRL.
[0010] The present disclosure is to provide an image encoding / decoding method and apparatus for generating samples using a planar pattern within a template region of a template-based intra mode derivation (TIMD) mode.
[0011] The present disclosure is to provide a non-transitory computer-readable recording medium for storing a bit stream generated by an image encoding method or apparatus according to the present disclosure.
[0012] The present disclosure is to provide a non-transitory computer-readable recording medium for storing a bit stream received and decoded by an image decoding apparatus according to the present disclosure and used for image reconstruction.
[0013] The present disclosure is to provide a method for transmitting a bit stream generated by an image encoding method or apparatus according to the present disclosure.
[0014] Technical problems to be achieved in the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not described can be clearly understood by ordinary technicians in the art from the following description.
[0015] Technical Solution
[0016] According to an embodiment of the present disclosure, a method for decoding an image performed by an image decoding device includes the following steps: determining a reference sample that is n samples away from a current block and a reference sample that is m samples away from the current block (the reference sample that is n samples away includes a top reference sample and an upper right reference sample, and the reference sample that is m samples away includes a left reference sample and a lower left reference sample), and generating a prediction block of the current block based on a weighted sum of at least two reference samples among the determined reference samples, wherein the weight used in the weighted sum is determined based on n or m, and n and m can be natural numbers.
[0017] According to an embodiment of the present disclosure, generating a prediction block of a current block may include: generating the prediction block of the current block based on a weighted sum of reference samples, and generating a final prediction block of the current block based on the prediction block.
[0018] According to an embodiment of the present disclosure, the reference sample may further include: a reference sample generated by modifying a reference sample that is n samples away based on n, and a reference sample generated by modifying a reference sample that is m samples away based on m.
[0019] According to an embodiment of the present disclosure, the reference sample may further include a reference sample generated in a plane mode based on at least two of a top reference sample, a left reference sample, an upper right reference sample, a lower left reference sample, an upper left reference sample adjacent to the current block, a reference sample adjacent to the lower left reference sample, or a reference sample adjacent to the upper right reference sample, wherein the plane mode may include a horizontal plane mode and a vertical plane mode.
[0020] According to an embodiment of the present disclosure, the reference sample may further include a reference sample generated by averaging sample values of multiple reference samples generated in a plane mode based on at least two of a top reference sample, a left reference sample, an upper right reference sample, a lower left reference sample, an upper left reference sample adjacent to the current block, a reference sample adjacent to the lower left reference sample, or a reference sample adjacent to the upper right reference sample, wherein the plane mode may include a horizontal plane mode and a vertical plane mode.
[0021] According to an embodiment of the present disclosure, an image decoding method performed by an image decoding device includes the following steps: deriving a first prediction mode and a second prediction mode for a current block based on a template matching cost (the template matching cost calculated based on the first prediction mode is less than the template matching cost calculated based on the second prediction mode); generating a first prediction block based on the first prediction mode and a first reference sample line and generating a second prediction block based on the second prediction mode and the second reference sample line; and generating a final prediction block of the current block based on the first prediction block and the second prediction block, wherein at least one of the first reference sample line or the second reference sample line can be determined based on whether the first prediction mode or the second prediction mode is a planar mode.
[0022] According to an embodiment of the present disclosure, based on the first prediction mode being the planar mode, the first reference sample line may be determined as the first reference sample line adjacent to the current block, and the second reference sample line may be determined as the rth reference sample line adjacent to the current block.
[0023] According to an embodiment of the present disclosure, based on the second prediction mode being the planar mode, the first reference sample line may be determined as the rth reference sample line adjacent to the current block, and the second reference sample line may be determined as the first reference sample line adjacent to the current block.
[0024] According to an embodiment of the present disclosure, based on a first prediction mode and a second prediction mode including a planar mode and a non-planar mode, a final prediction block can be generated by a weighted sum of a first prediction block generated in the planar mode based on a first reference sample line adjacent to a current block, a second prediction block generated in the planar mode based on an rth reference sample line adjacent to the current block, and a third prediction block generated in the non-planar mode based on the rth reference sample line.
[0025] According to an embodiment of the present disclosure, a template matching cost is calculated based on a prediction block caused by predicting a template area of a current block in a planar mode, wherein a top reference sample and an upper right reference sample used for predicting the template area can be obtained from a first reference sample line adjacent to the top template area, and a left reference sample and a lower left reference sample used for predicting the template area can be obtained from a first reference sample line adjacent to the left template area.
[0026] According to an embodiment of the present disclosure, the template area may include a top template area and a left template area, wherein the lower left reference sample used to predict the top template area may be the same as the lower left reference sample of the left template area, and the upper right reference sample used to predict the left template area may be the same as the upper right reference sample of the top template area.
[0027] According to an embodiment of the present disclosure, the template area may include a top template area and a left template area, wherein the lower left reference sample used to predict the top template area may have the same y coordinate as the lower left sample of the top template area, and the upper right reference sample used to predict the left template area may have the same x coordinate as the upper right sample of the left template area.
[0028] According to an embodiment of the present disclosure, the value of the left reference sample and the value of the lower left reference sample used to predict the top template area can be modified based on the width of the left template area, and the value of the top reference sample and the value of the upper right reference sample used to predict the left template area can be modified based on the height of the top template area.
[0029] According to an embodiment of the present disclosure, an image encoding method performed by an image encoding device includes: deriving a first prediction mode and a second prediction mode for a current block based on a template matching cost, generating a first prediction block based on the first prediction mode and a first reference sample line, generating a second prediction block based on the second prediction mode and the second reference sample line, and generating a final prediction block of the current block based on the first prediction block and the second prediction block, wherein at least one of the first reference sample line or the second reference sample line can be determined based on whether the first prediction mode or the second prediction mode is a planar mode.
[0030] According to an embodiment of the present disclosure, a computer-readable recording medium may store a bit stream generated by an image encoding method.
[0031] According to an embodiment of the present disclosure, in a method for sending a bit stream generated by an image encoding method, the image encoding method includes: deriving a first prediction mode and a second prediction mode for a current block based on a template matching cost, generating a first prediction block based on the first prediction mode and a first reference sample line, generating a second prediction block based on the second prediction mode and the second reference sample line, and generating a final prediction block of the current block based on the first prediction block and the second prediction block, wherein at least one of the first reference sample line or the second reference sample line can be determined based on whether the first prediction mode or the second prediction mode is a planar mode.
[0032] Beneficial Effects
[0033] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0034] According to the present disclosure, an image encoding / decoding method and apparatus for performing an intra prediction mode may be provided.
[0035] According to the present disclosure, an image encoding / decoding method and apparatus for performing an intra prediction mode using a multiple reference line (MRL) may be provided.
[0036] According to the present disclosure, an image encoding / decoding method and apparatus for fusing a plurality of prediction blocks generated using MRL may be provided.
[0037] According to the present disclosure, a method and apparatus for performing image encoding / decoding of a planar mode using MRL may be provided.
[0038] According to the present disclosure, an image encoding / decoding method and apparatus for generating samples in a planar mode within a template region of a template-based intra mode derivation (TIMD) mode may be provided.
[0039] According to the present disclosure, a non-transitory computer-readable recording medium for storing a bit stream generated by the image encoding method or apparatus according to the present disclosure may be provided.
[0040] According to the present disclosure, there may be provided a non-transitory computer-readable recording medium for storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used for image reconstruction.
[0041] According to the present disclosure, a method of transmitting a bit stream generated by the image encoding method or apparatus according to the present disclosure may be provided.
[0042] Effects obtainable by the present disclosure are not limited to the above-described effects, and other effects that are not described can be clearly understood by a person of ordinary skill in the art from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram showing a video coding system to which an embodiment of the present disclosure can be applied is shown.
[0044] Figure 2 A schematic diagram showing an image encoding device to which an embodiment of the present disclosure can be applied.
[0045] Figure 3 A schematic diagram showing an image decoding device to which an embodiment of the present disclosure can be applied is shown.
[0046] Figure 4 A flow chart showing a video / image encoding method based on intra-frame prediction.
[0047] Figure 5 An example diagram showing a configuration of an intra predictor according to the present disclosure.
[0048] Figure 6 A flow chart showing a video / image decoding method based on intra-frame prediction.
[0049] Figure 7 An example diagram showing a configuration of an intra predictor according to the present disclosure.
[0050] Figure 8 and Fig. 9 A diagram showing reference sample lines used for multiple reference line (MRL) based intra prediction according to the present disclosure.
[0051] Fig.10 Diagram showing template regions and reference samples used for template-based intra mode derivation (TIMD) according to the present disclosure.
[0052] Fig.11 A diagram illustrating a method of constructing a Histogram of Gradients (HoG) in a decoder-side Intra Mode Derivation (DIMD) mode according to the present disclosure.
[0053] Fig.12 A diagram illustrating a method of constructing a prediction block when a decoder-side intra mode derivation (DIMD) mode is applied according to the present disclosure.
[0054] Fig.13 A diagram showing an MRL-based intra prediction method according to an embodiment of the present disclosure.
[0055] Fig.14 A diagram showing reference samples used to generate prediction samples according to the present disclosure.
[0056] Fig.15 A diagram illustrating reference samples used to generate prediction samples according to another embodiment of the present disclosure.
[0057] Fig.16 A diagram illustrating reference samples used to generate prediction samples according to another embodiment of the present disclosure.
[0058] Fig.17 A diagram illustrating modified reference samples used to generate prediction samples according to an embodiment of the present disclosure.
[0059] Fig.18 A diagram illustrating modified reference samples used to generate prediction samples according to another embodiment of the present disclosure.
[0060] Fig.19 A flowchart of an image encoding / decoding method according to the present disclosure is shown.
[0061] Fig. 20 A diagram showing reference sample lines used to generate a prediction block according to the present disclosure.
[0062] Fig.21 A diagram illustrating reference samples used to generate a prediction block according to a template-based intra mode derivation (TIMD) mode according to an embodiment of the present disclosure.
[0063] Fig. 22A diagram illustrating reference samples used to generate a prediction block according to a template-based intra mode derivation (TIMD) mode according to another embodiment of the present disclosure.
[0064] Fig.23 A diagram illustrating reference samples used to generate a prediction block according to a template-based intra mode derivation (TIMD) mode according to another embodiment of the present disclosure.
[0065] Fig.24 A flowchart of an image encoding / decoding method according to the present disclosure is shown.
[0066] Fig.25 An exemplary diagram showing a content streaming system to which embodiments of the present disclosure can be applied. DETAILED DESCRIPTION
[0067] Hereinafter, in order for those skilled in the art to easily implement them, embodiments of the present disclosure will be described in detail by referring to the accompanying drawings. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.
[0068] When describing the embodiments of the present disclosure, when well-known configurations or functions are considered to obscure the main points of the present disclosure, their detailed explanation will be omitted. In addition, parts irrelevant to the description of the present disclosure are omitted from the drawings, and similar reference numerals have been assigned to similar parts.
[0069] In the present disclosure, when some components are described as being "connected," "coupled," or "linked" to another component, this may include not only a direct connection but also an indirect connection with another component in between. In addition, when a certain component is described as "including" or "having" another component, this means that, unless explicitly stated otherwise, it does not exclude other components but may further include additional components.
[0070] In the present disclosure, unless otherwise explicitly stated, the terms first, second, etc. are only used to distinguish one component from another component, and do not limit the order or importance of the components. Therefore, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0071] In the present disclosure, distinguishable components are described to clearly explain their corresponding characteristics, and do not necessarily mean that the components are separate. In other words, multiple components can be integrated into a single hardware or software unit, or a single component can be distributed across multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of the present disclosure without explicitly describing them.
[0072] In the present disclosure, the components described in the various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of the present disclosure. In addition, embodiments including additional components in addition to the components described in the various embodiments are also included in the scope of the present disclosure.
[0073] The present disclosure relates to encoding and decoding of images, and terms used herein may have ordinary meanings commonly used in the technical field to which the present disclosure belongs, unless the terms are newly defined in the present disclosure.
[0074] In this disclosure, "video" may refer to a collection of images in sequence over time.
[0075] In this disclosure, a "picture" generally refers to a unit representing a single image at a specific point in time. A slice / tile is a coding unit that constitutes a part of a picture, and a picture may consist of one or more slices / tiles. In addition, a slice / tile may include one or more coding tree units (CTUs).
[0076] In the present disclosure, "pixel" or "picture element" may refer to the smallest unit constituting a picture (or image). In addition, the term "sample" may be used as a corresponding term for a pixel. A sample may generally represent a pixel or a value of a pixel, and may indicate only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.
[0077] In the present disclosure, "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific area of a picture or information related to the area. Depending on the context, the term "unit" may be used interchangeably with "sample array", "block", "area", etc. In general, an M×N block may include a set (or array) of samples (or sample array) or a set (or array) of transform coefficients, which consists of M columns and N rows.
[0078] In the present disclosure, the term "current block" may refer to one of "current coding block", "current coding unit", "encoding target block", "decoding target block" or "processing target block". When prediction is performed, the "current block" may refer to "current prediction block" or "prediction target block". When transform (inverse transform) / quantization (dequantization) is performed, the "current block" may refer to "current transform block" or "transform target block". When filtering is performed, the "current block" may refer to "filtering target block".
[0079] In the present disclosure, unless explicitly stated as a chrominance block, the term "current block" may refer to a block including both a luma component block and a chrominance component block, or may refer to a "luma block of the current block". The luma component block of the current block may be explicitly expressed with terms such as "luma block" or "current luma block", clearly indicating it as a luma component block. In addition, the chrominance component block of the current block may be explicitly expressed with terms such as "chrominance block" or "current chrominance block", clearly indicating it as a chrominance component block.
[0080] In the present disclosure, " / " and "," may refer to "and / or". For example, "A / B" and "A, B" may refer to "A and / or B". In addition, "A / B / C" and "A, B, C" may refer to "at least one of A, B and / or C".
[0081] In the present disclosure, "or" may mean "and / or". For example, "A or B" may mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" may also mean "in addition or alternatively".
[0082] In the present disclosure, "at least one of A, B, and C" may refer to "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B and / or C" may refer to "at least one of A, B, and C".
[0083] The brackets used in the present disclosure may refer to "for example". For example, when describing "prediction (intra-frame prediction)", "intra-frame prediction" may be proposed as an example of "prediction". In other words, "prediction" in the present disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" may be proposed as an example of "prediction". In addition, when describing "prediction (ie, intra-frame prediction)", "intra-frame prediction" may also be proposed as an example of "prediction".
[0084] Overview of the video compilation system
[0085] Figure 1 A schematic diagram showing a video coding system to which an embodiment of the present disclosure can be applied is shown.
[0086] The video coding system according to an embodiment may include an encoder device 10 and a decoder device 20. The encoder device 10 may transmit encoded video and / or image information or data to the decoder device 20 in the form of a file or a stream through a digital storage medium or a network.
[0087] The encoder device 10 according to the embodiment may include a video source generator 11, an encoder 12, and a transmitter 13. The decoder device 20 according to the embodiment may include a receiver 21, a decoder 22, and a renderer 23. The encoder 12 may be referred to as a video / image encoder, and the decoder 22 may be referred to as a video / image decoder. The transmitter 13 may be included in the encoder 12. The receiver 21 may be included in the decoder 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0088] The video source generator 11 may obtain the video / image by a process of capturing, synthesizing or generating the video / image. The video source generator 11 may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive containing previously captured videos / images, etc. The video / image generating device includes, for example, a computer, a tablet computer or a smart phone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer or the like, and in this case, the video / image capturing process may be replaced by a process of generating relevant data.
[0089] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of processes such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in the form of a bit stream.
[0090] The transmitter 13 can obtain the encoded video / image information or data output in the form of a bit stream and send it to the receiver 21 of the decoder device 20 or another external object in the form of a file or a stream through a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 may include an element for generating a media file in a predetermined file format and an element for transmitting through a broadcast / communication network. The transmitter 13 may be provided as a transmission device separate from the encoder 120, in which case the transmission device may include at least one processor for obtaining the encoded video / image information or data in the form of a bit stream and a transmitter for delivering it in the form of a file or a stream. The receiver 21 may extract / receive the bit stream from the storage medium or the network and send it to the decoder 22.
[0091] The decoder 22 may decode a video / image by performing a series of processes corresponding to the operations of the encoder 12 , such as dequantization, inverse transformation, prediction, and the like.
[0092] The renderer 23 may render the decoded video / image. The rendered video / image may be displayed through a display unit.
[0093] Overview of image encoding apparatus
[0094] Figure 2 A schematic diagram showing an image encoding device to which an embodiment of the present disclosure can be applied.
[0095] like Figure 2 As described in the above, the image encoding device 100 may include an image partitioner 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as a "predictor". The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may further include a subtractor 115.
[0096] Depending on the embodiment, all or at least some of the components constituting the image encoding device 100 may be implemented as a single hardware component (ie, the image encoding device 100 or a processor). In addition, the memory 170 may include a decoded picture buffer (DPB) and may be implemented by a digital storage medium.
[0097] The image partitioner 110 may partition an input image (or picture, frame) input to the image encoding device 100 into at least one processing unit. As an example, a processing unit may be referred to as a coding unit (CU). A coding unit may be obtained by recursively partitioning a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree, binary tree, or ternary tree (QT / BT / TT) structure. For example, a coding unit may be divided into coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the partitioning of a coding unit, a quadtree structure may be applied first, and then a binary tree structure and / or a ternary tree structure may be applied. The coding process according to the present disclosure may be performed based on a final coding unit without further partitioning the final coding unit. The maximum coding unit may be directly used as the final coding unit, or a coding unit of a deeper depth obtained by partitioning the maximum coding unit may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit for a coding process may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may each be divided or partitioned from the final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or deriving a residual signal from the transform coefficient.
[0098] The predictor (inter predictor 180 or intra predictor 185) may perform prediction on the target block (current block) and generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block or coding unit (CU). The predictor may generate various information related to the prediction of the current block and send it to the entropy encoder 190. The prediction related information may be encoded by the entropy encoder 190 and may be output in the form of a bitstream.
[0099] The intra-frame predictor 185 can predict the current block by referring to the samples in the current picture. The referenced samples may be located in the neighboring area of the current block, or may be located at a farther position, depending on the intra-frame prediction mode and / or the intra-frame prediction method. The intra-frame prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional mode may include, for example, a DC mode and a plane mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the granularity of the prediction direction. However, this is only an example, and a greater or lesser number of directional prediction modes may be used depending on the configuration. The intra-frame predictor 185 may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0100] The inter-frame predictor 180 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted at a block, sub-block or sample level based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information about the inter-frame prediction direction (ie, L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block or collocated coding unit (colCU). The reference picture containing the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 180 may construct a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and, for example, in skip mode and merge mode, the inter predictor 180 may use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may refer to the difference between the motion vector of the current block and the motion vector predictor.
[0101] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques described later. For example, the predictor may apply intra prediction or inter prediction to the prediction of the current block, and may also apply both intra prediction and inter prediction at the same time. The prediction method that simultaneously applies intra prediction and inter prediction to the prediction of the current block may be referred to as combined inter and intra prediction (CIIP). In addition, the predictor may perform intra block copy (IBC) on the prediction of the current block. Intra block copy may be used, for example, for screen content coding (SCC) in applications such as game content image / video coding. IBC is a method for predicting the current block by using a pre-reconstructed reference block within the current picture, which is located at a predetermined distance from the current block. When IBC is applied, the position of the reference block within the current picture may be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but because it derives the reference block within the current picture, it may operate similarly to inter prediction. In other words, IBC may use at least one of the inter prediction methods described in the present disclosure.
[0102] The prediction signal generated by the predictor can be used to generate a reconstruction signal or to generate a residual signal. The subtractor 115 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the predictor from the input image signal (original block, original sample array). The generated residual signal can be sent to the transformer 120.
[0103] The transformer 120 may generate a transform coefficient by applying a transform method to the residual signal. For example, the transform method may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. The transform process may be applied to pixel blocks of the same square size or non-square blocks of variable size.
[0104] The quantizer 130 may quantize the transform coefficients and send them to the entropy encoder 190. The entropy encoder 190 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 130 may rearrange the block-shaped quantized transform coefficients into a one-dimensional vector based on a coefficient scanning order, and may generate information about the quantized transform coefficients based on the one-dimensional vector of the quantized transform coefficients.
[0105] The entropy encoder 190 may perform various encoding methods, such as exponential Golomb, context adaptive variable length coding (CAVLC), or context adaptive binary arithmetic coding (CABAC). The entropy encoder 190 may encode not only the quantized transform coefficients, but also the information necessary for video / image reconstruction (i.e., the value of the syntax element) together with or separately from the quantized transform coefficients. The encoded information (i.e., the encoded video / image information) may be transmitted or stored in a network abstraction layer (NAL) unit in the form of a bitstream. The video / image information may further include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The signaling information, the transmitted information, and / or the syntax elements described in the present disclosure may be encoded and included in the bitstream through the above-mentioned encoding process.
[0106] The bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting a signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be provided as an internal / external element of the image encoding device 100, or the transmitter may be configured as a component of the entropy encoder 190.
[0107] The quantized transform coefficients output from the quantizer 130 may be used to generate a residual signal. For example, the residual signal (residual block or residual sample) may be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.
[0108] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-frame predictor 180 or the intra-frame predictor 185. When there is no residual for the target block, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current picture, and as described later, after filtering, it can also be used for inter-frame prediction of the next picture.
[0109] Meanwhile, luminance mapping and chrominance scaling (LMCS) may be applied during image encoding and / or reconstruction.
[0110] The filter 160 may apply filtering to the reconstructed signal to enhance the subjective / objective quality. For example, the filter 160 may apply various filtering methods to the reconstructed image to generate a modified reconstructed picture, and the modified reconstructed picture may be stored in the memory 170, specifically in the DPB of the memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 160 may generate various filtering related information, as described in the explanation of each filtering method later, and may send it to the entropy encoder 190. The filtering related information may be encoded by the entropy encoder 190 and output in the form of a bit stream.
[0111] The modified reconstructed picture transmitted to the memory 170 may be used as a reference picture in the inter predictor 180. When inter prediction is applied in this case, the image encoding device 100 may avoid prediction mismatch between the image encoding device 100 and the image decoding apparatus, and may improve encoding efficiency.
[0112] The DPB in the memory 170 may store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 180. The memory 170 may store the motion information of the block for which the motion information has been derived (or encoded) in the current picture and / or the motion information of the block in the reconstructed picture. The stored motion information may be sent to the inter-frame predictor 180 for use as the motion information of the spatial neighboring block or the temporal neighboring block. The memory 170 may store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 185.
[0113] Overview of image decoding device
[0114] Figure 3 A schematic diagram showing an image decoding device to which an embodiment of the present disclosure can be applied is shown.
[0115] like Figure 3 As shown in , the image decoding device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame predictor 265. The inter-frame predictor 260 and the intra-frame predictor 265 may be collectively referred to as a "predictor". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0116] All or at least some of the plurality of components constituting the image decoding device 200 may be implemented as a single hardware component (ie, the image decoding device 200 or a processor), depending on the embodiment. In addition, the memory 170 may include a DPB and may be implemented by a digital storage medium.
[0117] The image decoding apparatus 200 receiving a bit stream containing video / image information may perform the same Figure 2 The image may be reconstructed by a process corresponding to the process performed by the image encoding device 100 in the image decoding device 200. For example, the image decoding device 200 may perform decoding using a processing unit applied in the image encoding device 100. Therefore, the processing unit for decoding may be, for example, a coding unit. The coding unit may be a coding tree unit, or may be obtained by splitting a maximum coding unit. In addition, the reconstructed image signal decoded and output by the image decoding device 200 may be played by a playback device (not shown).
[0118] The image decoding apparatus 200 may receive the image in the form of a bit stream from Figure 2The received signal may be decoded by the entropy decoder 210. For example, the entropy decoder 210 may parse the bitstream to extract information (i.e., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The image decoding device 200 may additionally use the information about the parameter set and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements described in the present disclosure may be obtained from the bitstream by decoding through a decoding process. For example, the entropy decoder 210 may decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output syntax element values necessary for image reconstruction and quantized values of transform coefficients associated with the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to a syntax element in a bitstream, can use information of a decoding target syntax element, decoding information of a neighboring block and a decoding target block, or information of a previously decoded symbol / bin to determine a context model, can predict the probability of bin occurrence according to the determined context model, and can perform arithmetic decoding on the bin to generate a symbol corresponding to each syntax element. In this case, the CABAC entropy decoding method can update the context model for the next symbol / bin context model using the decoded symbol / bin information after determining the context model. Among the decoded information from the entropy decoder 210, prediction related information can be provided to the predictor (inter-frame predictor 260 and intra-frame predictor 265), and the residual value entropy decoded by the entropy decoder 210, in other words, the quantized transform coefficient and related parameter information, can be input to the dequantizer 220. In addition, among the decoded information from the entropy decoder 210, filtering related information can be provided to the filter 240. Meanwhile, a receiver (not shown) receiving a signal output from the image encoding device 100 may be additionally configured as an internal / external element of the image decoding device 200 , or the receiver may be configured as a component of the entropy decoder 210 .
[0119] Meanwhile, the image decoding device 200 according to the present disclosure may also be referred to as a video / image / picture decoding device. The image decoding device 200 may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210, and the sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or an intra-frame predictor 265.
[0120] The dequantizer 220 may dequantize the quantized transform coefficient and output the transform coefficient. The dequantizer 220 may rearrange the quantized transform coefficient into a two-dimensional block. In this case, the rearrangement may be performed based on the coefficient scanning order applied in the image encoding device 100. The dequantizer 220 may perform dequantization on the quantized transform coefficient using a quantization parameter (ie, quantization step size information), and may obtain the transform coefficient.
[0121] The inverse transformer 230 may perform inverse transform on the transform coefficients to obtain a residual signal (a residual block or a residual sample array).
[0122] The predictor may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the prediction related information output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction method).
[0123] The predictor can generate a prediction signal based on various prediction methods (techniques) to be described later, which is the same as described in the explanation of the predictor in the image encoding device 100 .
[0124] The intra predictor 265 may predict the current block by referring to samples within the current picture. The explanation of the intra predictor 185 may also be applied to the intra predictor 265 in the same manner.
[0125] The inter-frame predictor 260 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted at a block, sub-block or sample level based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information about the inter-frame prediction direction (ie, L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, the neighboring blocks may include spatial neighboring blocks within the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 260 may construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction may be performed based on various prediction modes (methods), and the prediction-related information may include information indicating the inter-frame prediction mode (method) applied to the current block.
[0126] The adder 235 can generate a reconstruction signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 260 and / or the intra-frame predictor 265). When there is no residual for the target block, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The explanation of the adder 155 can also be applied to the adder 235 in the same manner. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstruction signal can be used for intra-frame prediction of the next target block in the current picture, and as described later, it can also be used for inter-frame prediction of the next picture after filtering.
[0127] The filter 240 may apply filtering to the reconstructed signal to enhance the subjective / objective quality. For example, the filter 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and the modified reconstructed picture may be stored in the memory 250, specifically in the DPB of the memory 250. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0128] The (modified) reconstructed picture stored in the DPB of the memory 250 may be used as a reference picture in the inter-frame predictor 260. The memory 250 may store the motion information of the block for which the motion information has been derived (or decoded) in the current picture and / or the motion information of the block in the reconstructed picture. The stored motion information may be sent to the inter-frame predictor 260 to be used as the motion information of the spatial neighboring block or the temporal neighboring block. The memory 250 may store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 265.
[0129] In the present specification, the embodiments described for the filter 160, the inter-frame predictor 180, and the intra-frame predictor 185 of the image encoding device 100 may be applied to the filter 240, the inter-frame predictor 260, and the intra-frame predictor 265 of the image decoding device 200 in the same or corresponding manner.
[0130] Overview of Intra Prediction
[0131] Hereinafter, intra prediction according to the present disclosure will be described.
[0132] Intra-frame prediction may refer to a prediction method for generating prediction samples for the current block based on reference samples within a picture to which the current block belongs (hereinafter referred to as the current picture). When intra-frame prediction is applied to the current block, neighboring reference samples to be used for intra-frame prediction of the current block may be derived. The neighboring reference samples of the current block may include a total of 2×nH samples adjacent to the left boundary of the current block of nW×nH size / adjacent to the left boundary of the current block of nW×nH size and adjacent to the lower left corner of the current block of nW×nH size, a total of 2×nW samples adjacent to the upper boundary of the current block and adjacent to the upper right corner of the current block, and one sample adjacent to the upper left corner of the current block. Alternatively, the neighboring reference samples of the current block may include multiple columns of top neighboring samples and multiple rows of left neighboring samples. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of nWxnH size, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the lower right corner of the current block.
[0133] However, some neighboring reference samples of the current block may not have been decoded or may not be available. In this case, the image decoding device 200 can construct neighboring reference samples for prediction by replacing unavailable samples with available samples. Alternatively, the neighboring reference samples for prediction can be constructed by interpolating available samples.
[0134] When deriving neighboring reference samples, (i) the prediction sample may be derived based on an average or interpolation of neighboring reference samples of the current block, and (ii) the prediction sample may be derived based on a reference sample located in a specific (prediction) direction for the prediction sample among neighboring reference samples of the current block. Case (i) may be referred to as a non-directional mode or a non-angular mode, and case (ii) may be referred to as a directional mode or an angular mode.
[0135] In addition, the prediction sample can be generated by interpolating between the first neighboring sample located in the prediction direction of the intra prediction mode of the current block based on the prediction target sample of the current block and the second neighboring sample located in the opposite direction among the neighboring reference samples. The above situation can be called linear interpolation intra prediction (LIP).
[0136] In addition, a linear model can be used to generate chrominance prediction samples based on luma samples. This case can be called linear model (LM) mode.
[0137] In addition, the temporary prediction sample of the current block can be derived based on the filtered neighboring reference sample, and the prediction sample of the current block can be derived by calculating the weighted sum of at least one of the reference sample derived according to the intra prediction mode and the temporary prediction sample among the conventional neighboring reference samples (i.e., the unfiltered neighboring reference samples). This case is called position-dependent intra prediction (PDPC).
[0138] In addition, a reference sample line with the highest prediction accuracy among multiple adjacent reference sample lines of the current block can be selected, and the reference sample located in the prediction direction in the corresponding line can be used to derive the prediction sample. In this case, information about the reference sample line used (ie, intra_luma_ref_idx) can be encoded and signaled in the bitstream. In this case, it is called multi-reference line intra prediction (MRL) or MRL-based intra prediction. When MRL is not applied, reference samples can be derived from reference sample lines directly adjacent to the current block, and in this case, information about the reference sample line may not be signaled.
[0139] In addition, the current block can be divided into vertical or horizontal sub-partitions, and intra-frame prediction can be performed based on the same intra-frame prediction mode for each sub-partition. In this case, neighboring reference samples for intra-frame prediction can be derived for each sub-partition unit. In other words, the reconstructed samples of the previous sub-partition in the encoding / decoding order can be used as neighboring reference samples for the current sub-partition. In this case, the intra-frame prediction mode for the current block is applied to the sub-partition in the same way, and neighboring reference samples are derived and used for each sub-partition unit, so that the intra-frame prediction performance can be improved in some cases. This prediction method is called intra sub-partitioning (ISP) or intra-frame prediction based on ISP.
[0140] The above-mentioned intra-frame prediction method can be referred to by various terms, such as intra-frame prediction type or additional intra-frame prediction mode, to distinguish it from a directional or non-directional intra-frame prediction mode. For example, the intra-frame prediction method (i.e., intra-frame prediction type or additional intra-frame prediction mode, etc.) may include at least one of the above-mentioned LIP, LM, PDPC, MRL or ISP. A general intra-frame prediction method excluding a specific intra-frame prediction type such as LIP, LM, PDPC, MRL, ISP, etc. may be referred to as a normal intra-frame prediction type. When the above-mentioned specific intra-frame prediction type is not used, a normal intra-frame prediction type may generally be applied, and prediction may be performed based on the above-mentioned intra-frame prediction mode. At the same time, post-processing filtering may be performed on the derived prediction samples when necessary.
[0141] Specifically, the intra prediction process may include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra prediction mode / type. In addition, a post-filtering step may be performed on the derived prediction samples if necessary.
[0142] Meanwhile, in addition to the above-mentioned intra prediction types, affine linear weighted intra prediction (ALWIP) can also be used. ALWIP can also be referred to as linear weighted intra prediction (LWIP) or matrix weighted intra prediction or matrix-based intra prediction (MIP). When MIP is applied to the current block, the prediction sample for the current block can be derived by i) using neighboring reference samples to which an averaging process has been performed, ii) performing a matrix vector multiplication process, and iii) further performing a horizontal / vertical interpolation process when necessary. The intra prediction mode used for MIP can be configured differently from those used in the above-mentioned LIP, PDPC, MRL, ISP intra prediction or normal intra prediction. The intra prediction mode for MIP can be referred to as MIP intra prediction mode, MIP prediction mode or MIP mode. For example, the matrix and offset used in the matrix vector multiplication can be set differently depending on the intra prediction mode for MIP. Here, the matrix can be referred to as a (MIP) weight matrix, and the offset can be referred to as a (MIP) offset vector or a (MIP) offset vector. A specific MIP method will be described later.
[0143] Will refer to it later Figure 4 and Figure 5 A block reconstruction process based on intra prediction and an intra predictor in an encoding device is described.
[0144] Figure 4 The present invention is a flowchart of a video / image encoding method based on intra-frame prediction.
[0145] Figure 4 The encoding method can be Figure 2 The image encoding device 100 of the present invention is performed. Specifically, step S410 can be performed by the intra predictor 185, and step S420 can be performed by the residual processor. Specifically, step S420 can be performed by the subtractor 115. Step S430 can be performed by the entropy encoder 190. The prediction information in step S430 can be derived by the intra predictor 185, and the residual information in step S430 can be derived by the residual processor. The residual information refers to information about the residual sample. The residual information may include information about the quantized transform coefficient of the residual sample. As described above, the residual sample is derived as a transform coefficient by the transformer 120 of the image encoding device 100, and the transform coefficient can be derived as a quantized transform coefficient by the quantizer 130. The information about the quantized transform coefficient can be encoded in the entropy encoder 190 through the residual encoding process.
[0146] The image encoding device 100 may perform intra prediction on the current block S410. The image encoding device 100 may determine an intra prediction mode / type for the current block, may derive neighboring reference samples of the current block, and may generate prediction samples within the current block based on the intra prediction mode / type and the neighboring reference samples. Here, the processes of determining the intra prediction mode / type, deriving neighboring reference samples, and generating prediction samples may be performed simultaneously, or one process may be performed before the other process.
[0147] Figure 5 An exemplary diagram showing a configuration of the intra predictor 185 according to the present disclosure.
[0148] like Figure 5 As shown in , the intra-frame predictor 185 of the image encoding device 100 may include an intra-frame prediction mode / type determiner 186, a reference sample deriver 187 and / or a prediction sample deriver 188. The intra-frame prediction mode / type determiner 186 may determine the intra-frame prediction mode / type for the current block. The reference sample deriver 187 may derive the neighboring reference samples of the current block. The prediction sample deriver 188 may derive the prediction sample of the current block. Meanwhile, although not shown, when performing the prediction sample filtering process to be described later, the intra-frame predictor 185 may further include a prediction sample filter (not shown).
[0149] The image encoding device 100 may determine an intra prediction mode / type applied to a current block among a plurality of intra prediction modes / types. The image encoding device 100 may compare rate-distortion costs (RD costs) of intra prediction modes / types and may determine an optimal intra prediction mode / type for the current block.
[0150] At the same time, the image encoding device 100 may perform a prediction sample filtering process. The prediction sample filtering may be referred to as post-filtering. Through the prediction sample filtering process, part or all of the prediction samples may be filtered. In some cases, the prediction sample filtering process may be omitted.
[0151] Reference again Figure 4 , the image encoding device 100 may generate residual samples for the current block based on the prediction samples or the filtered prediction samples S420. The image encoding device 100 may derive the residual samples by subtracting the prediction samples from the original samples of the current block. In other words, the image encoding device 100 may derive the residual sample values by subtracting the corresponding prediction sample values from the original sample values.
[0152] The image encoding device 100 may encode image information including information about intra prediction (prediction information) and residual information about residual samples S430. The prediction information may include intra prediction mode information and / or intra prediction method information. The image encoding device 100 may output the encoded image information in the form of a bit stream. The output bit stream may be sent to the image decoding device 200 via a storage medium or a network.
[0153] The residual information may include a residual coding syntax, which will be described later. The image encoding apparatus 100 may derive a quantized transform coefficient by transforming / quantizing the residual samples. The residual information may include information on the quantized transform coefficient.
[0154] At the same time, as described above, the image encoding device 100 can generate a reconstructed picture (including reconstructed samples and reconstructed blocks). The image encoding device 100 can perform dequantization / inverse transformation on the quantized transform coefficients to derive (modified) residual samples. The reason for applying dequantization / inverse transformation after transforming / quantizing the residual samples is to derive residual samples that are the same as the residual samples derived in the image decoding device 200. The image encoding device 100 can generate a reconstructed block including reconstructed samples for the current block based on the predicted samples and the (modified) residual samples. A reconstructed picture for the current picture can be generated based on the reconstructed block. As described above, the loop filtering process can be further applied to the reconstructed picture.
[0155] Figure 6 A flow chart showing a video / image decoding method based on intra-frame prediction.
[0156] The image decoding device 200 may perform operations corresponding to those performed in the image encoding device 100 .
[0157] Figure 6 The decoding method in can be Figure 3 The image decoding device 200 in the embodiment of the present invention is performed. Steps S610 to S630 may be performed by the intra predictor 265, and the prediction information in step S610 and the residual information in step S640 may be obtained from the bitstream by the entropy decoder 210. The residual processor of the image decoding device 200 may derive residual samples S640 for the current block based on the residual information. Specifically, the dequantizer 220 of the residual processor may perform dequantization on the quantized transform coefficients derived based on the residual information to obtain transform coefficients, and the inverse transformer 230 of the residual processor may perform inverse transform on the transform coefficients to derive residual samples for the current block. Step S650 may be performed by the adder 235 or the reconstructor.
[0158] Specifically, the image decoding device 200 may derive an intra prediction mode / type for the current block based on the received prediction information (i.e., intra prediction mode / type information) S610. In addition, the image decoding device 200 may derive neighboring reference samples of the current block S620. The image decoding device 200 may generate prediction samples within the current block based on the intra prediction mode / type and the neighboring reference samples S630. In this case, the image decoding device 200 may perform a prediction sample filtering process. Prediction sample filtering may be referred to as post filtering. Through this prediction sample filtering process, some or all of the prediction samples may be filtered. In some cases, the prediction sample filtering process may be omitted.
[0159] The image decoding device 200 may generate residual samples for the current block based on the received residual information S640. The image decoding device 200 may generate reconstructed samples for the current block based on the predicted samples and the residual samples, and derive a reconstructed block including the reconstructed samples S650. A reconstructed picture for the current picture may be generated based on the reconstructed block. As described above, the loop filtering process may be further applied to the reconstructed picture.
[0160] Figure 7 An exemplary diagram showing a configuration of the intra predictor 265 according to the present disclosure.
[0161] like Figure 7 As shown in , the intra-frame predictor 265 of the image decoding device 200 may include an intra-frame prediction mode / type determiner 266, a reference sample deriver 267, and a prediction sample deriver 268. The intra-frame prediction mode / type determiner 266 may determine the intra-frame prediction mode / type for the current block based on the intra-frame prediction mode / type information generated in the intra-frame prediction mode / type determiner 186 of the image encoding device 100 and transmitted by a signal, and the reference sample deriver 266 may derive the neighboring reference samples of the current block from the reconstructed reference area in the current picture. The prediction sample deriver 268 may derive the prediction sample of the current block. Meanwhile, although not shown, when performing the above-mentioned prediction sample filtering process, the intra-frame predictor 265 may further include a prediction sample filter (not shown).
[0162] The intra-frame prediction mode information may include, for example, flag information (i.e., intra_luma_mpm_flag) indicating whether the most probable mode (MPM) is applied to the current block or the residual mode is applied, and when the MPM is applied to the current block, the intra-frame prediction mode information may further include index information (i.e., intra_luma_mpm_idx) indicating one of the intra-frame prediction mode candidates (MPM candidates). The intra-frame prediction mode candidates (MPM candidates) may be configured as an MPM candidate list or an MPM list. In addition, when the MPM is not applied to the current block, the intra-frame prediction mode information may further include residual mode information (i.e., intra_luma_mpm_remainder) indicating one of the remaining intra-frame prediction modes other than the intra-frame prediction mode candidates (MPM candidates). The image decoding device 200 may determine the intra-frame prediction mode of the current block based on the intra-frame prediction mode information.
[0163] In addition, the intra prediction method information may be implemented in various forms. As an example, the intra prediction method information may include intra prediction method index information indicating one of the intra prediction methods. As another example, the intra prediction method information may include reference sample line information (i.e., intra_luma_ref_idx) indicating whether MRL is applied to the current block and which reference sample line is used when MRL is applied, ISP flag information (i.e., intra_subpartitions_mode_flag) indicating whether ISP is applied to the current block, ISP type information (i.e., intra_subpartitions_split_flag) indicating the split type of the sub-partition when ISP is applied, flag information indicating whether PDPC is applied, or at least one of flag information indicating whether LIP is applied. In addition, the intra prediction type information may include a MIP flag indicating whether MIP is applied to the current block. In the present disclosure, the ISP flag information may be referred to as an ISP application indicator.
[0164] The intra-frame prediction mode information and / or the intra-frame prediction method information may be encoded / decoded by the coding method described in the present disclosure. For example, the intra-frame prediction mode information and / or the intra-frame prediction method information may be encoded / decoded by entropy coding based on truncated (Rice) binary code (ie, CABAC, CAVLC).
[0165] Meanwhile, in addition to the planar mode, DC mode, and directional intra prediction mode, the intra prediction mode may further include a cross component linear model (CCLM) mode for chroma samples. Depending on whether the left sample, the top sample, or both are considered for deriving CCLM parameters, the CCLM mode may be classified into L_CCLM, T_CCLM, and LT_CCLM, and it is applied only to the chroma component.
[0166] For example, the intra prediction modes may be indexed as shown in Table 1 below.
[0167] [Table 1]
[0168]
[0169] Meanwhile, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the above-mentioned LIP, PDPC, MRL, ISP, or MIP. The intra prediction type may be indicated based on the intra prediction type information, and the intra prediction type information may be implemented in various forms. As an example, the intra prediction type information may include an intra prediction type index indicating one of the intra prediction types. In another example, the intra prediction type information may include reference sample line information (i.e., intra_luma_ref_idx) indicating whether MRL is applied to the current block and which reference sample line is used when MRL is applied, ISP flag information (i.e., intra_subpartitions_mode_flag) indicating whether ISP is applied to the current block, ISP type information (i.e., intra_subpartitions_split_flag) indicating the partition type of the sub-partition when ISP is applied, flag information indicating whether PDPC is applied, or at least one of flag information indicating whether LIP is applied. In addition, the intra prediction type information may include a MIP flag (which may be referred to as intra_mip_flag) indicating whether MIP is applied to the current block.
[0170] Overview of Multiple Reference Lines (MRL)
[0171] In conventional intra prediction, only the neighboring samples of the first line above the current block and the first line to the left of the current block are used as reference samples for intra prediction. However, in the MRL method, the image encoding device 100 and / or the image decoding device 200 can perform intra prediction by using neighboring samples located in a sample line located at a distance of one to three samples from the top and / or left of the current block as reference samples.
[0172] Figure 8 and Fig. 9 Illustrated is a reference sample line used for MRL-based intra prediction according to the present disclosure. Figure 8 Can be an example of multiple reference lines. Figure 8, at least one of reference sample line 0 (reference line 0), reference sample line 1 (reference line 1), reference sample line 2 (reference line 2), or reference sample line 3 (reference line 3) may be used for prediction of the current block. Here, the multiple reference line index (ie, mrl_idx) may be information indicating a reference sample line used for intra prediction. For example, the multiple reference line index may be signaled through a coding unit syntax as shown in Table 2 below. The multiple reference line index may be configured in the form of an intra_luma_ref_idx syntax element.
[0173] [Table 2]
[0174]
[0175] Here, intra_luma_ref_idx[x0][y0] may represent the intra reference line index IntraLumaRefLineIdx[x0][y0], as shown in Table 3 below. When intra_luma_ref_idx[x0][y0] is not present (ie, not signaled), it may be inferred to be 0. intra_luma_ref_idx may be referred to as a (intra) reference sample line index or mrl_idx. Additionally, intra_luma_ref_idx may be referred to as intra_luma_ref_line_idx.
[0176] [Table 3]
[0177]
[0178] MRL may not be used for blocks in the first line (row) within a coding tree unit (CTU). This may prevent the use of extended reference lines outside the current CTU line. In other words, this may prevent the use of external reference samples not included in the current CTU. In addition, when the above-mentioned additional reference lines are used, position-dependent intra prediction (PDPC) may not be used. In other words, when extended reference samples are used, PDPC may not be applied.
[0179] The MRL method can use neighboring samples in a sample line located at a distance of one to three samples from the top and / or left of the current block as reference samples to perform intra prediction. However, in extended MRL, neighboring samples in a sample line located at a distance of up to twelve samples from the top and / or left of the current block can be used as reference samples to perform intra prediction. Fig. 9 , multiple reference sample lines adjacent to the current block can be configured as an extended MRL candidate list. For example, the reference sample line index in the extended MRL list can be configured as {1, 3, 5, 7, 12}.
[0180] When extended MRL is used, intra_luma_ref_idx may be configured as shown in Table 4 below. intra_luma_ref_idx[x0][y0] may represent the intra reference line index IntraLumaRefLineIdx[x0][y0]. When intra_luma_ref_idx[x0][y0] does not exist (i.e., is not signaled), intra_luma_ref_idx[x0][y0] may be inferred to be 0. intra_luma_ref_idx may be referred to as a (intra) reference sample line index or mrl_idx. Additionally, intra_luma_ref_idx may be referred to as intra_luma_ref_line_idx.
[0181] [Table 4]
[0182]
[0183] Overview of Template-based Intra Mode Derivation (TIMD)
[0184] Fig.10 A diagram showing a template region and reference samples used for TIMD according to the present disclosure. For intra prediction modes (IPM) of adjacent intra blocks and inter blocks, TIMD can select a mode with a minimum SATD as the intra mode of the current block by calculating the sum of absolute transform differences (SATD) between the predicted block predicted from the template region 1010 and the actual reconstructed sample.
[0185] According to another embodiment of the present disclosure, the present disclosure may select two modes with minimum SATD, and then generate a prediction block for each of the two selected prediction modes. The prediction block of the current block may be generated by weighting and mixing the two generated prediction blocks. Here, the mixing of the two modes may be performed when the conditions in the following formula 1 are met.
[0186] [Formula 1]
[0187]
[0188] Here, costMode1 may refer to a mode having the smallest SATD. In addition, costMode2 may refer to a mode having the second smallest SATD.
[0189] When the condition in Formula 1 is met, the final prediction block can be generated by mixing the prediction blocks generated by the two modes. Otherwise, the final prediction block can be generated using only the mode with the minimum SATD value.
[0190] The weighting ratio applied when mixing the prediction blocks generated using the two prediction modes with the minimum SATD may be as shown in Formula 2 below.
[0191] [Formula 2]
[0192]
[0193] Here, weight1 may refer to a weight applied to a prediction block generated based on a mode having the smallest SATD. In addition, weight2 may refer to a weight applied to a prediction block generated based on a mode having the second smallest SATD.
[0194] Overview of decoder-side intra mode derivation (DIMD)
[0195] Fig.11 The method for constructing a gradient histogram (HoG) in the DIMD mode according to the present disclosure is illustrated. The DIMD mode according to the present disclosure can derive intra prediction mode information from the image encoding device 100 and the image decoding device 200 and use the information without directly transmitting the information. The DIMD mode can be performed by obtaining horizontal gradients and vertical gradients from second neighboring reference columns and rows adjacent to the current block and constructing the HoG based on them.
[0196] refer to Fig.11 , HoG can be obtained by applying a Sobel filter using L-shaped columns and rows of three neighboring pixels 1110 around the current block. In this case, when the boundary of the block exists in different CTUs, the neighboring pixels of the current block may not be used for texture analysis.
[0197] Meanwhile, the Sobel filter may be referred to as a Sobel operator and may be an effective filter for detecting edges. When the Sobel filter is used, two types of Sobel filters may be used, a Sobel filter for a vertical direction and a Sobel filter for a horizontal direction.
[0198] Fig.12 FIG. 1 illustrates a method for constructing a prediction block when a DIMD mode is applied according to the present disclosure. Fig.12 , the DIMD mode may be performed by selecting two intra modes with the highest histogram magnitude 1210 and generating a final prediction block 1250 by mixing the prediction blocks predicted by the two selected intra modes (1220, 1230) and the prediction block predicted by the planar mode (1240). In this case, the weight applied when mixing the prediction blocks may be derived from the histogram magnitude 1260. In addition, a DIMD flag may be sent on a block unit to determine whether DIMD is used.
[0199] Hereinafter, an image encoding / decoding method according to various embodiments of the present disclosure will be described in detail.
[0200] Example 1
[0201] The present disclosure may relate to a method for performing a planar mode using an MRL method. In other words, the present disclosure may describe a method for performing a planar mode by using a pre-reconstructed sample located on a sample line at a distance of k samples from the top and / or left side of a current block as a reference sample. Here, k may be a natural number. In this case, a valid planar prediction block may be generated by considering the distance between reference samples that are not adjacent to sample positions within the current block when performing the planar mode.
[0202] FIG13 illustrates an intra-frame prediction method based on MRL according to the present disclosure. Fig.13 , a prediction sample p(x, y) 1320 at a position (x, y) within the current block 1310 may be generated using the top reference sample T, the upper right reference sample TR, the left reference sample L, and / or the lower left reference sample BL.
[0203] According to an embodiment of the present disclosure, the encoder device 100 and / or the decoder device 200 may use a horizontal plane mode and / or a vertical plane mode to generate a prediction sample 1320. When the position of the upper left sample of the current block is defined as (0, 0), the horizontal plane mode may be to use a left reference sample L located at (-k-1, y) and an upper right reference sample TR located at (W, -k-1) to generate a prediction sample at a position (x, y). Here, k may represent the distance between the current block and the reference sample line. In addition, W may represent the width of the current block. The prediction sample may be calculated using the following formula 3.
[0204] [Formula 3]
[0205]
[0206] In formula 3, and may be the weights applied to L and TR respectively. and The value of can be calculated using the following formula 4. In other words, it can be determined based on the distance between the current block and the reference sample line, the size of the current block and / or the position of the prediction sample. and .
[0207] [Formula 4]
[0208]
[0209] When the position of the upper left pixel of the current block is set to (0, 0), the vertical planar mode may be to generate a prediction sample at position (x, y) using the top reference sample T located at (x, -k-1) and the bottom left reference sample BL located at (-k-1, H) Here, k can represent the distance between the current block and the reference sample line. In addition, H can represent the height of the current block. The prediction sample can be calculated using the following formula 5 .
[0210] [Formula 5]
[0211]
[0212] In formula 5, and can be the weights applied to T and BL respectively. and The value of can be calculated using the following formula 6. In other words, and It may be determined based on the distance between the current block and the reference sample line, the size of the current block and / or the position of the prediction sample.
[0213] [Formula 6]
[0214]
[0215] The final prediction sample according to the present disclosure can be calculated using Formula 7. For example, the final prediction sample can be a prediction sample generated using a horizontal plane mode Alternatively, the final prediction samples may be prediction samples generated using the vertical plane mode Alternatively, the final prediction sample may be a prediction sample generated using a horizontal plane pattern. and prediction samples generated using vertical planar patterns The average value of .
[0216] [Formula 7]
[0217]
[0218] According to another embodiment of the present disclosure, by using the prediction samples generated by the horizontal plane mode and the prediction samples generated by using the vertical plane pattern Both can be calculated using Formula 8.
[0219] [Formula 8]
[0220]
[0221] In formula 8, This can be done by The values are obtained by approximating integers and can have a range from 0 to In addition, It can be The value obtained is approximately an integer and can have a range from 0 to Here, i can be a natural number.
[0222] Fig.14 1 shows a reference sample used to generate a prediction sample according to an embodiment of the present disclosure. The present disclosure may use a reference sample located at ( The horizontal plane mode is performed by using TR samples at -k, -k-1) on the left side, which is horizontally shifted by ,like Fig.14 In this case, the image encoding device 100 and / or the image decoding device 200 may generate a prediction sample 1420 within the current block 1410 using Formula 3. Here, can be the weight applied to L. In addition, Can be a weight to apply to TR. and It can be calculated using the following formula 9.
[0223] [Formula 9]
[0224]
[0225] refer to Fig.14 , the present disclosure may use the position (-k-1, -k) to perform the vertical planar pattern, which is vertically displaced from the top reference sample line by In this case, the image encoding device 100 and / or the image decoding device 200 may generate a prediction sample 1420 within the current block 1410 using Formula 5. Here, can be the weight applied to T, and It can be a weight applied to BL. and This can be calculated using the following formula 10.
[0226] [Formula 10]
[0227]
[0228] Fig.15 The figure in FIG. 1 illustrates reference samples used to generate prediction samples according to another embodiment of the present disclosure. Fig.15, the present disclosure may perform planar mode prediction by using a pre-reconstructed sample on a reference sample line located k samples away from the left side of the current block 1510 and a reference sample line located 1 sample away from the top of the current block 1510 as reference samples. Here, k may be a natural number. In order to perform the horizontal planar mode, the present disclosure may generate the prediction sample 1520 using Formula 3. In other words, in order to perform the horizontal planar mode, the image encoding device 100 and / or the image decoding device 200 may use a left reference sample L located k samples away from the current block 1510 and an upper right reference sample TR located 1 sample away from the current block 1510. In this case, can be the weight applied to L. In addition, Can be a weight to apply to TR. and It can be calculated using the following formula 11.
[0229] [Formula 11]
[0230]
[0231] The present disclosure may generate the prediction sample 1520 using Formula 5 to perform the vertical plane mode. In other words, in order to perform the vertical plane mode, the image encoding device 100 and / or the image decoding device 200 may use the top reference sample T located 1 sample away from the current block 1510 and the bottom left reference sample BL located k samples away from the current block 1510. In this case, can be the weight applied to T. In addition, It can be a weight applied to BL. and This can be calculated using the following formula 12.
[0232] [Formula 12]
[0233]
[0234] Fig.16 FIG. 2 illustrates a reference sample used to generate a prediction sample according to another embodiment of the present disclosure. Fig.16, the present disclosure may perform planar mode prediction by using a reference sample line located 1 sample to the left of the current block 1610 and a pre-reconstructed sample on a reference sample line located k samples above the current block 1610 as reference samples. Here, k may be a natural number. In order to perform a horizontal planar mode, the present disclosure may generate a prediction sample 1620 using Formula 3. In other words, in order to perform a horizontal planar mode, the image encoding device 100 and / or the image decoding device 200 may use a left reference sample L located 1 sample to the left of the current block 1610 and an upper right reference sample TR located k samples above the current block 1610. In this case, can be the weight applied to L. In addition, Can be a weight to apply to TR. and It can be calculated using the following formula 13.
[0235] [Formula 13]
[0236]
[0237] The present disclosure may generate the prediction sample 1620 using Formula 5 to perform the vertical plane mode. In other words, in order to perform the vertical plane mode, the image encoding device 100 and / or the image decoding device 200 may use the top reference sample T located at k samples away from the current block 1610 and the bottom left reference sample BL located at 1 sample away from the current block 1610. In this case, can be the weight applied to T. In addition, It can be a weight applied to BL. and It can be calculated using the following formula 14.
[0238] [Formula 14]
[0239]
[0240] Fig.17 Illustrated is a modified reference sample for generating a prediction sample according to an embodiment of the present disclosure. Fig.17 The upper left position of the current block 1710 is defined as (0, 0), and the left and top areas of the current block 1710 may be pre-reconstructed areas. Fig.17, when using the pre-reconstructed samples located at a reference sample line k+1 samples away from the left and a reference sample line 1 sample away from the top of the current block 1710 as reference samples, the present disclosure may perform a more accurate planar prediction mode by using a left reference sample L' adjacent to the current block 1710 and a lower left reference sample BL' adjacent to the current block 1710. Here, L' and BL' may be reference samples included in the pre-reconstructed region and may be defined using various methods.
[0241] For example, L' may be a pre-reconstructed sample located at (-1, y), and BL' may be a pre-reconstructed sample located at (-1, H). Here, y may be the y coordinate of the prediction sample 1720 to be predicted. In addition, H may be the height of the current block 1710. In another example, L' may be a pre-reconstructed sample L located at (-1-k, y). In other words, L and L' may be the same. In addition, BL' may be a pre-reconstructed sample BL located at (-1-k, H). In other words, BL and BL' may be the same. In another example, L' and BL' may be defined based on a difference d (=qp) between a reference sample p located at (-k-1, -1) and a reference sample q located at (-1, -1). In other words, L' may be L+d, and BL' may be BL+d. Here, d may represent the distance between the reference sample p and the reference sample q. Alternatively, d may be the difference between the sample values of the reference sample p and the reference sample q.
[0242] As another example, L' and BL' may be generated by using a left reference sample L and an upper right reference sample TR using a horizontal plane mode. In another example, L' and BL' may be generated by using a vertical plane mode using a reference sample q and a lower left reference sample BL. In another example, BL' may be generated by using a vertical plane mode using a reference sample q and a reference sample located at (-k-1, H+1).
[0243] As another example, L' and BL' may be generated as an average of a horizontal plane pattern using the left reference sample L and the upper right reference sample TR and a vertical plane pattern using the reference sample q and the lower left reference sample BL. In other words, L' and BL' may be generated as an average of sample values of a sample generated by a horizontal plane pattern using the left reference sample L and the upper right reference sample TR and a sample value of a sample generated by a vertical plane pattern using the reference sample q and the lower left reference sample BL.
[0244] As another example, BL' may be generated as an average of a horizontal plane pattern using the left reference sample L and the upper right reference sample TR and a vertical plane pattern using the reference sample q and the reference sample at (-k-1, H+1). In other words, BL' may be generated as an average of sample values of a sample generated by a horizontal plane pattern using the left reference sample L and the upper right reference sample TR and a sample value of a sample generated by a vertical plane pattern using the reference sample q and the reference sample at (-k-1, H+1).
[0245] To generate prediction samples 1720 using L' and BL' in the horizontal planar mode, the following formula 15 may be used. In formula 15, can be the weight applied to L'. In addition, Can be a weight to apply to TR. and It can be calculated using the following formula 16.
[0246] [Formula 15]
[0247]
[0248] [Formula 16]
[0249]
[0250] To generate prediction samples 1720 using L' and BL' using the vertical planar mode, the following formula 17 may be used. In formula 17, can be the weight applied to T. In addition, M may be a weight applied to BL'. and It can be calculated using the following formula 18.
[0251] [Formula 17]
[0252]
[0253] [Formula 18]
[0254]
[0255] Fig.18 FIG. 2 illustrates a modified reference sample used to generate a prediction sample according to another embodiment of the present disclosure. Fig.18 The upper left position of the current block 1810 is defined as (0, 0), and the left and top areas of the current block 1810 may be pre-reconstruction areas. Fig.18, when using a reference sample line located at a one sample distance to the left of the current block 1810 and a pre-reconstructed sample on a reference line located at a k+1 sample distance above the current block 1810 as reference samples, the present disclosure may perform a more accurate planar prediction mode using a top reference sample T' adjacent to the current block 1810 and an upper right reference sample TR' adjacent to the current block 1810. Here, T' and TR' may be reference samples included in the pre-reconstructed region and may be defined in various methods.
[0256] For example, T' may be a pre-reconstructed sample located at (x, -1), and TR' may be a pre-reconstructed sample located at (W, -1). Here, x may be the x coordinate of the prediction sample 1820 to be predicted, and W may be the width of the current block 1810. In another example, T' may be a pre-reconstructed sample T located at (x, -1-k). In other words, T' and T may be the same. In addition, TR' may be a pre-reconstructed sample TR located at (W, -1-k). In other words, TR' and TR may be the same.
[0257] In another example, T' and TR' may be defined based on a difference d (=qp) between a reference sample p located at (-1, -k-1) and a reference sample q located at (-1, -1). In other words, T' may be T+d and TR' may be TR+d. Here, d may represent the distance between the reference sample p and the reference sample q. In addition, d may be the difference between the sample values of the reference sample p and the reference sample q.
[0258] In another example, T' and TR' may be generated in horizontal plane mode using reference sample q and top right reference sample TR. In another example, T' and TR' may be generated in vertical plane mode using top reference sample T and bottom left reference sample BL. In another example, TR' may be generated in horizontal plane mode using reference sample q and a reference sample located at (W+1, -k-1).
[0259] In another example, T' and TR' may be generated as an average of a horizontal plane mode using the reference sample q and the upper right reference sample TR, and a vertical plane mode using the top reference sample T and the lower left reference sample BL. In other words, T' and TR' may be generated as an average of sample values of a sample generated using the reference sample q and the upper right reference sample TR in the horizontal plane mode and sample values of a sample generated using the top reference sample T and the lower left reference sample BL in the vertical plane mode.
[0260] As another example, TR' may be generated as an average of a horizontal plane mode using the reference sample q and the reference sample at the position (W+1, -k-1) and a vertical plane mode using the top reference sample T and the bottom left reference sample BL. In other words, TR' may be generated as an average of sample values of a sample generated in the horizontal plane mode using the reference sample q and the reference sample at the position (W+1, -k-1) and sample values of a sample generated in the vertical plane mode using the top reference sample T and the bottom left reference sample BL.
[0261] To generate prediction samples 1820 in the horizontal planar mode using T' and TR', the following formula 19 may be used. In formula 19, can be the weight applied to L. In addition, Can be a weight to be applied to TR'. and It can be calculated using the following formula 20.
[0262] [Formula 19]
[0263]
[0264] [Formula 20]
[0265]
[0266] To generate prediction samples 1820 in the vertical planar mode using T' and TR', the following formula 21 may be used. In formula 21, can be the weight applied to T'. In addition, It can be a weight applied to BL. and This can be calculated using the following formula 22.
[0267] [Formula 21]
[0268]
[0269] [Formula 22]
[0270]
[0271] Fig.19 is a flowchart of an image encoding / decoding method according to the present disclosure. Fig.19, the video encoding device (100) and / or the video decoding device (200) may determine a reference sample at a distance of n samples from the current block and a reference sample at a distance of m samples from the current block S1910. In this case, the reference sample at a distance of n samples may include a top reference sample and an upper right reference sample. In addition, the reference sample at a distance of m samples may include a left reference sample and a lower left reference sample. Here, n and m may be natural numbers.
[0272] The image encoding device 100 and / or the image decoding device 200 may generate a prediction block of the current block based on a weighted sum of at least two reference samples among the determined reference samples S1920. The weight used in the weighted sum may be determined based on n and / or m. In other words, the weight may be determined based on the distance between the current block and the reference sample.
[0273] According to an embodiment of the present disclosure, the image encoding device 100 and / or the image decoding device 200 may generate a plurality of prediction blocks based on a weighted sum of at least two reference samples among the determined reference samples. In this case, a final prediction block of the current block may be generated based on the generated prediction blocks. For example, a prediction block may be generated in a horizontal plane mode using a left reference sample and an upper right reference sample. In addition, a prediction block may be generated in a vertical plane mode using an upper reference sample and a lower left reference sample. A final prediction block may be generated based on a weighted sum of a prediction block generated in a horizontal plane mode and a prediction block generated in a vertical plane mode.
[0274] According to another embodiment of the present disclosure, the reference samples used to generate the prediction block may include reference samples generated by modifying the reference samples with a distance of n samples based on n. In addition, the reference samples used to generate the prediction block may include reference samples generated by modifying the reference samples with a distance of m samples based on m.
[0275] For example, the top reference sample and the upper right reference sample may be modified based on n. The prediction block of the current block may be generated based on at least two of the modified top reference sample, the modified upper right reference sample, the left reference sample, or the lower left reference sample. As another example, the left reference sample and the lower left reference sample may be modified based on m. The prediction block of the current block may be generated based on at least two of the modified left reference sample, the modified lower left reference sample, the top reference sample, or the upper right reference sample. When multiple prediction blocks are generated, the final prediction block may be generated by a weighted sum of the multiple prediction blocks.
[0276] According to another embodiment of the present disclosure, a reference sample for generating a prediction block of the current block may be generated in a planar mode based on at least two of a top reference sample, a left reference sample, an upper right reference sample, a lower left reference sample, an upper left reference sample adjacent to the current block, a reference sample adjacent to the lower left reference sample, or a reference sample adjacent to the upper right reference sample.
[0277] According to another embodiment of the present disclosure, the reference sample used to generate the prediction block of the current block may further include a reference sample generated by averaging sample values of a plurality of reference samples generated in a plane mode based on at least two of a top reference sample, a left reference sample, an upper right reference sample, a lower left reference sample, an upper left reference sample adjacent to the current block, a reference sample adjacent to the lower left reference sample, or a reference sample adjacent to the upper right reference sample. The plane mode described in the embodiment of the present disclosure may include a horizontal plane mode and a vertical plane mode.
[0278] Example 2
[0279] The present disclosure may relate to a method for mixing a prediction block predicted in a plane mode with another prediction block when performing an MRL mode. When a plane prediction block used for mixing is generated using a pre-reconstructed reference sample close to a current block, a more accurate plane prediction block may be generated. In other words, the prediction accuracy of the plane prediction block generated using the pre-reconstructed reference sample close to the current block may be high.
[0280] For example, when a reference sample line 0 adjacent to the current block is used, a more accurate plane prediction block can be generated. Here, the plane prediction block can be a prediction block predicted in the plane mode. In addition, by mixing the generated plane prediction block with another prediction block, a new predictor can be generated, thereby improving coding efficiency.
[0281] According to an embodiment of the present disclosure, when a planar mode is derived in a template-based intra mode derivation (TIMD) mode or a decoder-side intra mode derivation (DIMD) mode and an MRL mode is used, a blend may be performed between a prediction block predicted in the planar mode and another prediction block. Here, a reference sample line used to generate a prediction block using the MRL mode may be a reference sample line having an index greater than 0, such as a reference sample line having indexes of 1, 3, 5, 7, and 12.
[0282] Fig. 20 is a diagram illustrating a reference sample line for generating a prediction block according to the present disclosure. Fig. 20 In order to generate a prediction sample 2020 within the current block 2010, the image encoding device 100 and / or the image decoding device 200 may mix a prediction block predicted in a planar mode with a prediction block predicted using a reference sample line r.
[0283] According to an embodiment of the present disclosure, when the first mode derived in the TIMD mode or the DIMD mode is a planar mode, a prediction block can be generated using reference sample line 0 regardless of the MRL index sent by the signal used. Here, reference sample line 0 may refer to the first reference sample line adjacent to the current block. In addition, the second mode derived in the TIMD mode or the DIMD mode may generate a prediction block using reference sample line r. In other words, a prediction block may be generated using reference sample line r in the second mode derived in the TIMD mode or the DIMD mode. Here, r may be determined by an index explicitly sent by a signal. Therefore, a final prediction block may be generated by mixing prediction blocks generated using the above-mentioned first mode and prediction blocks generated using the second mode.
[0284] According to another embodiment of the present disclosure, when the second mode of the TIMD mode or the DIMD mode is a planar mode, a prediction block can be generated using reference sample line 0 regardless of the MRL index sent by the signal. In addition, for the first mode derived in the TIMD mode or the DIMD mode, a prediction block can be generated using reference sample line r. In other words, a prediction block can be generated using reference sample line r based on the first mode derived in the TIMD mode or the DIMD mode. Here, r can be determined by an index explicitly sent by the signal. Therefore, the final prediction block can be generated by mixing the prediction blocks generated by the above-mentioned first mode and the prediction blocks generated using the second mode.
[0285] According to another embodiment of the present disclosure, when both the first mode and the second mode derived in the TIMD mode or the DIMD mode are planar modes, the first prediction block may be generated in the planar mode using the reference sample line 0 regardless of the MRL index sent by the signal. In addition, the second prediction block may be generated in the planar mode using the reference sample line r. Here, r may be determined by an index explicitly sent by the signal.
[0286] According to another embodiment of the present disclosure, when the first mode and the second mode derived in the TIMD mode or the DIMD mode are different from each other and one of them is a planar mode, the final prediction block can be generated by mixing three prediction blocks. The first prediction block can be generated in the planar mode using reference sample line 0 regardless of the MRL index sent by the signal. The second prediction block can be generated in the planar mode using reference sample line r. The third prediction block can be generated in the non-planar mode using reference sample line r.
[0287] In various embodiments of the present disclosure, the weight used for the mixed prediction block may be an average value between prediction blocks, may be a predefined weight, or may be determined by a selective combination of these methods. In other words, the weight may be determined based on the average value of the prediction values between prediction blocks or at least one of the predefined weights. In addition, the weight may be determined based on whether the first mode of the TIMD mode or DIMD mode is a planar mode. Alternatively, the weight may be determined based on whether the second mode of the TIMD mode or DIMD mode is a planar mode. Alternatively, the weight may be determined based on whether the mode derived in the TIMD mode or DIMD mode includes a planar mode. Alternatively, the weight may be determined by a selective combination of these methods.
[0288] According to an embodiment of the present disclosure, the plane prediction block used for mixing may be a horizontal plane prediction block, a vertical plane prediction block, or an average value of a horizontal plane prediction block and a vertical plane prediction block. Here, the plane prediction block may be a prediction block predicted in a plane mode. The vertical plane prediction block may be a prediction block predicted in a vertical plane mode. The horizontal plane prediction block may be a prediction block predicted in a horizontal plane mode.
[0289] Example 3
[0290] The present disclosure may relate to a method for generating samples in a top template area and a left template area of a current block in a planar mode when using a TIMD mode. In other words, the present disclosure may relate to a method for generating prediction samples in a top template area or a left template area of a current block in a planar mode.
[0291] The TIMD mode may be a mode that uses a mode with a lowest template cost as an intra prediction mode of a current block after calculating a template cost between a prediction block predicted from a template region and an actual reconstructed sample. In this case, when the mode with the lowest template cost is a planar mode, in order to predict a sample at a position (x, y) within a current block having a width of W and a height of H in the planar prediction mode, a lower left pre-reconstructed sample at a position (-1, H), an upper right pre-reconstructed sample at a position (W, -1), a top pre-reconstructed sample at a position (x, -1), and / or a left pre-reconstructed sample at a position (-1, y) may be used.
[0292] Therefore, the present disclosure can obtain the template cost for the planar mode more accurately by performing the planar mode in the template area using the pre-reconstructed reference pixels close to the reference pixels used in the current block. In other words, in order to calculate a more accurate template cost in the template area, the planar mode can be performed using the pre-reconstructed reference samples close to the reference samples used in the current block.
[0293] Fig.21is a diagram illustrating reference samples used for prediction block generation based on the TIMD mode according to an embodiment of the present disclosure. Fig.21 The upper left position of the current block 2110 is defined as (0, 0). Fig.21 , the prediction sample may be generated in the planar mode using the top right reference sample TR of the top template region 2130 located at (W, -TH-1), the bottom left reference sample BL of the left template region 2140 located at (-TH-1, H), the top reference sample T of the sample to be predicted 2120 within the template region, and / or the left reference sample L of the sample to be predicted 2120 within the template region. Here, TH may be the height of the top template region 2130. In this case, the prediction sample may be generated based on the various formulas used in the above-described embodiment 1.
[0294] Fig. 22 is a diagram illustrating reference samples used for prediction block generation based on the TIMD mode according to another embodiment of the present disclosure. Fig. 22 The upper left position of the top template area 2230 is defined as (0, 0). Fig. 22 , the accurate planar prediction samples 2220 may be generated using reference samples L′ and BL′ located adjacent to the top template region 2230 of the current block 2210 .
[0295] In this case, L' and BL' can be defined in various ways. For example, L' can be a pre-reconstructed sample at position (-1, y). In addition, BL' can be a pre-reconstructed sample at position (-1, TH). Here, TH can be the height of the top template area 2230. As another example, L' and BL' can be determined based on the difference d (=pq) between the reference sample p located at (-TW-1, -1) relative to the top template area 2230 and the reference sample q located at (-1, -1). L' can be L+d and BL' can be BL+d. Here, TW can be the width of the left template area or the distance between the current block and the left reference sample line. In addition, d can be the distance between the reference sample p and the reference sample q, or d can be the difference between the sample values of the reference sample p and the reference sample q. The predicted sample 2220 can be generated using Formula 15 and / or Formula 17.
[0296] As another example, L' and BL' may be generated in horizontal plane mode using left reference sample L and upper right reference sample TR. As another example, L' and BL' may be generated in vertical plane mode using reference sample q and lower left reference sample BL. As another example, BL' may be generated in vertical plane mode using reference sample q and a sample at (-TW-1, TH+1).
[0297] As another example, L' and BL' may be generated as an average of a horizontal plane mode using the left reference sample L and the upper right reference sample TR and a vertical plane mode using the reference sample q and the lower left reference sample BL. In other words, L' and BL' may be generated as an average of sample values of a sample generated in the horizontal plane mode using the left reference sample L and the upper right reference sample TR and sample values of a sample generated in the vertical plane mode using the reference sample q and the lower left reference sample BL.
[0298] As another example, BL' may be generated as an average of a horizontal plane mode using the lower left reference sample BL and the upper right reference sample TR and a vertical plane mode using the reference sample q and the reference sample at the position (-TW-1, TH+1). In other words, BL' may be generated as an average of sample values of a sample generated in the horizontal plane mode using the lower left reference sample BL and the upper right reference sample TR and sample values of a sample generated in the vertical plane mode using the reference sample q and the reference sample at the position (-TW-1, TH+1).
[0299] Fig.23 is a diagram illustrating reference samples used for prediction block generation based on the TIMD mode according to another embodiment of the present disclosure. Fig.23 The upper left position of the left template area 2330 is defined as (0, 0). Fig.23 , the accurate planar prediction samples 2320 may be generated using reference samples T′ and TR′ located adjacent to the left template region 2330 of the current block 2310 .
[0300] T' and TR' may be defined in various ways. For example, T' may be a pre-reconstructed sample at position (x, -1). In addition, TR' may be a pre-reconstructed sample at position (TW, -1). Here, TW may be the width of the left template area 2330. As another example, T' and TR' may be determined based on the difference d (=pq) between a reference sample p located at (-1, -TH-1) and a reference sample q located at (-1, -1) relative to the left template area 2330. T' may be T+d, and TR' may be TR+d. Here, TH may be the height of the top template area or the distance between the current block and the top reference sample line. In addition, d may be the distance between reference sample p and reference sample q, or the difference between the sample values of reference sample p and reference sample q. The predicted sample 2320 may be generated in planar mode using T', TR', L, and BL, and may be generated using Formula 19 and / or Formula 21.
[0301] As another example, T' and TR' may be generated in horizontal plane mode using reference sample q and top right reference sample TR. As another example, T' and TR' may be generated in vertical plane mode using top reference sample T and bottom left reference sample BL. As another example, TR' may be generated in horizontal plane mode using reference sample q and a sample at position (TW+1, -TH-1).
[0302] As another example, T' and TR' may be generated as an average of a horizontal plane mode using the reference sample q and the upper right reference sample TR and a vertical plane mode using the top reference sample T and the lower left reference sample BL. In other words, T' and TR' may be generated as an average of sample values of samples generated in the horizontal plane mode using the reference sample q and the upper right reference sample TR and sample values of samples generated in the vertical plane mode using the top reference sample T and the lower left reference sample BL.
[0303] As another example, TR' may be generated as an average of a horizontal plane mode using the reference sample q and a sample at a position (TW+1, -TH-1) and a vertical plane mode using the top reference sample T and the lower left reference sample BL. In other words, TR' may be generated as an average of sample values of samples generated in the horizontal plane mode using the reference sample q and the sample at the position (TW+1, -TH-1) and sample values of samples generated in the vertical plane mode using the top reference sample T and the lower left reference sample BL.
[0304] The above embodiment may include a horizontal plane mode or a vertical plane mode as an intra-mode candidate for calculating a mode with a minimum template cost in a left and / or top template region of a TIMD mode. In addition, when the mode with the minimum template cost is a horizontal plane mode, the current block may be predicted in the horizontal plane mode. Alternatively, when the mode with the minimum template cost is a vertical plane mode, the current block may be predicted in the vertical plane mode.
[0305] Fig.24 is a flowchart of an image encoding / decoding method according to the present disclosure. Fig.24 , the image encoding device 100 and / or the image decoding device 200 may derive the first prediction mode and the second prediction mode for the current block based on the template matching cost S2410. In this case, the template matching cost calculated based on the first prediction mode may be less than the template matching cost calculated based on the second prediction mode.
[0306] The template matching cost may be calculated based on a prediction block resulting from predicting a template region of the current block in a planar mode. A top reference sample and an upper right reference sample used to predict the template region may be obtained from a first reference sample line adjacent to the top template region, and a left reference sample and a lower left reference sample used to predict the template region may be obtained from a first reference sample line adjacent to the left template region.
[0307] According to an embodiment of the present disclosure, the template region may include a top template region and a left template region. In this case, the lower left reference sample used to predict the top template region may be the same as the lower left reference sample of the left template region, and the upper right reference sample used to predict the left template region may be the same as the upper right reference sample of the top template region.
[0308] According to another embodiment of the present disclosure, the template area may include a top template area and a left template area. In this case, the lower left reference sample used to predict the top template area may have the same y coordinate as the lower left sample of the top template area, and the upper right reference sample used to predict the left template area may have the same x coordinate as the upper right sample of the left template area. The values of the left reference sample and the lower left reference sample used to predict the top template area may be modified based on the width of the left template area, and the values of the top reference sample and the upper right reference sample used to predict the left template area may be modified based on the height of the top template area.
[0309] The image encoding device 100 and / or the image decoding device 200 may generate a first prediction block based on the first prediction mode and the first reference sample line, and generate a second prediction block based on the second prediction mode and the second reference sample line S2420. Here, at least one of the first reference sample line and the second reference sample line may be determined based on whether the first prediction mode or the second prediction mode is a planar mode.
[0310] According to an embodiment of the present disclosure, when the first prediction mode is the planar mode, the first reference sample line may be determined as the first reference sample line adjacent to the current block, and the second reference sample line may be determined as the rth reference sample line adjacent to the current block. According to another embodiment of the present disclosure, when the second prediction mode is the planar mode, the first reference sample line may be determined as the rth reference sample line adjacent to the current block, and the second reference sample line may be determined as the first reference sample line adjacent to the current block.
[0311] The image encoding device 100 and / or the image decoding device 200 may generate a final prediction block of the current block based on the first prediction block and the second prediction block S2430. According to an embodiment of the present disclosure, when the first prediction mode and the second prediction mode include a planar mode and a non-planar mode, the final prediction block may be generated by performing a weighted sum of a first prediction block generated in the planar mode based on a first reference sample line adjacent to the current block, a second prediction block generated in the planar mode based on an r-th reference sample line adjacent to the current block, and a third prediction block generated in the non-planar mode based on an r-th reference sample line. The weighted sum according to the present disclosure may be performed based on the formula used in the above-mentioned embodiment.
[0312] For the sake of clarity of explanation, the exemplary method of the present disclosure is described as a series of operations, but this is not intended to limit the order in which the steps are performed, and each step can be performed simultaneously or in a different order when necessary. In order to implement the method according to the present disclosure, in addition to the steps illustrated, additional steps may be included, some steps may be omitted while the remaining steps are included, or some steps may be omitted while the additional steps are included.
[0313] In the present disclosure, the image encoding device 100 or the image decoding device 200 that performs a predetermined operation (step) may perform the operation (step) to check the execution condition or situation of the corresponding operation (step). For example, when it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device 100 or the image decoding device 200 may perform an operation to check whether the predetermined condition is satisfied before performing the predetermined operation.
[0314] The various embodiments of the present disclosure are not a list of all possible combinations but are provided to illustrate representative aspects of the present disclosure, and the elements described in the various embodiments may be applied independently or in combination of two or more elements.
[0315] In addition, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. When implemented in hardware, various embodiments may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, microcontrollers, or microprocessors.
[0316] In addition, the image decoding device 200 and the image encoding device 100 to which the embodiments of the present disclosure are applied may be included in various devices such as multimedia broadcast transmission / reception devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication devices, mobile streaming devices, storage media, cameras, video on demand (VoD) service providing devices, over-the-top video (OTT) devices, Internet streaming service providing devices, three-dimensional (3D) video devices, video phone devices, and medical video devices, and may be used to process video signals or data signals. For example, over-the-top video (OTT) devices may include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smart phones, tablet computers, and digital video recorders (DVRs), etc.
[0317] Fig.25 An exemplary diagram showing a content streaming system to which embodiments of the present disclosure can be applied.
[0318] like Fig.25 As shown in , a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0319] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data, generates a bitstream, and sends it to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server can be omitted.
[0320] A bitstream may be generated by applying the video encoding method and / or the image encoding device 100 according to an embodiment of the present disclosure, and a streaming server may temporarily store the bitstream during a process of transmitting or receiving the bitstream.
[0321] The streaming server can send multimedia data to the user device based on the user request through the web server, and the web server can act as a medium to inform the user of available services. When the user requests the required service from the web server, the web server can send the request to the streaming server, and the streaming server can transmit the multimedia data to the user. In this case, the content streaming system can include a separate control server, and in this case, the control server can play a role in controlling the command / response exchange between the devices within the content streaming system.
[0322] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, in order to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.
[0323] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (i.e., smart watches, smart glasses, head-mounted displays (HMDs)), digital televisions, desktop computers, and digital signage.
[0324] Each server in the content streaming system may operate as a distributed server, in which case data received by each server may be processed in a distributed manner.
[0325] The scope of the present disclosure includes software or machine-executable instructions (i.e., operating systems, applications, firmware, programs, etc.) that enable operations of methods according to various embodiments to be performed on a device or computer, as well as non-transitory computer-readable media in which such software or instructions are stored and executable on a device or computer.
[0326] [Industrial Applicability]
[0327] The embodiments of the present disclosure may be used to encode / decode images.
Claims
1. An image decoding method performed by an image decoding device, comprising: determining a reference sample located at a n-sample distance from a current block and a reference sample located at a m-sample distance from the current block, wherein the reference samples located at the n-sample distance include a top reference sample and an upper-right reference sample, and the reference samples located at the m-sample distance include a left reference sample and a lower-left reference sample; and generating a prediction block of the current block based on a weighted sum of at least two reference samples among the determined reference samples, wherein the weight used in the weighted sum is determined based on the n or the m, and Wherein, the n and the m are natural numbers.
2. The image decoding method according to claim 1, wherein: Generating a prediction block of the current block includes: generating a prediction block of the current block based on the weighted sum of the reference samples; and A final prediction block of the current block is generated based on the prediction block.
3. The image decoding method according to claim 2, wherein: The reference samples further include: a reference sample generated by modifying the reference sample located at the n-sample distance based on the n, and a reference sample generated by modifying the reference sample located at the m-sample distance based on the m.
4. The image decoding method according to claim 2, wherein: The reference samples further include reference samples generated in a planar mode based on at least two of the top reference sample, the left reference sample, the upper right reference sample, the lower left reference sample, an upper left reference sample adjacent to the current block, a reference sample adjacent to the lower left reference sample, or a reference sample adjacent to the upper right reference sample, and The plane mode includes a horizontal plane mode and a vertical plane mode.
5. The image decoding method according to claim 2, wherein: The reference sample further includes a reference sample generated by averaging sample values of a plurality of reference samples generated in a planar mode based on at least two of the top reference sample, the left reference sample, the upper right reference sample, the lower left reference sample, an upper left reference sample adjacent to the current block, a reference sample adjacent to the lower left reference sample, or a reference sample adjacent to the upper right reference sample, and The plane mode includes a horizontal plane mode and a vertical plane mode.
6. An image decoding method performed by an image decoding device, comprising: deriving a first prediction mode and a second prediction mode for the current block based on a template matching cost, wherein the template matching cost calculated based on the first prediction mode is less than the template matching cost calculated based on the second prediction mode; Generate a first prediction block based on the first prediction mode and a first reference sample line, and generate a second prediction block based on the second prediction mode and a second reference sample line; and generating a final prediction block of the current block based on the first prediction block and the second prediction block, At least one of the first reference sample line or the second reference sample line is determined based on whether the first prediction mode or the second prediction mode is a planar mode.
7. The image decoding method according to claim 6, wherein: Based on the first prediction mode being the planar mode, the first reference sample line is determined as a first reference sample line adjacent to the current block, and the second reference sample line is determined as an rth reference sample line adjacent to the current block.
8. The image decoding method according to claim 6, wherein: Based on the second prediction mode being the planar mode, the first reference sample line is determined as the rth reference sample line adjacent to the current block, and the second reference sample line is determined as the first reference sample line adjacent to the current block.
9. The image decoding method according to claim 6, wherein: Based on that the first prediction mode and the second prediction mode include a planar mode and a non-planar mode, the final prediction block is generated by performing a weighted sum of a first prediction block generated in the planar mode based on a first reference sample line adjacent to the current block, a second prediction block generated in the planar mode based on an rth reference sample line adjacent to the current block, and a third prediction block generated in the non-planar mode based on the rth reference sample line.
10. The image decoding method according to claim 6, wherein: The template matching cost is calculated based on a prediction block resulting from a template region predicting the current block in the planar mode, wherein the top reference sample and the upper right reference sample used for prediction of the template region are obtained from a first reference sample line adjacent to the top template region, and The left reference sample and the lower left reference sample used for prediction of the template region are obtained from a first reference sample line adjacent to the left template region.
11. The image decoding method according to claim 10, wherein: The template area includes a top template area and a left template area. wherein the lower left reference sample used for prediction of the top template region is the same as the lower left reference sample of the left template region, and The upper right reference sample used for prediction of the left template region is the same as the upper right reference sample of the top template region.
12. The image decoding method according to claim 10, wherein: The template area includes a top template area and a left template area. wherein the lower left reference sample used for prediction of the top template region has the same y coordinate as the lower left sample of the top template region, and The upper right reference sample used for prediction of the left template region has the same x coordinate as the upper right sample of the left template region.
13. The image decoding method according to claim 12, wherein: Based on the width of the left template region, modifying the values of the left reference samples and the values of the bottom left reference samples used for prediction of the top template region, and Wherein, based on the height of the top template area, the value of the top reference sample and the value of the upper right reference sample used for prediction of the left template area are modified.
14. An image encoding method performed by an image encoding device, comprising: deriving a first prediction mode and a second prediction mode for the current block based on the template matching cost; generating a first prediction block based on the first prediction mode and a first reference sample line; generating a second prediction block based on the second prediction mode and the second reference sample line; as well as generating a final prediction block of the current block based on the first prediction block and the second prediction block, At least one of the first reference sample line or the second reference sample line is determined based on whether the first prediction mode or the second prediction mode is a planar mode.
15. A computer-readable recording medium storing a bit stream generated by the image encoding method according to claim 14.
16. A method for transmitting a bit stream generated by an image encoding method, in, The image encoding method comprises: deriving a first prediction mode and a second prediction mode for the current block based on the template matching cost; generating a first prediction block based on the first prediction mode and a first reference sample line; generating a second prediction block based on the second prediction mode and the second reference sample line; and generating a final prediction block of the current block based on the first prediction block and the second prediction block, At least one of the first reference sample line or the second reference sample line is determined based on whether the first prediction mode or the second prediction mode is a planar mode.