A video encoding / decoding method based on intra-template matching, a method for transmitting a bitstream, and a recording medium storing the bitstream.
The video encoding/decoding method with intra-template matching addresses the high data volume challenge of high-resolution images by improving encoding/decoding efficiency and reducing transmission/storage costs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2024-04-22
- Publication Date
- 2026-05-19
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to higher transmission and storage costs due to increased data volume, necessitating a more efficient image compression technique.
A video encoding/decoding method utilizing intra-template matching for improved encoding/decoding efficiency, including a template search area and methods to derive values of unavailable samples within reference blocks.
Enhances predictive performance and coding efficiency through improved intra-template matching, enabling efficient storage and transmission of high-resolution images.
Smart Images

Figure 2026515848000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a video encoding / decoding method, a method for transmitting a bitstream, and a recording medium storing a bitstream, and relates to prediction based on intra-template matching.
Background Art
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared with conventional image data. The increase in the amount of information or bits to be transmitted brings about an increase in transmission costs and storage costs.
[0003] Thus, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images.
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present disclosure is to provide a video encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to propose a template search area for effectively generating a template matching prediction block.
[0006] Another object of the present disclosure is to propose a method for deriving values of samples that are not available within a reference block or a reference template area.
[0007] Another object of the present disclosure is to provide a non-temporary computer-readable recording medium storing a bitstream generated by the video encoding method according to the present disclosure.
[0008] Furthermore, this disclosure aims to provide a non-temporary computer-readable recording medium that stores a bitstream received and decoded by the video decoding device relating to this disclosure and used for video restoration.
[0009] Furthermore, this disclosure aims to provide a method for transmitting a bitstream generated by the video encoding method relating to this disclosure.
[0010] The technical challenges that this disclosure seeks to address are not limited to those mentioned above, and any other technical challenges not mentioned can be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Means for solving the problem]
[0011] A video decoding method according to one aspect of the present disclosure is a video decoding method performed by a video decoding device, which includes (composes; constructs; sets; includes; encompasses; contains; contains; has) a reference block of the current block based on the error.
[0012] Other aspects of the present disclosure may be video encoding methods performed by a video encoding device, the methods comprising: determining a search region associated with a current block; inducing an error between a reference template region which is the template region of each reference block in the search region and a current template region which is the template region of the current block; and determining a reference block of the current block based on the error.
[0013] Computer-readable recording media relating to other aspects of this disclosure can store bitstreams generated by the video encoding method or apparatus of this disclosure.
[0014] A transmission method relating to yet another aspect of the present disclosure may transmit a bitstream generated by the video encoding method or apparatus of the present disclosure.
[0015] The features of this disclosure briefly summarized above are merely illustrative aspects of the detailed description of this disclosure described below and do not limit the scope of this disclosure. [Effects of the Invention]
[0016] According to this disclosure, a video encoding / decoding method and apparatus with improved encoding / decoding efficiency may be provided.
[0017] Furthermore, this disclosure suggests that the predictive performance of intra-template matching may be improved.
[0018] Furthermore, according to this disclosure, coding efficiency may improve as the predictive performance of intra-template matching improves.
[0019] Furthermore, this disclosure may provide a non-temporary computer-readable recording medium for storing a bitstream generated by the video encoding method relating to this disclosure.
[0020] Furthermore, this disclosure may provide a non-temporary computer-readable recording medium for storing a bitstream that is received and decoded by the video decoding device relating to this disclosure and used for restoring video.
[0021] Furthermore, this disclosure may provide a method for transmitting a bitstream generated by a video encoding method.
[0022] The effects that can be obtained in the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those with ordinary knowledge in the technical field to which the present disclosure belongs from the following description.
Brief Description of Drawings
[0023] [Figure 1] It is a diagram schematically showing a video coding system to which an embodiment according to the present disclosure can be applied. [Figure 2] It is a diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure can be applied. [Figure 3] It is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied. [Figure 4] It is a flowchart showing a video coding method based on intra prediction. [Figure 5] It is a drawing schematically showing an intra prediction unit in a video coding apparatus. [Figure 6] It is a flowchart showing a video coding method based on intra prediction. [Figure 7] It is a drawing schematically showing an intra prediction unit in a video decoding apparatus. [Figure 8-11] It is a drawing for explaining an example of a search area for intra template matching and block vector candidates. [Figure 12-13] It is a drawing for explaining an example of a reference block outside the search area. [Figure 14-15] It is a flowchart showing a video coding method and a video decoding method according to an embodiment of the present disclosure. [Figure 16] It is a drawing for explaining an example of a search position in integer sample units and a search position in fractional sample units. [Figure 17] It is a drawing for explaining an example of a block vector for a search area. [Figure 18-25] It is a drawing for explaining an example of inducing non-available sample values. [Figure 26-27]This is a flowchart showing a video encoding method and a video decoding method according to another embodiment of the present disclosure. [Figure 28] This figure illustrates a content streaming system to which the embodiments described herein can be applied. [Modes for carrying out the invention]
[0024] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings, so that they can be easily implemented by a person with ordinary skill in the art to which the present disclosure pertains. However, the present disclosure can be implemented in a variety of different forms and is not limited to the embodiments described herein.
[0025] In describing embodiments of this disclosure, if it is determined that a specific description of a known configuration or function would obscure the gist of this disclosure, such detailed description will be omitted. In the drawings, parts unrelated to the description of this disclosure will be omitted, and similar parts will be denoted by the same reference numerals.
[0026] In this disclosure, when one component is described as being “connected,” “joined,” or “linked” to another component, this can include not only direct connections but also indirect connections where another component exists between them. Furthermore, when one component is described as “containing” or “having” another component, this means, unless otherwise stated to the contrary, that it may include another component rather than excluding it.
[0027] In this disclosure, terms such as "first," "second," etc., are used solely for the purpose of distinguishing one component from another, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be called a second component in another embodiment, and similarly, a second component in one embodiment may be called a first component in another embodiment.
[0028] In this disclosure, components that are distinguished from each other are used to clearly describe their respective characteristics and do not necessarily mean that the components are separate. In other words, multiple components may be integrated to constitute a single hardware or software unit, or a single component may be distributed to constitute multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of this disclosure, without needing to be specifically mentioned.
[0029] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of this disclosure. Furthermore, embodiments that include additional components in addition to the components described in various embodiments are also included in the scope of this disclosure.
[0030] This disclosure relates to the encoding and decoding of images, and the terms used in this disclosure may have their ordinary meanings in the art to which this disclosure pertains, unless otherwise defined herein.
[0031] In this disclosure, "picture" generally means a unit representing any one image within a specific time period, and "slice / tile" is an encoding unit that constitutes part of a picture, and a single picture can consist of one or more slices / tiles. Furthermore, a slice / tile may contain one or more CTUs (coding tree units).
[0032] In this disclosure, “pixel” or “pel” may mean the smallest unit that constitutes a picture (or image). The term “sample” may also be used as a counterpart to pixel. A sample may generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0033] In this disclosure, “unit” can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. A unit may be used interchangeably with terms such as “sample array,” “block,” or “area,” as it may be used. Generally, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0034] In this disclosure, “current block” can mean any one of the following: “current coding block,” “current coding unit,” “block to encode,” “block to decode,” or “block to process.” If prediction is performed, “current block” can mean “current prediction block” or “block to predict.” If transformation (inverse transformation) / quantization (inverse quantization) is performed, “current block” can mean “current transformation block” or “block to transform.” If filtering is performed, “current block” can mean “block to filter.”
[0035] Furthermore, in this disclosure, "current block" may mean the block containing all of the rumor component blocks and chroma component blocks, or the "rumor block of the current block," unless there is an explicit mention of a chroma block. The rumor component block of the current block may be expressed with an explicit mention of a rumor component block, such as "rumor block" or "current rumor block." Similarly, the chroma component block of the current block may be expressed with an explicit mention of a chroma component block, such as "chroma block" or "current chroma block."
[0036] In this disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."
[0037] In this disclosure, “or” may be interpreted as “and / or.” For example, “A or B” may mean 1) “A” only, 2) “B” only, or 3) “A and B.” Alternatively, in this disclosure, “or” may mean “additionally or alternatively.”
[0038] Overview of the video coding system
[0039] Figure 1 is a schematic diagram showing a video coding system to which the embodiments of this disclosure can be applied.
[0040] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 via a digital storage medium or network in file or streaming format.
[0041] An encoding device 10 according to one embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to one embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be called a video / image encoding unit, and the decoding unit 22 may be called a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may also include a display unit, which may be configured as a separate device or external component.
[0042] The video source generation unit 11 can acquire video / images through processes such as video / image capture, synthesis, or generation. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. The video / image generation device may include, for example, a computer, tablet, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by a process in which the relevant data is generated.
[0043] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of steps such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in bitstream format.
[0044] The transmission unit 13 can acquire encoded video / image information or data output in bitstream format and transmit it in file or streaming format to the receiving unit 21 of the decoding device 20 or other external object via a digital storage medium or network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmission unit 13 can include elements for generating media files via a predetermined file format and elements for transmission via a broadcast / communication network. The transmission unit 13 can be provided as a transmission device separate from the encoding device 12, in which case the transmission device can include at least one processor that acquires encoded video / image information or data output in bitstream format, and a transmission unit that transmits it in file or stream format. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.
[0045] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, corresponding to the operation of the encoding unit 12.
[0046] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0047] Overview of Image Encoding Devices
[0048] Figure 2 is a schematic diagram showing an image encoding device to which the embodiments of this disclosure can be applied.
[0049] As shown in Figure 2, the image coding device 100 may include an image splitting unit 110, a subtraction unit 115, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter-prediction unit 180, an intra-prediction unit 185, and an entropy coding unit 190. The inter-prediction unit 180 and the intra-prediction unit 185 can together be called the "prediction unit". The transformation unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.
[0050] All or at least some of the multiple components constituting the image encoding device 100 can be implemented by a single hardware component (e.g., an encoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB (decoded picture buffer) and can be implemented by a digital storage medium.
[0051] The image splitting unit 110 can split an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). Coding units can be obtained by recursively splitting a coding tree unit (CTU) or the largest coding unit (LCU) using a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, a single coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure and / or a ternary-tree structure. For the splitting of coding units, a quad-tree structure may be applied first, followed by a binary-tree structure and / or a ternary-tree structure. Based on the final coding unit that cannot be further split, the coding procedure according to this disclosure can be performed. The largest coding unit can be used as the final coding unit, or a lower-depth coding unit obtained by dividing the largest coding unit can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, as described later. As another example, the processing units of the coding procedure may be prediction units (PU) or transformation units (TU). The prediction unit and the transformation unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.
[0052] The prediction unit (inter-prediction unit 180 or intra-prediction unit 185) can make predictions for the block to be processed (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or on a CU basis. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy coding unit 190. The prediction information can be encoded by the entropy coding unit 190 and output in bitstream format.
[0053] The intra-prediction unit 185 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) or at a distance from the current block, according to the intra-prediction mode and / or intra-prediction technique. The intra-prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0054] The interprediction unit 180 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, collocated CU (colCU), etc. The reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 180 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 180 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, the residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding block is used as the motion vector predictor, and the motion vector of the current block can be signaled by encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference can represent the difference between the motion vector of the current block and the motion vector predictor.
[0055] The prediction unit can generate a prediction signal based on various prediction methods and / or techniques described later. For example, the prediction unit can apply intra-prediction or inter-prediction to predict the current block, and can also apply intra-prediction and inter-prediction simultaneously. A prediction method that applies intra-prediction and inter-prediction simultaneously to predict the current block can be called CIIP (combined inter and intra prediction). The prediction unit can also perform intra-block copy (IBC) to predict the current block. Intra-block copy can be used for content image / video coding such as in games, for example, in SCC (screen content coding). IBC is a method of predicting the current block using a reference block that has already been restored in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but it can be performed similarly to inter-prediction in that it derives the reference block within the current picture. In other words, IBC can use at least one of the interpretation techniques described in this disclosure.
[0056] The predicted signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtraction unit 115 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal output from the prediction unit (predicted block, predicted sample array) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit 120.
[0057] The transformation unit 120 can generate transformation coefficients by applying transformation techniques to the residual signal. For example, the transformation technique may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph, where the relationship information between pixels is represented by a graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels. The transformation process can be applied to pixel blocks of the same size and square shape, or to non-square, variable-sized blocks.
[0058] The quantization unit 130 can quantize the conversion coefficients and transmit them to the entropy coding unit 190. The entropy coding unit 190 can encode the quantized signal (information about the quantized conversion coefficients) and output it in bitstream format. The information about the quantized conversion coefficients can be called residual information. The quantization unit 130 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector format based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector format of the quantized conversion coefficients.
[0059] The entropy coding unit 190 can perform various coding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy coding unit 190 can also encode information necessary for video / image restoration (e.g., the values of syntax elements) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream format in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. The signaling information, transmitted information and / or syntax elements referred to in this disclosure may be encoded via the encoding procedure described above and included in the bitstream.
[0060] The bitstream can be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storage unit (not shown) for storing it may be provided as internal / external elements of the image encoding device 100, or the transmission unit may be provided as a component of the entropy encoding unit 190.
[0061] The quantized conversion coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be reconstructed.
[0062] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 180 or the intra-prediction unit 185. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 155 may be called the reconstruction unit or the reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.
[0063] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 160 can generate various filtering-related information, as will be described later in the explanation of each filtering method, and transmit it to the entropy coding unit 190. The filtering-related information can be encoded by the entropy coding unit 190 and output in bitstream format.
[0064] The corrected restored picture transmitted to memory 170 can be used as a reference picture in the interpretation unit 180. When interpretation is applied via this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.
[0065] The DPB in memory 170 can store the modified restored picture for use as a reference picture in the inter-prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 180 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 185.
[0066] Overview of the image decoding device
[0067] Figure 3 is a schematic diagram showing an image decoding apparatus to which the embodiments of this disclosure can be applied.
[0068] As shown in Figure 3, the image decoding device 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an additive unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265. The inter-prediction unit 260 and the intra-prediction unit 265 can together be called the "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in the residual processing unit.
[0069] All or at least some of the multiple components constituting the image decoding device 200 can be implemented by a single hardware component (e.g., a decoder or processor) according to the embodiment. Furthermore, the memory 170 may include a DPB and can be implemented by a digital storage medium.
[0070] An image decoding device 200, having received a bitstream containing video / image information, can restore the image by executing a process corresponding to the process performed in the image encoding device 100 in Figure 2. For example, the image decoding device 200 can perform decoding using the processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The restored image signal decoded and output via the image decoding device 200 can then be reproduced via a playback device (not shown).
[0071] The image decoding device 200 can receive the signal output from the image encoding device 2 in bitstream format. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The image decoding device may further use the parameter set information and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements referred to in this disclosure can be obtained from the bitstream by decoding via the decoding procedure. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image reconstruction and the quantized values of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the blocks to be decoded, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction unit (inter-prediction unit 260 and intra-prediction unit 265), and the residual values that have undergone entropy decoding in the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, can be input to the inverse quantization unit 220. In addition, of the information decoded by the entropy decoding unit 210, information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives signals output from the image coding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.
[0072] On the other hand, the image decoding device according to this disclosure may be called a video / image / picture decoding device. The image decoding device may also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265.
[0073] The inverse quantization unit 220 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 220 can rearrange the quantized transformation coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed by the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) to obtain the transformation coefficients.
[0074] The inverse conversion unit 230 can inversely convert the conversion coefficients to obtain residual signals (residual blocks, residual sample arrays).
[0075] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode (prediction technique).
[0076] As described in the explanation of the prediction unit of the image coding device 100, the prediction unit can generate prediction signals based on various prediction methods (techniques) described later.
[0077] The intra-prediction unit 265 can predict the current block by referring to the samples in the current picture. The description of the intra-prediction unit 185 can also be applied to the intra-prediction unit 265.
[0078] The interprediction unit 260 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in block, sub-block, or sample units based on the correlation of motion information between surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In interprediction, surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 260 can construct a motion information candidate list based on surrounding blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes (techniques), and the prediction information may include information indicating the mode (technique) of interprediction for the current block.
[0079] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 260 and / or intra-prediction unit 265). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 is sometimes called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or for inter-prediction of the next picture via filtering, as described later.
[0080] The filtering unit 240 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0081] The restored picture stored (modified) in the DPB of memory 250 can be used as a reference picture in the inter-prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 250 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 265.
[0082] In this specification, the embodiments described for the filtering unit 160, inter-prediction unit 180, and intra-prediction unit 185 of the image coding device 100 can be applied similarly or in a corresponding manner to the filtering unit 240, inter-prediction unit 260, and intra-prediction unit 265 of the image decoding device 200, respectively.
[0083] Intra Prediction
[0084] Intra prediction can indicate a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter, the current picture). When the intra example is applied to the current block, surrounding reference samples to be used for intra prediction of the current block can be derived. The surrounding reference samples of the current block may include a total of 2 x nH samples adjacent to the left boundary of the current block of size nW x nH and adjacent to the bottom-left, a total of 2 x nW samples adjacent to the top boundary of the current block and adjacent to the top-right, and one sample adjacent to the top-left of the current block. Alternatively, the surrounding reference samples of the current block may include upper surrounding samples in multiple columns and left surrounding samples in multiple rows. Furthermore, the surrounding reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block, which is nW x nH in size, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right side of the current block.
[0085] However, some of the surrounding reference samples in the current block may not yet be decoded or available. In this case, the video decoder 200 can substitute the unavailable samples with available samples to construct the surrounding reference samples to be used for prediction. Alternatively, it can construct the surrounding reference samples to be used for prediction through interpolation of available samples.
[0086] If peripheral reference samples are derived, (i) a predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, or (ii) the predicted sample can be derived based on a reference sample among the peripheral reference samples of the current block that is located in a specific (predicted) direction relative to the predicted sample. Case (i) may be called a non-directional mode or non-angular mode, and case (ii) may be called a directional mode or angular mode. Alternatively, the predicted sample may be generated by interpolation between the first peripheral sample and a second peripheral sample located in the opposite direction to the prediction direction of the intra-prediction mode of the current block, with respect to the predicted sample of the current block. In the above case, it may be called linear interpolation intra-prediction (LIP). Alternatively, a chroma predicted sample may be generated based on a chroma sample using a linear model. In this case, it may be called LM mode. Alternatively, a temporary predicted sample for the current block may be derived based on filtered peripheral reference samples, and the predicted sample for the current block may be derived by weighting the existing peripheral reference samples, i.e., at least one reference sample derived by the intra-prediction mode from among the unfiltered peripheral reference samples, with the temporary predicted sample. In the above case, this may be called PDPC (Position dependent intra prediction). Alternatively, intra-predictive coding can be performed by selecting the reference sample line with the highest prediction accuracy among the peripheral multi-reference sample lines of the current block, deriving the predicted sample using the reference sample located in the prediction direction on that line, and then instructing (signaling) the video decoding device 200 of the reference sample line used. In the above case, this may be called multi-reference line (MRL) intra prediction or MRL-based intra-prediction.Furthermore, while intra-prediction is performed based on the same intra-prediction mode for dividing the current block into vertical or horizontal subpartitions, surrounding reference samples can be derived and utilized on a subpartition-by-subpartition basis. That is, in this case, although the intra-prediction mode for the current block is applied identically to the subpartition, the intra-prediction performance can be improved in some cases by deriving and utilizing surrounding reference samples on a subpartition-by-subpartition basis. Such a prediction method may be called intra-subpartitions (ISP) or ISP-based intra-prediction. The intra-prediction method described above may be called an intra-prediction type to distinguish it from the intra-prediction mode in section 1.2 of the table of contents. The intra-prediction type may be referred to by various terms such as intra-prediction technique or additional intra-prediction mode. For example, the intra-prediction type (or additional intra-prediction mode, etc.) may include at least one of the aforementioned LIP, PDPC, MRL, and ISP. General intra-prediction methods excluding specific intra-prediction types such as LIP, PDPC, MRL, and ISP may be called normal intra-prediction types. The normal intra-prediction type can be generally applied when the specific intra-prediction types described above are not applicable, and predictions can be performed based on the intra-prediction modes described above. On the other hand, post-processing filtering may be performed on the derived prediction samples as needed.
[0087] Specifically, the intra-prediction procedure may include an intra-prediction mode / type determination stage, a peripheral reference sample derivation stage, and an intra-prediction mode / type-based prediction sample derivation stage. Furthermore, a post-filtering stage may be performed on the derived prediction samples as needed.
[0088] On the other hand, in addition to the intra-prediction types described above, affine linear weighted intra-prediction (ALWIP) may also be used. ALWIP may also be called LWIP (linear weighted intra-prediction) or MIP (matrix weighted intra-prediction or matrix-based intra-prediction). When MIP is applied to a current block, prediction samples for the current block can be derived by i) using surrounding reference samples from which an averaging procedure has been performed, ii) performing a matrix-vector-multiplication procedure, and iii) further performing horizontal / vertical interpolation procedures as needed. The intra-prediction mode used for MIP may be configured differently from the intra-prediction modes used in LIP, PDPC, MRL, ISP intra-prediction, and normal intra-prediction. The intra-prediction mode for MIP may be called MIP intra-prediction mode, MIP prediction mode, or MIP mode. For example, the matrix and offset used in the matrix-vector-multiplication may be set differently depending on the intra-prediction mode for MIP. Here, the matrix may be called the (MIP) weighted matrix, and the offset may be called the (MIP) offset vector or (MIP) bias vector. The specific MIP method will be described later.
[0089] The block reconstruction procedure based on intra prediction and the intra prediction unit 185 within the video encoding device 100 can schematically include, for example, Figures 4 and 5.
[0090] S400 may be performed by the intra-prediction unit 185 of the video encoding device 100, and S410 may be performed by the residual processing unit of the video encoding device 100. Specifically, S410 may be performed by the subtraction unit 115 of the video encoding device 100. In S420, the prediction information is derived by the intra-prediction unit 185 and encoded by the entropy encoding unit 190. In S420, the residual information is derived by the residual processing unit and encoded by the entropy encoding unit 190. The residual information is information about the residual sample. The residual information may include information about the quantized conversion coefficients for the residual sample. As described above, the residual sample is derived into conversion coefficients through the conversion unit 120 of the video encoding device 100, and the conversion coefficients may be derived into quantized conversion coefficients through the quantization unit 130. Information about the quantized conversion coefficients may be encoded by the entropy encoding unit 190 through the residual coding procedure.
[0091] The video encoding device 100 performs intra-prediction for the current block (S400). The video encoding device 100 derives an intra-prediction mode / type for the current block, derives peripheral reference samples for the current block, and generates predicted samples within the current block based on the intra-prediction mode / type and the peripheral reference samples. Here, the intra-prediction mode / type determination, peripheral reference sample derivation, and predicted sample generation procedures may be performed simultaneously, or one procedure may be performed before the others. For example, the intra-prediction unit 185 of the video encoding device 100 may include an intra-prediction mode / type determination unit 186, a reference sample derivation unit 187, and a predicted sample derivation unit 188. The intra-prediction mode / type determination unit 186 can determine the intra-prediction mode / type for the current block, the reference sample derivation unit 187 can derive peripheral reference samples for the current block, and the predicted sample derivation unit 188 can derive predicted samples for the current block. On the other hand, even if not shown, if a prediction sample filtering procedure described later is performed, the intra prediction unit 185 may further include a prediction sample filter unit (not shown). The video encoding device 100 can determine which mode / type from a plurality of intra prediction modes / types is to be applied to the current block. The video encoding device 100 can compare the RD costs for the intra prediction modes / types and determine the optimal intra prediction mode / type for the current block.
[0092] On the other hand, the video encoding device 100 may perform a predictive sample filtering procedure. Predictive sample filtering may be called post-filtering. Some or all of the predictive samples may be filtered by the predictive sample filtering procedure. In some cases, the predictive sample filtering procedure may be omitted.
[0093] The video encoding device 100 generates a residual sample for the current block based on the (filtered) predicted sample (S410). The video encoding device 100 can derive the residual sample by comparing the predicted sample with the original sample of the current block on a phase basis.
[0094] The video encoding device 100 can encode video information including information relating to the intra prediction (prediction information) and residual information relating to the residual sample (S420). The prediction information may include the intra prediction mode information and the intra prediction type information. The video encoding device 100 can output the encoded video information in bitstream form. The output bitstream can be transmitted to the video decoding device 200 via a storage medium or network.
[0095] The residual information may include the residual coding syntax described later. The video encoding device 100 can convert / quantize the residual samples to derive quantized conversion coefficients. The residual information may include information relating to the quantized conversion coefficients.
[0096] On the other hand, as mentioned above, the video encoding device 100 can generate a restored picture (including restored samples and restored blocks). To this end, the video encoding device 100 can perform inverse quantization / inverse transformation again on the quantized conversion coefficients to derive (corrected) residual samples. The reason for performing inverse quantization / inverse transformation again on the residual samples after transformation / quantization is, as mentioned above, to derive the same residual samples as those derived by the video decoding device 200. Based on the predicted samples and the (corrected) residual samples, the video encoding device 100 can generate a restored block containing restored samples for the current block. Based on the restored block, a restored picture for the current picture can be generated. As mentioned above, in-loop filtering procedures and the like may be further applied to the restored picture.
[0097] The intra-prediction unit within the video / image decoding device 200, which is based on intra-prediction, may include, for example, the following:
[0098] The video decoding device 200 can perform operations corresponding to those performed by the video encoding device 100.
[0099] Steps S600 to S620 can be performed by the intra-prediction unit 265 of the video decoding device 200, and the prediction information in S600 and the residual information in S630 can be obtained from the bitstream by the entropy decoding unit 210 of the video decoding device 200. The residual processing unit of the video decoding device 200 can derive a residual sample for the current block based on the residual information. Specifically, the inverse quantization unit 220 of the residual processing unit performs inverse quantization based on the quantized conversion coefficients derived from the residual information to derive conversion coefficients, and the inverse transformation unit 230 of the residual processing unit performs an inverse transformation on the conversion coefficients to derive a residual sample for the current block. Step S640 can be performed by the addition unit 235 or the reconstruction unit of the video decoding device 200.
[0100] Specifically, the video decoder 200 can derive the intra-prediction mode / type for the current block based on the received prediction information (intra-prediction mode / type information) (S600). The video decoder 200 can derive the surrounding reference samples for the current block (S610). The video decoder 200 generates prediction samples within the current block based on the intra-prediction mode / type and the surrounding reference samples (S620). In this case, the video decoder 200 can perform a prediction sample filtering procedure. Prediction sample filtering may be called post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted.
[0101] The video decoding device 200 generates a residual sample for the current block based on the received residual information. The video decoding device 200 generates a restored sample for the current block based on the predicted sample and the residual sample, and can derive a restored block containing the restored sample (S630). A restored picture for the current picture can be generated based on the restored block. As previously mentioned, in-loop filtering procedures and the like may be further applied to the restored picture.
[0102] Here, the intra-prediction unit 265 of the video decoding device 200 may include an intra-prediction mode / type determination unit 266, a reference sample derivation unit 267, and a prediction sample derivation unit 268. The intra-prediction mode / type determination unit 266 determines the intra-prediction mode / type for the current block based on the intra-prediction mode / type information generated and signaled by the intra-prediction mode / type determination unit 186 of the video encoding device 100. The reference sample derivation unit 266 derives the surrounding reference samples of the current block, and the prediction sample derivation unit 267 derives the prediction samples of the current block. On the other hand, even if the prediction sample filtering procedure described above is performed, the intra-prediction unit 265 may further include a prediction sample filter unit (not shown).
[0103] The intra-prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether MPM (most probable mode) or remaining mode is applied to the current block. If MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) indicating one of the intra-prediction mode candidates (MPM candidates). The intra-prediction mode candidates (MPM candidates) may consist of an MPM candidate list or an MPM list. If MPM is not applied to the current block, the intra-prediction mode information may further include remaining mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra-prediction modes excluding the intra-prediction mode candidates (MPM candidates). The video decoding device 200 can determine the intra-prediction mode of the current block based on the intra-prediction mode information. A separate MPM list may be configured for the aforementioned MIP.
[0104] Furthermore, the intra-prediction type information can be embodied in various forms. For example, the intra-prediction type information may include intra-prediction type index information that indicates one of the intra-prediction types. As another example, the intra-prediction type information may include reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if so, which reference sample line is used; ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block; ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartition if the ISP is applied; and at least one of the following flag information: flag information indicating whether PDCP is applied or flag information indicating whether LIP is applied. In addition, the intra-prediction type information may include an MIP flag indicating whether MIP is applied to the current block.
[0105] The intra-prediction mode information and / or the intra-prediction type information may be encoded / decoded through the coding methods described in this document. For example, the intra-prediction mode information and / or the intra-prediction type information may be encoded / decoded through entropy coding (e.g., CABAC, CAVLC) coding based on truncated (rice) binary code.
[0106] Intra template matching prediction(IntraTMP)
[0107] Intra-Template Matching Prediction (IntraTMP) is a special intra-prediction mode in which the optimal prediction block is copied (guided) from the reconstructed portion of the current frame where the L-shaped template matches the current template. For a given search area, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and can use the corresponding block (Matching block in Figure 8) as the prediction block. The encoder can transmit the use of IntraTMP mode, and the same prediction operation can be performed on the decoder side.
[0108] The prediction signal can be generated by matching the L-shaped casual adjacent to the current block with other blocks in a predefined search area, as shown in Figure 8. In Figure 8, R1 represents the current CTU, R2 represents the top-left CTU, R3 represents the above CTU, and R4 represents the left CTU.
[0109] SAD (sum of absolute differences) can be used as the error function. Within each region, the decoder searches for the template with the smallest SAD for the current block, and the block corresponding to the found template can be used as the prediction block.
[0110] The size of all regions (SearchRange_w, SearchRange_h) can be set proportionally to the size of the block (BlkW, BlkH) so that there is a fixed number of SAD comparisons per pixel. That is, the size of all regions can be set as shown in mathematical formula 1.
[0111]
number
[0112] In equation 1, α is a constant that controls the trade-off between gain and complexity. For example, α can be identical to 5.
[0113] To speed up the template matching process, the search range of all search regions can be subsampled by a factor of two. This can reduce the template matching search by a factor of four. After the optimal match is found, a refinement process can be performed. The refinement process can be carried out through a second template matching search around the optimal match with a reduced range. The reduced range can be defined as min(BlkW, BlkH) / 2.
[0114] The intra-template matching tool can be activated for CUs with a width and height of 64 or less. The maximum CU size for intra-template matching can be variable.
[0115] The intra-template matching prediction mode can be signaled at the CU level via a dedicated flag when DIMD (decoder-side intra-mode derivation) is not currently used for the CU. Here, the dedicated flag can indicate whether or not the intra-template matching prediction mode is applied.
[0116] IntraTMP-derived block vector candidates for IBC
[0117] Block vectors (IntraTMP BVs) derived from intra-template matching predictions can be used for IBC. The preserved IntraTMP BVs of surrounding blocks can be used as spatial BV candidates in the IBC candidate list construction along with the IBC BV.
[0118] IntraTMP BVs can be stored in the IBC's block vector buffer. As shown in Figure 9, an IBC block can now use all of the surrounding blocks' IBC BVs and IntraTMP BVs as BV candidates for the IBC BV candidate list.
[0119] IntraTMP BVs may be added to the IBC BV candidate list as spatial candidates.
[0120] [Examples]
[0121] This disclosure proposes a template search region in IntraTMP technology. Furthermore, this disclosure proposes a method for generating or inducing new samples for unavailable (unavailable) sample locations within a reference block and its template region (reference template region). The embodiments proposed through this disclosure may improve the performance of IntraTMP.
[0122] In IntraTMP, as shown in Figure 10, template search regions can be set as specific regions R1, R2, R3, and R4, and a search can be performed for each search region. In this case, the range of each search region may change depending on the width and height of the current block, or the position of the current block. Furthermore, the range of each search region may be determined by an agreement between the encoder and decoder.
[0123] When template search is performed using search regions, constraints may be applied to the range of each search region (i.e., the search range). For example, as shown in Figure 11, when a search region is represented by positions (StartX, StartY) and (EndX, EndY), (StartX, StartY) may be restricted to have a value greater than or equal to the (minimum) template size. For example, the template size may be a predetermined value of 4. (StartX, StartY) indicates the position of the upper-left corner of the search region, StartX indicates the horizontal position of the upper-left corner of the search region, and StartY indicates the vertical position of the upper-left corner of the search region. (EndX, EndY) indicates the lower-right corner of the search region, EndX indicates the horizontal position of the lower-right corner of the search region, and EndY indicates the vertical position of the lower-right corner of the search region. As another example, as shown in Figure 11, when a search region is represented by positions (StartX, StartY) and (EndX, EndY), (EndX, EndY) may be restricted to have a value less than or equal to picture width-block width and picture height-block height, respectively.
[0124] Such restrictions can help reduce complexity in the encoding and decoding processes by preventing excessive template searching. However, conversely, such restrictions can degrade the predictive performance of IntraTMP. For example, as shown in Figure 12, a reference block that is to the left of the search range may be the block most similar to the current block, so such restrictions on the search region (i.e., the search range) can degrade the predictive performance of IntraTMP. As another example, as shown in Figure 13, a reference block that is to the right of the search range may be the block most similar to the current block, so such restrictions on the search region (i.e., the search range) can degrade the predictive performance of IntraTMP.
[0125] The following describes various embodiments that can solve the problem of reduced prediction performance of IntraTMP due to limitations on the search area. The video encoding method described below can be performed by a video encoding device 100, and the video decoding method described below can be performed by a video decoding device 200.
[0126] Figure 14 is a diagram illustrating a video encoding / decoding method according to one embodiment of the present disclosure.
[0127] Referring to Figure 14, the search region associated with the current block can be determined (S1410). There may be multiple search regions, and these can be determined identically in advance by the video encoding device 100 and the video decoding device 200. The size of the search region, i.e., the search range, can also be determined identically in advance by the video encoding device 100 and the video decoding device 200. For example, the search range can be determined based on the size of the current block, as shown in mathematical formula 1.
[0128] An error may be induced between the reference template region and the current template region (S1420). The reference template region is the template region for each reference block located within each search region, and may be the template region corresponding to each reference block or the template region of each reference block. The current template region is the template region of the current block, and may be the template region corresponding to the current block. The error may be induced based on the difference between the current template region and the reference template region. The error may be referred to as "cost," "error," "difference," "template cost," etc. In other words, in this disclosure, "error," "cost," "error," "difference," "template cost," etc. may have the same meaning. The error may be induced based on at least one of SAD (sum of difference), SATD (sum of transformed difference), SSE (sum of squared error), MR-SAD (mean-removed sum of difference), MR-SSE (mean-removed sum of squared error), or MR-SATD (mean-removed sum of transformed difference).
[0129] The reference block of the current block may be determined based on the error (S1430). For example, the reference template region with the smallest error from the current template region (the optimal reference template region) may be determined, and the reference block corresponding to the determined reference template region may be determined as the reference block of the current block. Here, the reference block of the current block may be the reference block used for predicting the current block.
[0130] Figure 15 is a diagram illustrating an example of a video encoding / decoding method for searching for the optimal reference template region.
[0131] Referring to Figure 15, a flag indicating whether IntraTMP is applied (IntraTMP flag) may be encoded in the bitstream, and the IntraTMP flag may be retrieved from the bitstream. If the IntraTMP flag indicates that IntraTMP is applied (S1510), a search for the search region may be performed.
[0132] It can be determined whether the search has been completed for all predetermined search areas (S1520). If there are search areas that have not been searched, a search range may be set for those search areas (S1530). The setting of the search range may change depending on the width and height of the current block, the current block's position, etc. Alternatively, the search range may be predetermined by agreement between the video encoding device 100 and the video decoding device 200.
[0133] An error can be induced between the template adjacent to the current block (current template region) and the template adjacent to the reference block (reference template region) while moving the reference block in subsampling units within a defined search range (S1540). The S1540 process may be a process in which an error is induced between the reference template region corresponding to the reference block at each movement position and the current template region while moving in multiple sample units. During the search process, the bv positions of the reference block may be added to a candidate list in order of increasing error. The size of the candidate list may be defined by an agreement between the video encoding device 100 and the video decoding device 200. For example, the size of the candidate list may be 1, 4, 15, 19, or 30. The error may be induced based on at least one of SAD, SATD, SSE, MR-SAD, MR-SSE, or MR-SATD.
[0134] Once the search for all predetermined search areas is complete, a search for position refinement may be performed for the surrounding positions of each candidate in the candidate list (S1550, S1560). For example, the error between the template adjacent to the current block (current template area) and the template adjacent to the candidate reference block (reference template area) may be induced by moving the reference block one pixel at a time within the search range defined between the video encoding device 100 and the video decoding device 200. Such a refinement search can have a search range of eight adjacent positions centered on the candidate reference block position, as shown in Figure 16(a). In addition, during the search process, the bv positions of the reference blocks may be added to the candidate list in order of increasing error. The error may be induced based on at least one of SAD, SATD, SSE, MR-SAD, MR-SSE, or MR-SATD.
[0135] A search for the surrounding subpixel positions may be performed for candidates in the candidate list (S1570). As shown in Figure 16(b), the subpixel search may be performed in units of 1 / 4, 1 / 2, and 3 / 4 positions, and the search direction may be one of eight adjacent directions centered on the candidate reference block position. Whether or not a subpixel search is performed may be determined by signaling or by an agreement between the video encoding device 100 and the video decoding device 200.
[0136] Example 1
[0137] Example 1 describes a method for changing or removing the restrictions on the search area of IntraTMP.
[0138] There may be no restrictions on the starting position of the IntraTMP search area. For example, in Figure 11, StratX and StratY can each have values less than a predetermined template size. On the other hand, there may also be no restrictions on the ending position of the IntraTMP search area. For example, in Figure 11, EndX can have a value greater than or equal to picture width - block width, and EndY can have a value greater than or equal to picture height - block height.
[0139] As an example for Example 1, the StartX, StartY, EndX, and EndY of the search region may be defined as shown in Table 1.
[0140] [Table 1]
[0141] In Table 1 and Tables 2-4 described below, iTemplateSize indicates a predetermined template size, iCurrY indicates the current block's y-position (vertical position), and iCurrX indicates the current block's x-position (horizontal position). iTemplateSize can be 4. Also, iBlkHeight indicates the current block's height, iBlkWidth indicates the current block's width, and TMP_SEARCH_RANGE_MULT_FACTOR can indicate a value necessary to determine the search range. For example, TMP_SEARCH_RANGE_MULT_FACTOR can be 5. Furthermore, searchRangeHeight is TMP_SEARCH_RANGE_MULT_FACTOR*uiBlkWidth, searchRangeWidth is TMP_SEARCH_RANGE_MULT_FACTOR*uiBlkHeight, PicHeight indicates the picture's height, and PicWidth indicates the picture's width. Additionally, ctuRsAddrY currently indicates the y-position (vertical position) of CTU, ctuRsAddrX currently indicates the x-position (vertical position) of CUT, offsetLCUY represents iCurrY - ctuRsAddrY, and offsetLCUX represents iCurrY - ctuRsAddrX.
[0142] In Table 1, StartX can be std::max(α, iCurrX - searchRangeWidth). Therefore, the upper-left corner position of the search region (StartX) can be determined based on the current block position (iCurrX) and a given search range (searchRangeWidth). Since α can be a real number such that 0 ≤ α ≤ iTemplateSize, StartX, which is the horizontal position of the upper-left corner of the search region, can have a value less than a given template size.
[0143] In Table 1, EndX can be std::min(iCurrX+searchRangeWidth, PicWidth-β). Therefore, the lower right corner of the search region (EndX) can be determined based on the current block's position (iCurrX) and a given search range (searchRangeWidth). Since β can be a real number such that 0≦β≦iBlkWidth, EndX, which is the horizontal position of the lower right corner of the search region, can have a value that exceeds the size difference between the current picture and the current block.
[0144] As another example for Example 1, the StartX, StartY, EndX, and EndY of the search region may be defined as shown in Table 2.
[0145] [Table 2]
[0146] In Table 2, StartX can be std::max(α, iCurrX - searchRangeWidth). Therefore, the upper-left corner position of the search region (StartX) can be determined based on the current block position (iCurrX) and a given search range (searchRangeWidth). Since α can be a real number such that 0 ≤ α ≤ iTemplateSize, StartX, which is the horizontal position of the upper-left corner of the search region, can have a value less than a given template size.
[0147] In Table 2, StartY can be std::max(α, iCurrY-iBlkHeight-offsetLCUY). Therefore, the upper-left corner position of the search region (StartY) can be determined based on the current block position (iCurrY), the current block size (iBlkHeight), and the current CTU position (offsetLCUY=iCurrY-ctuRsAddrY). Since α can be a real number such that 0≦α≦iTemplateSize, StartY, which is the vertical position of the upper-left corner of the search region, can have a value less than a given template size.
[0148] As yet another example from Example 1, the StartX, StartY, EndX, and EndY of the search region may be defined as shown in Table 3.
[0149] [Table 3]
[0150] In Table 3, EndY can be std::min(PicHeight-β, iCurrY-offsetLCUY+getCTUSize-iBlkHeight). Therefore, the lower right corner of the search region (EndY) can be determined based on the current block's position (iCurrY), the current block's size (iBlkHeight), the current picture's height (PicHeight), the current CTU's position (offsetLCUY=iCurrY-ctuRsAddrY), and the current CTU's size (getCTUSize). Since β can be a real number such that 0≦β≦iBlkWidth, EndY, which is the vertical position of the lower right corner of the search region, can have a value that exceeds the size difference between the current picture and the current block.
[0151] In Table 3, StartX can be std::max(α, iCurrX - searchRangeWidth). Therefore, the upper-left corner position of the search region (StartX) can be determined based on the current block position (iCurrX) and a given search range (searchRangeWidth). Since α can be a real number such that 0 ≤ α ≤ iTemplateSize, StartX, which is the horizontal position of the upper-left corner of the search region, can have a value less than a given template size.
[0152] As yet another example from Example 1, the StartX, StartY, EndX, and EndY of the search region may be defined as shown in Table 4.
[0153] [Table 4]
[0154] In Table 4, StartY can be std::max(α, iCurrY - offsetLCUY - iBlkHeight + 1). Therefore, the upper-left corner position of the search region (StartY) can be determined based on the current block position (iCurrY), the current block size (iBlkHeight), and the current CTU position (offsetLCUY = iCurrY - ctuRsAddrY). Since α can be a real number such that 0 ≤ α ≤ iTemplateSize, StartY, which is the vertical position of the upper-left corner of the search region, can have a value less than a given template size.
[0155] In Table 4, StartX can be std::max(α, iCurrX - offsetLCUX - iBlkWidth + 1). Therefore, the upper-left corner position of the search region (StartX) can be determined based on the current block position (iCurrX), the current block size (iBlkWeight), and the current CTU position (offsetLCUX iCurrX - ctuRsAddrX). Since α can be a real number such that 0 ≤ α ≤ iTemplateSize, StartX, which is the horizontal position of the upper-left corner of the search region, can have a value less than a given template size.
[0156] For StartX, StartY, EndX, and EndY defined by the examples in Tables 1 to 4, mvMin, mvMax, mvMax, and mvMin can be derived as shown in Table 5.
[0157] [Table 5]
[0158] In Table 5 and Figure 17, (mvXMin, mvYMin) can be a bv indicating the starting position (upper left corner) of the search region, and (mvXMax, mvYMax) can be a bv indicating the ending position (lower right corner) of the search region.
[0159] Example 2
[0160] When the restrictions on the search area are changed or removed according to Example 1, it is possible that there may be cases where no pre-recovered samples exist in the process of inducing the error between the template area and the reference template area. That is, as shown in Figure 18, there may be samples within the reference template area or reference block that fall outside the picture boundary or CTU boundary, and there may be cases where no pre-recovered sample values exist for such samples. Here, samples for which pre-recovered sample values exist may be referred to as "available samples," and samples for which no pre-recovered sample values exist may be referred to as "not available samples."
[0161] Example 2 describes a method for deriving the value of a non-usable sample from the value of a usable sample.
[0162] As an example for Example 2, if some samples within a reference template region and / or reference block are unavailable samples, available samples can be padded to the locations of the unavailable samples to induce values for the unavailable samples.
[0163] For example, as shown in Figure 19(a), if an unavailable sample is located above the reference template area, the value of the unavailable sample can be derived by vertically padding the nearest available sample. As another example, as shown in Figure 19(b), if an unavailable sample is located to the left of the reference template area, the value of the unavailable sample can be derived by horizontally padding the nearest available sample. As yet another example, as shown in Figure 19(c), if an unavailable sample is located to the left and above the reference template area, the value of the unavailable sample can be derived by vertically and horizontally padding the nearest available sample. The value of an unavailable sample located at the top left can be derived by padding the nearest available sample. As yet another example, as shown in Figure 20(a), if an unavailable sample is located to the right of the reference template area and reference block, the value of the unavailable sample can be derived by horizontally padding the nearest available sample. Here, the available samples used for padding can be located in the reference template area and reference block. As another example, as shown in Figure 20(b), if an unused sample exists below the reference template region and reference block, the value of the unused sample can be induced by vertically padding the nearest available sample. Here, the available sample used for padding can exist in the reference template region and reference block. Although padding has been described above focusing on L-shaped templates, the same padding can be applied to Above-shaped templates or Left-shaped templates.
[0164] As another example relating to Example 2, if some samples within a reference template region and / or reference block are unavailable samples, the values of the unavailable samples can be derived by averaging the values of the available samples.
[0165] For example, as shown in Figure 21(a), if an unavailable sample exists above the reference template region, the value of the unavailable sample can be derived from the mean of the available samples. In Figures 21 to 25, "a" represents the mean of the available samples and may be a real number. "b" may represent the available sample used in calculating the mean. The mean may be the mean of the available samples within the reference template region or the mean of the available samples above the reference template region (Figure 21(b)). Alternatively, the mean may be the mean of the available sample closest to the unavailable sample (Figure 21(c)).
[0166] As another example, as shown in Figure 22(a), if there are unavailable samples to the left of the reference template area, the value of the unavailable sample can be derived from the mean of the available samples. The mean could be the mean of the available samples within the reference template area or the mean of the available samples to the left of the reference template area (Figure 22(b)). Alternatively, the mean could be the mean of the available sample closest to the unavailable sample (Figure 22(c)).
[0167] As yet another example, as shown in Figure 23(a), if there are unavailable samples on the left and above the reference template region, the value of the unavailable sample can be derived from the mean of the available samples. The mean could be the mean of the available samples within the reference template region or the mean of the available sample closest to the unavailable sample (Figure 23(b)).
[0168] As yet another example, as shown in Figure 24(a), if there are unavailable samples to the right of the reference template region and reference block, the value of the unavailable sample can be derived from the mean of the available samples. The mean could be the mean of the available samples within the reference template region and reference block, or the mean of the available samples to the right of the reference template region and reference block (Figure 24(b)). Alternatively, the mean could be the mean of the available sample closest to the unavailable sample (Figure 24(c)).
[0169] As yet another example, as shown in Figure 25(a), if there are unavailable samples below the reference template region and reference block, the value of the unavailable samples can be derived from the mean of the available samples. The mean may be the mean of the available samples within the reference template region and reference block, or the mean of the available samples below the reference template region and reference block (Figure 25(b)). Alternatively, the mean may be the mean of the available samples closest to the unavailable sample (Figure 25(c)).
[0170] The above explanation focused on padding in the context of L-shaped templates, but the same padding can be applied to Above-shaped templates or Left-shaped templates as well.
[0171] Figure 26 is a diagram illustrating an example of a video encoding / decoding method for Example 2.
[0172] Referring to Figure 26, it can be determined whether there is at least one unavailable sample (unavailable sample) within the reference template area (S2610). If an unavailable sample exists within the reference template area, the value of the unavailable sample can be derived based on at least one available sample (available sample) within the reference template area (S2620). For example, the value of the unavailable sample can be derived by padding the available sample or based on the mean value of the available sample. The available sample used to derive the value of the unavailable sample can be any available sample within the reference template area, or an available sample located at a specific position within the reference template area (upper, left, upper left, right, or lower). Alternatively, the available sample used to derive the value of the unavailable sample can be the available sample closest to the unavailable sample.
[0173] Figure 27 is a diagram illustrating another example of the video encoding / decoding method for Example 2.
[0174] Referring to Figure 27, it can be determined whether the sample to the right or below within the reference block is an unavailable sample (S2710). If the sample to the right or below within the reference block is an unavailable sample, the value of the unavailable sample can be derived based on the reference template area and / or the available samples within the reference block (S2720). For example, the value of the unavailable sample can be derived by padding the available samples or based on the mean value of the available samples. When the value of the unavailable sample is derived through padding, the available samples within the reference block can be used for padding. When the value of the unavailable sample is derived through mean induction, the reference template area and the available samples within the reference block can be used to induce the mean value. The available samples used to induce the value of the unavailable sample may be all available samples within the reference template area and / or the reference block, or available samples located at a specific position (or below) within the reference template area and / or the reference block. Alternatively, the available samples used to induce the value of the unavailable sample may be the available sample closest to the unavailable sample.
[0175] Example 2 describes how to induce the value of an unavailable sample for all cases where the unavailable sample is located above the reference template area, to the left of the reference template area, to the right of the reference template area and reference block, and below the reference template area and reference block. However, the method for inducing the value of an unavailable sample may be limited to performing the procedure only when the unavailable sample is located above the reference template area and to the left of the reference template area.
[0176] Figure 28 illustrates a content streaming system to which the embodiments of this disclosure can be applied.
[0177] As shown in Figure 28, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.
[0178] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmits this bitstream to the streaming server. In another example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted.
[0179] The bitstream can be generated by an image encoding method and / or image encoding apparatus to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0180] The streaming server transmits multimedia data to the user's device based on the user's request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server can play a role in controlling the commands and responses between the devices within the content streaming system.
[0181] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0182] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices such as smartwatches, smart glasses, HMDs (head-mounted displays), digital TVs, desktop computers, and digital signage.
[0183] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received from each server can be processed in a distributed manner.
[0184] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that enable the operation of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands etc. are stored and can be executed on a device or computer.
[0185] [Industrial applicability] The embodiments described herein can be used for encoding / decoding images.
[0186] [Claims when filing an international application] [Claim 1] A video decoding method performed by a video decoding device, The current stage is to determine the search area associated with the block; A step of inducing an error between the reference template region, which is the template region of each reference block within the search region, and the current template region, which is the template region of the current block; and A video decoding method comprising the step of determining a reference block for the current block based on the aforementioned error. [Claim 2] The video decoding method according to claim 1, wherein the upper left end position of the search area is determined based on the position of the current block and a predetermined search range. [Claim 3] The video decoding method according to claim 2, wherein the upper left end position of the search area is determined based on the size of the current block and the position of the current coding tree unit (CTU). [Claim 4] The video decoding method according to claim 1, wherein at least one of the horizontal position and vertical position of the upper left end of the search area has a value less than a predetermined template size. [Claim 5] The video decoding method according to claim 1, wherein the lower right end of the search area is determined based on the position of the current block and a predetermined search range. [Claim 6] The video decoding method according to claim 5, wherein the lower right end of the search area is determined based on the current position of the coding tree unit (CTU) and the size of the CTU. [Claim 7] The video decoding method according to claim 1, wherein at least one of the horizontal and vertical positions of the lower right end of the search area exceeds the size difference between the current picture and the current block. [Claim 8] The video decoding method according to claim 1, wherein, based on the fact that at least one sample in the reference template region is unavailable, the value of the unavailable at least one sample is derived based on at least one available sample in the reference template region. [Claim 9] The video decoding method according to claim 8, wherein at least one available sample within the reference template region is the sample closest to at least one unavailable sample. [Claim 10] The video decoding method according to claim 8, wherein, based on the fact that the right-hand or lower-hand sample within the reference block is unavailable, the value of the unavailable right-hand or lower-hand sample is derived based on the available samples within the reference block. [Claim 11] The video decoding method according to claim 8, wherein the value of at least one unavailable sample is derived based on the average value of the available samples within the reference template region. [Claim 12] A video encoding method performed by a video encoding device, The current stage is to determine the search area associated with the block; A step of inducing an error between the reference template region, which is the template region of each reference block within the search region, and the current template region, which is the template region of the current block; and A video encoding method comprising the step of determining a reference block for the current block based on the aforementioned error; [Claim 13] A computer-readable recording medium for storing a bitstream generated by the video encoding method described in claim 12. [Claim 14] A method for transmitting a bitstream generated by a video encoding method, The aforementioned video encoding method is The current stage is to determine the search area associated with the block; A step of inducing an error between the reference template region, which is the template region of each reference block within the search region, and the current template region, which is the template region of the current block; and A method comprising the step of determining a reference block of the current block based on the aforementioned error;
Claims
1. A video decoding method performed by a video decoding device, The current stage is to determine the search area associated with the block; A step of inducing an error between the reference template region, which is the template region of each reference block within the search region, and the current template region, which is the template region of the current block; and A video decoding method comprising the step of determining a reference block for the current block based on the aforementioned error;
2. The video decoding method according to claim 1, wherein the upper left end position of the search area is determined based on the position of the current block and a predetermined search range.
3. The video decoding method according to claim 2, wherein the upper left end position of the search area is determined based on the size of the current block and the position of the current coding tree unit (CTU).
4. The video decoding method according to claim 1, wherein at least one of the horizontal position and vertical position of the upper left end of the search area has a value less than a predetermined template size.
5. The video decoding method according to claim 1, wherein the lower right end of the search area is determined based on the position of the current block and a predetermined search range.
6. The video decoding method according to claim 5, wherein the lower right end of the search area is determined based on the current position of the coding tree unit (CTU) and the size of the CTU.
7. The video decoding method according to claim 1, wherein at least one of the horizontal and vertical positions of the lower right end of the search area exceeds the size difference between the current picture and the current block.
8. The video decoding method according to claim 1, wherein, based on the fact that at least one sample in the reference template region is unavailable, the value of the at least one unavailable sample is derived based on at least one available sample in the reference template region.
9. The video decoding method according to claim 8, wherein at least one available sample within the reference template region is the sample closest to at least one unavailable sample.
10. The video decoding method according to claim 8, wherein, based on the fact that the right-hand sample or lower sample within the reference block is unavailable, the value of the unavailable right-hand sample or lower sample is derived based on the available samples within the reference block.
11. The video decoding method according to claim 8, wherein the value of at least one unavailable sample is derived based on the average value of the available samples within the reference template region.
12. A video encoding method performed by a video encoding device, The current stage is to determine the search area associated with the block; A step of inducing an error between the reference template region, which is the template region of each reference block within the search region, and the current template region, which is the template region of the current block; and A video encoding method comprising the step of determining a reference block for the current block based on the aforementioned error;
13. A computer-readable recording medium for storing a bitstream generated by the video encoding method described in claim 12.
14. A method for transmitting a bitstream generated by a video encoding method, The aforementioned video encoding method is The current stage is to determine the search area associated with the block; A step of inducing an error between the reference template region, which is the template region of each reference block within the search region, and the current template region, which is the template region of the current block; and A method comprising the step of determining a reference block of the current block based on the aforementioned error;