Image encoding / decoding method and device, and recording medium storing bit stream
By combining the weighted sum operation of intra-frame template matching and intra-frame block copy mode, the encoding/decoding efficiency of high-resolution and high-quality images is improved, the problem of insufficient accuracy of intra-frame prediction mode in the existing technology is solved, and the transmission and storage costs are reduced.
Patent Information
- Application Number
- CN202480014276.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-07
- Filing Date
- 2024-04-05
- Publication Date
- 2025-10-03
AI Technical Summary
In the prior art, the encoding/decoding efficiency of high-resolution and high-quality images is low, especially the prediction accuracy of the intra-frame prediction mode is insufficient, resulting in increased encoding/decoding costs.
A method combining intra-frame template matching mode and intra-frame block copy mode is adopted to determine the final prediction block of the current block through weighted sum operation, and the weight information and the prediction mode of the reference block are used to improve the prediction accuracy.
The accuracy of intra-frame prediction is improved, thereby improving the overall encoding/decoding efficiency and reducing transmission and storage costs.
Smart Images

Figure CN120752916A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream. More particularly, the present invention relates to an image encoding / decoding method and apparatus using an intra-frame prediction method, and a recording medium storing a bitstream. Background Art
[0002] Recently, in various application fields, the demand for high-resolution, high-quality images such as ultra-high definition (UHD) images has increased. As the resolution and quality of image data become higher, the amount of data increases relatively compared to existing image data. Therefore, when image data is transmitted using media such as existing wired and wireless broadband lines or stored using existing storage media, the transmission and storage costs increase. In order to solve these problems that arise as the resolution and quality of image data become higher, efficient image encoding / decoding technology is required for images with higher resolution and quality.
[0003] There are various discussions on improving the prediction accuracy of intra-frame prediction modes. In particular, for intra-frame pictures that only use intra-frame prediction, inter-frame prediction, which has a relatively high prediction accuracy, is unavailable, resulting in low encoding and decoding efficiency. Accordingly, by improving the prediction accuracy of intra-frame prediction, the overall encoding / decoding performance can be significantly improved. Specifically, various methods are being discussed to improve intra-frame prediction modes, such as intra-frame block copy mode and intra-frame template matching mode. Summary of the Invention
[0004] Technical issues
[0005] An object of the present invention is to provide a method and apparatus for encoding / decoding an image with improved encoding / decoding efficiency.
[0006] Another object of the present invention is to provide a recording medium for storing a bit stream generated by the method or apparatus for decoding an image provided by the present invention.
[0007] Technical Solution
[0008] A method for decoding an image according to an embodiment of the present invention may include: determining a first prediction block of a current block according to an intra template matching mode, determining a second prediction block of the current block according to an intra block copy mode, and determining a final prediction block of the current block based on a weighted sum of the first prediction block and the second prediction block.
[0009] According to an embodiment, based on weight information obtained from a bitstream, a first weight of the first prediction block and a second weight of the second prediction block for calculating a weighted sum of the first prediction block and the second prediction block may be selected from a plurality of weight candidates.
[0010] According to an embodiment, a first weight of the first prediction block and a second weight of the second prediction block used to calculate a weighted sum of the first prediction block and the second prediction block may be determined according to a prediction mode applied to a predetermined reference block adjacent to the current block.
[0011] According to an embodiment, when a predetermined reference block is predicted using a prediction mode other than the intra block copy mode or the intra template matching mode, the predetermined reference block may be excluded from determination of the first weight and the second weight.
[0012] According to an embodiment, when a predetermined reference block is predicted using a prediction mode other than the intra block copy mode or the intra template matching mode, during determination of the first weight and the second weight, the predetermined reference block may be considered to be predicted by one of the intra block copy mode and the intra template matching mode.
[0013] According to an embodiment, the predetermined reference block may include an upper block at a predetermined upper position of the current block and a left block at a predetermined left position of the current block, and the first weight and the second weight may be determined according to a prediction mode applied to the upper block and a prediction mode applied to the left block.
[0014] According to an embodiment, when both the upper side block and the left side block are predicted according to the intra template matching mode, the first weight may have a value greater than the second weight, when both the upper side block and the left side block are predicted according to the intra block copy mode, the second weight may have a value greater than the first weight, and when one of the upper side block and the left side block is predicted according to the intra template matching mode and the other of the upper side block and the left side block is predicted according to the intra block copy mode, the first weight and the second weight may be set to have the same value.
[0015] According to an embodiment, the predetermined reference block may include one or more upper blocks adjacent to the upper side of the current block and one or more left blocks adjacent to the left side of the current block, and the first weight and the second weight may be determined according to a prediction mode applied to the one or more upper blocks and the one or more left blocks.
[0016] According to an embodiment, the first weight and the second weight may be determined to be proportional to the number of intra template matching mode blocks and the number of intra block copy mode blocks applied to the one or more upper blocks and the one or more left blocks, respectively.
[0017] According to an embodiment, a first weight of the first prediction block and a second weight of the second prediction block for calculating a weighted sum of the first prediction block and the second prediction block may be determined according to a first distortion value of the first prediction block and a second distortion value of the second prediction block.
[0018] According to an embodiment, a first distortion value of the first prediction block may be derived based on a difference between the current block and the first prediction block or a difference between a current template of the current block and a reference template used to derive the first prediction block. Furthermore, a second distortion value of the second prediction block may be derived based on a difference between the current block and the second prediction block or a difference between the current template of the current block and a template of the second prediction block.
[0019] According to an embodiment, the first weight and the second weight may be determined to be proportional to the inverse of the first distortion value and the inverse of the second distortion value, respectively.
[0020] According to an embodiment, when a larger distortion value between the first distortion value and the second distortion value is greater than a predetermined limit value, a third prediction block predicted using a predetermined conventional intra prediction mode may be used to determine a final prediction block, rather than a prediction block corresponding to a larger distortion value between the first prediction block and the second prediction block.
[0021] According to an embodiment, the method for decoding an image may further include determining a third prediction block of the current block according to a conventional intra prediction mode, and may determine a final prediction block of the current block based on a weighted sum of the first prediction block, the second prediction block, and the third prediction block.
[0022] A method for encoding an image according to an embodiment of the present invention may include: determining a first prediction block of a current block according to an intra template matching mode, determining a second prediction block of the current block according to an intra block copy mode, and determining a final prediction block of the current block based on a weighted sum of the first prediction block and the second prediction block.
[0023] A non-transitory computer-readable recording medium according to an embodiment of the present invention may store a bitstream generated by a method for encoding an image.
[0024] The transmission method according to the embodiment of the present invention can transmit a bit stream generated by the method for encoding an image.
[0025] The features briefly summarized above regarding the present invention are provided merely as examples to explain the detailed description and should not be construed as limiting the scope of the present invention.
[0026] Beneficial effects
[0027] The present invention proposes various embodiments of a method for improving prediction accuracy of intra prediction by combining a prediction block derived based on intra template matching and a prediction block derived based on intra block copying.
[0028] In addition, the present invention proposes various embodiments of a method for improving prediction accuracy of intra prediction by combining a prediction block generated from an intra prediction mode, a prediction block derived based on intra template matching, and a prediction block derived based on intra block copying.
[0029] According to various embodiments, since the prediction accuracy of intra prediction is improved, the overall encoding and decoding efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a block diagram showing the configuration of an encoding device according to an embodiment of the present invention
[0031] Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment of the present invention.
[0032] Figure 3 FIG. 1 is a diagram schematically illustrating a video coding and decoding system to which the present invention is applicable.
[0033] Figure 4 A method for deriving a first prediction block based on intra template matching and a method for deriving a second prediction block based on intra block copying are shown.
[0034] Figure 5 Shows the current block and the reference blocks referred to by the current block.
[0035] Figure 6 is a flowchart illustrating an embodiment of an intra prediction method according to the present invention.
[0036] Figure 7 A content streaming system to which an embodiment according to the present invention is applicable is exemplarily shown. DETAILED DESCRIPTION
[0037] Best Mode for Carrying Out the Invention
[0038] A method for decoding an image according to an embodiment of the present invention may include: determining a first prediction block of a current block according to an intra template matching mode, determining a second prediction block of the current block according to an intra block copy mode, and determining a final prediction block of the current block based on a weighted sum of the first prediction block and the second prediction block.
[0039] Embodiments for carrying out the invention
[0040] The present invention may have various modifications and embodiments, and specific embodiments are shown in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but should be understood to include all modifications, equivalents or alternatives contained within the spirit and technical scope of the present invention. The same figure marks in the drawings indicate the same or similar functions in various aspects. For a clearer description, the shapes and sizes of the elements in the drawings can be provided by way of example. The detailed description of the exemplary embodiments described below refers to the drawings, which illustrate specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that the various embodiments are different from each other, but not necessarily mutually exclusive. For example, without departing from the spirit and scope of the present invention with reference to one embodiment, the specific shapes, structures and features described herein can be implemented in other embodiments. It should also be understood that without departing from the spirit and scope of the embodiment, the position or arrangement of each component within each disclosed embodiment can be changed. Accordingly, the detailed description set forth below is not intended to be restrictive, and the scope of the exemplary embodiments is limited only by the full range of equivalent forms given by the appended claims and these claims (if appropriately described).
[0041] In the present invention, the terms first, second, etc. may be used to describe various components, but the components should not be limited by the terms. The terms are used only to distinguish one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component without departing from the scope of the present invention. The terms are and / or include a combination of multiple related descriptive items or any item in multiple related descriptive items.
[0042] The components shown in the embodiments of the present invention are depicted independently to indicate different characteristic functions, and do not represent that each component is formed as a separate hardware or software configuration unit. That is, for ease of explanation, each component is listed and included as a separate component, and at least two of the components can be combined to form a single component, or a component can be divided into multiple components to perform functions. As long as it does not depart from the essence of the present invention, embodiments in which the components are integrated and embodiments in which each component is divided are also included in the scope of the present invention.
[0043] The terms used in the present invention are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In addition, some components of the present invention are not essential components for performing necessary functions in the present invention and can be optional components that are only used to improve performance. The present invention can be realized by only including essential components for realizing the gist of the present invention and not including components that are only used to improve performance, and the structure that only includes essential components and does not include optional components that are only used to improve performance is also included in the scope of the present invention.
[0044] In an embodiment, the term "at least one" may refer to one of a number greater than or equal to 1, such as 1, 2, 3, and 4. In an embodiment, the term "plurality" may refer to one of a number greater than or equal to 2, such as 2, 3, and 4.
[0045] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. When describing the embodiments of this specification, if it is determined that a detailed description of related known configurations or functions will obscure the subject matter of this specification, the detailed description will be omitted, the same reference numerals will be used for the same components in the drawings, and repeated description of the same components will be omitted.
[0046] Description of terms
[0047] Hereinafter, "image" may refer to a picture constituting a video, or may refer to the video itself. For example, "encoding and / or decoding of an image" may refer to "encoding and / or decoding of a video," or may refer to "encoding and / or decoding of one of the pictures constituting the video."
[0048] Hereinafter, "moving image" and "video" may be used interchangeably with each other in the same meaning. Furthermore, a target image may be an encoding target image that is the target of encoding and / or a decoding target image that is the target of decoding. Furthermore, a target image may be an input image input to an encoding device or an input image input to a decoding device. Here, the target image may have the same meaning as the current image.
[0049] Hereinafter, “image,” “picture,” “frame,” and “picture” may be used with the same meaning and may be used interchangeably.
[0050] Hereinafter, a "target block" may be an encoding target block that is a target of encoding and / or a decoding target block that is a target of decoding. In addition, a target block may be a current block that is a target of current encoding and / or decoding. For example, "target block" and "current block" may be used with the same meaning and may be used interchangeably.
[0051] In the following, "block" and "unit" can be used with the same meaning and can be used interchangeably. In addition, "unit" can mean a luma component block and its corresponding chroma component block to distinguish it from a block. For example, a coding tree unit (CTU) can consist of a luma component (Y) coding tree block (CTB) and its associated two chroma component (Cb, Cr) coding tree blocks.
[0052] Hereinafter, “sample”, “picture element”, and “pixel” may be used with the same meaning and may be used interchangeably. Herein, a sample may represent a basic unit constituting a block.
[0053] Hereinafter, “inter-frame” and “inter-screen” may be used with the same meaning and may be used interchangeably.
[0054] Hereinafter, “intra-frame” and “intra-screen” may be used with the same meaning and may be used interchangeably.
[0055] Figure 1 is a block diagram showing the configuration of an encoding device according to an embodiment of the present invention.
[0056] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images. The encoding device 100 may encode one or more images sequentially.
[0057] refer to Figure 1 The encoding device 100 may include: an image partitioning unit 110, an intra-frame prediction unit 120, a motion prediction unit 121, a motion compensation unit 122, a switch 115, a subtractor 113, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 117, a filtering unit 180 and a reference picture buffer 190.
[0058] In addition, the encoding device 100 can generate a bit stream including information encoded by encoding the input image and output the generated bit stream. The generated bit stream can be stored in a computer-readable recording medium or can be streamed through a wired / wireless transmission medium.
[0059] The image partitioning unit 110 can partition the input image into various forms to improve the efficiency of video encoding / decoding. That is, the input video consists of multiple pictures, and for compression efficiency, parallel processing, etc., a picture can be partitioned and processed hierarchically. For example, a picture can be partitioned into one or more tiles or slices, and then partitioned again into multiple codec tree units (CTUs). Alternatively, a picture can first be partitioned into multiple sub-pictures defined as rectangular slice groups, and each sub-picture can be partitioned into tiles / slices. Here, sub-pictures can be used to support the function of partially encoding / decoding and transmitting pictures independently. Since multiple sub-pictures can be reconstructed separately, this has the advantage of easy editing in applications where multi-channel input is configured as one picture. In addition, tiles can be divided horizontally to generate bricks. Here, bricks can be used as basic units for parallel processing within a picture. In addition, a CTU can be recursively partitioned into a quad tree (QT), and the terminal node of the partition can be defined as a coding unit (CU). The CU can be partitioned into a PU (Prediction Unit) as a prediction unit and a TU (Transform Unit) as a transform unit to perform prediction and partitioning. On the other hand, the CU can be used as a prediction unit and / or a transform unit itself. Here, for flexible partitioning, each CTU can be recursively partitioned into multi-type trees (MTT) and quadtrees (QT). Partitioning the CTU into multi-type trees can start from the terminal node of the QT, and the MTT can be composed of a binary tree (BT) and a triple tree (TT). For example, the MTT structure can be classified into a vertical binary split mode (SPLIT_BT_VER), a horizontal binary split mode (SPLIT_BT_HOR), a vertical ternary split mode (SPLIT_TT_VER), and a horizontal ternary split mode (SPLIT_TT_HOR). In addition, during partitioning, the minimum block size (MinQTSize) of the quadtree of the luma block can be set to 16 × 16, the maximum block size (MaxBtSize) of the binary tree can be set to 128 × 128, and the maximum block size (MaxTtSize) of the ternary tree can be set to 64 × 64. In addition, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the ternary tree can be specified as 4 × 4, and the maximum depth (MaxMttDepth) of the multi-type tree can be specified as 4. In addition, in order to improve the coding efficiency of the I slice, a dual tree that uses the CTU partition structure of the luma component and the chroma component differently can be applied.On the other hand, in P slices and B slices, the luma and chroma coding tree blocks (CTBs) within a CTU can be partitioned into a single tree that shares the coding tree structure.
[0060] The encoding device 100 can perform encoding on the input image in intra mode and / or inter mode. Alternatively, the encoding device 100 can perform encoding on the input image in a third mode other than intra mode and inter mode (for example, IBC mode, palette mode, etc.). However, if the third mode has functional characteristics similar to those of intra mode or inter mode, it can be classified as intra mode or inter mode for ease of explanation. In the present invention, the third mode is classified and described separately only when a specific description of the third mode is required.
[0061] When the intra mode is used as the prediction mode, the switch 115 can switch to intra, and when the inter mode is used as the prediction mode, the switch 115 can switch to inter. Here, the intra mode can represent the intra prediction mode, and the inter mode can represent the inter prediction mode. The encoding device 100 can generate a prediction block for the input block of the input image. In addition, the encoding device 100 can encode the residual block using the residual of the input block and the prediction block after generating the prediction block. The input image can be referred to as the current image that is the current encoding target. The input block can be referred to as the current block that is the current encoding target or the encoding target block.
[0062] When the prediction mode is intra mode, the intra prediction unit 120 can use samples of blocks that have been encoded / decoded around the current block as reference samples. The intra prediction unit 120 can perform spatial prediction of the current block by using the reference samples, or generate prediction samples of the input block through spatial prediction. In this article, intra prediction can refer to intra-frame prediction.
[0063] As the intra prediction method, non-directional prediction modes such as DC mode and planar mode and directional prediction modes (for example, 65 directions) can be applied. Here, the intra prediction method can be expressed as an intra prediction mode or an intra-screen prediction mode.
[0064] When the prediction mode is inter mode, the motion prediction unit 121 can retrieve the area that best matches the input block from the reference image during the motion prediction process and derive the motion vector by using the retrieved area. In this case, a search area can be used as the area. The reference image can be stored in the reference picture buffer 190. Here, when encoding / decoding the reference image, it can be stored in the reference picture buffer 190.
[0065] The motion compensation unit 122 may generate a prediction block of the current block by performing motion compensation using a motion vector. Herein, inter prediction may refer to inter-picture prediction or motion compensation.
[0066] When the value of the motion vector is not an integer, the motion prediction unit 121 and the motion compensation unit 122 may generate a prediction block by applying an interpolation filter to a partial area of the reference picture. In order to perform inter-frame prediction or motion compensation, it may be determined based on the codec unit whether the motion prediction and motion compensation mode of the prediction unit included in the codec unit is one of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and intra block copy (IBC) mode, and inter-frame prediction or motion compensation may be performed according to each mode.
[0067] In addition, based on the above-mentioned inter-frame prediction method, the affine mode based on sub-PU prediction, the subblock-based temporal motion vector prediction (SbTMVP) mode, the PU-based prediction and MVD merge (Merge with MVD, MMVD) mode and the geometric partitioning mode (Geometric Partitioning Mode, GPM) mode can be applied. In addition, in order to improve the performance of each mode, history based MVP (HMVP), pairwise average MVP (PAMVP), combined intra / inter prediction (CIIP), adaptive motion vector resolution (AMVR), bidirectional optical flow (BDOF), bidirectional prediction with CU weights (BCW), local illumination compensation (LIC), template matching (TM), overlapped block motion compensation (OBMC), etc. can be applied.
[0068] The subtractor 113 may generate a residual block by utilizing the difference between the input block and the prediction block. The residual block may be referred to as a residual signal. The residual signal may represent the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming or quantizing, or transforming and quantizing, the difference between the original signal and the prediction signal. The residual block may be a residual signal of a block unit.
[0069] The transform unit 130 may generate a transform coefficient by performing a transform on the residual block and output the generated transform coefficient. In this article, the transform coefficient may be a coefficient value generated by performing a transform on the residual block. When the transform skip mode is applied, the transform unit 130 may skip the transform of the residual block.
[0070] The quantization level may be generated by applying quantization to the transform coefficients or the residual signal. Hereinafter, the quantization level may also be referred to as the transform coefficient in the embodiments.
[0071] For example, the 4×4 luma residual block generated by intra prediction is transformed using a discrete sine transform (DST)-based basis vector, and the remaining residual block can be transformed using a discrete cosine transform (DCT)-based basis vector. In addition, the transform block is partitioned into a quadtree shape of one block using the residual quadtree (RQT) technology, and after transforming and quantizing each transform block partitioned by the RQT, a coded block flag (cbf) can be transmitted when all coefficients become 0 to improve coding efficiency.
[0072] As another alternative, a Multiple Transform Selection (MTS) technique that selectively uses multiple transform bases to perform transforms can be applied. That is, instead of partitioning the CU into TUs through RQT, a function similar to TU partitioning can be performed through sub-block transform (SBT) technology. Specifically, SBT is only applied to inter-frame prediction blocks, and unlike RQT, the current block can be partitioned into 1 / 2 or 1 / 4 sizes in the vertical or horizontal direction, and then the transform can be performed on only one of the blocks. For example, if the current block is partitioned vertically, the transform can be performed on the leftmost or rightmost block, and if the current block is partitioned horizontally, the transform can be performed on the topmost or bottommost block.
[0073] In addition, a low-frequency non-separable transform (LFNST) can be applied. This is a secondary transform technique that additionally transforms the transformed residual signal into the frequency domain through DCT or DST. LFNST additionally transforms the 4×4 or 8×8 low-frequency region on the upper left side so that the residual coefficients can be concentrated on the upper left side.
[0074] The quantization unit 140 may generate a quantization level by quantizing the transform coefficient or the residual signal according to a quantization parameter (QP) and output the generated quantization level. Herein, the quantization unit 140 may quantize the transform coefficient by using a quantization matrix.
[0075] For example, a quantizer with a QP value of 0 to 51 may be used. Alternatively, if the image size is large and high coding efficiency is required, a QP of 0 to 63 may be used. In addition, a dependent quantization (DQ) method using two quantizers instead of one quantizer may be applied. DQ performs quantization using two quantizers (e.g., Q0 and Q1), but even without signaling information about the use of a specific quantizer, a quantizer for the next transform coefficient may be selected based on the current state through a state transition model.
[0076] The entropy coding unit 150 may generate a bitstream by performing entropy coding on the values calculated by the quantization unit 140 or the codec parameter values calculated when performing coding according to the probability distribution, and output the bitstream. The entropy coding unit 150 may perform entropy coding on information about samples of an image and information for decoding the image. For example, the information for decoding the image may include syntax elements.
[0077] When entropy coding is applied, symbols are represented so that a smaller number of bits are assigned to symbols with a high probability of occurrence and a larger number of bits are assigned to symbols with a low probability of occurrence, thereby reducing the size of the bitstream for the symbol to be encoded. The entropy coding unit 150 can perform entropy coding using coding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. For example, the entropy coding unit 150 can perform entropy coding by utilizing a variable length coding / code (VLC) table. In addition, the entropy coding unit 150 can derive a binarization method of the target symbol and a probability model of the target symbol / binary, and perform arithmetic coding and decoding by utilizing the derived binarization method and context model.
[0078] In this regard, when CABAC is applied, in order to reduce the size of the probability table stored in the decoding device, the table probability update method can be changed to a table update method using a simple equation and applied. In addition, two different probability models can be used to obtain more accurate symbol probability values.
[0079] In order to encode the transformation coefficient level (quantization level), the entropy encoding unit 150 may change the coefficients in a two-dimensional block form into a one-dimensional vector form through a transformation coefficient scanning method.
[0080] The codec parameters may include information (flags, indexes, etc.) encoded in the encoding device 100 and signaled to the decoding device 200, such as syntax elements, and information derived during the encoding or decoding process, and may represent information required when encoding or decoding an image.
[0081] Herein, signaling a flag or an index may mean that the corresponding flag or index is entropy-encoded in an encoder and included in a bitstream, and may mean that the corresponding flag or index is entropy-decoded from a bitstream in a decoder.
[0082] The encoded current image can be used as a reference image for another image to be processed later. Therefore, the encoding device 100 can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference picture buffer 190.
[0083] The quantization level may be dequantized in the dequantization unit 160 or inversely transformed in the inverse transform unit 170. The dequantized and / or inversely transformed coefficients may be added to the prediction block by the adder 117. Herein, the dequantized and / or inversely transformed coefficients may represent coefficients on which at least one of dequantization and inverse transformation has been performed, and may represent a reconstructed residual block. The dequantization unit 160 and the inverse transform unit 170 may perform the inverse process of the quantization unit 140 and the transform unit 130.
[0084] The reconstructed block may pass through the filtering unit 180. The filtering unit 180 may utilize all or some filtering techniques to apply a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a bilateral filter (BIF), a luma mapping with chroma scaling (LMCS) filter, etc. to reconstruct the sample, reconstruct the block, or reconstruct the image. The filtering unit 180 may be referred to as an in-loop filter. In this case, the name in-loop filter may also be used without LMCS.
[0085] A deblocking filter can remove block distortion generated at the boundaries between blocks. To determine whether to apply a deblocking filter, a determination can be made based on samples included in a number of rows or columns included in the block whether to apply the deblocking filter to the current block. When applying a deblocking filter to a block, different filters can be applied depending on the desired deblocking filter strength.
[0086] To compensate for coding errors using sample adaptive offset, an appropriate offset value can be added to the sample value. Sample adaptive offset can correct the offset from the original image in units of samples for the deblocked image. A method can be used in which the samples included in the image are partitioned into a predetermined number of regions, the regions to which the offset is applied are determined, and the offset is applied to the determined regions, or a method can be used in which the offset is applied taking into account edge information about each sample.
[0087] The bilateral filter (BIF) can also correct the deviation from the original image on a sample-by-sample basis for the image on which deblocking has been performed.
[0088] The adaptive loop filter can perform filtering based on the comparison result of the reconstructed image and the original image. The samples included in the image can be partitioned into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information on whether to apply the ALF can be signaled by the codec unit (CU), and the form and coefficients of the adaptive loop filter to be applied to each block can be varied.
[0089] In luma mapping with chroma scaling (LMCS), luma mapping (LM) refers to remapping luma values using a piecewise linear model, and chroma scaling (CS) refers to a technique for scaling the residual values of chroma components according to the average luma value of the prediction signal. In particular, LMCS can be used as an HDR correction technique to reflect the characteristics of High Dynamic Range (HDR) images.
[0090] The reconstructed block or reconstructed image that has passed through the filtering unit 180 may be stored in the reference picture buffer 190. The reconstructed block that has passed through the filtering unit 180 may be part of a reference image. That is, the reference image is a reconstructed image composed of the reconstructed block that has passed through the filtering unit 180. The stored reference image may be used later for inter-frame prediction or motion compensation.
[0091] Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment of the present invention.
[0092] The decoding device 200 may be a decoder, a video decoding device, or an image decoding device.
[0093] refer to Figure 2 The decoding device 200 may include: an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 201, a switch 203, a filtering unit 260 and a reference picture buffer 270.
[0094] The decoding device 200 can receive the bitstream output from the encoding device 100. The decoding device 200 can receive the bitstream stored in a computer-readable recording medium, or can receive the bitstream streamed via a wired / wireless transmission medium. The decoding device 200 can decode the bitstream in intra-frame mode or inter-frame mode. In addition, the decoding device 200 can generate a reconstructed image or a decoded image generated by decoding, and output the reconstructed image or the decoded image.
[0095] When the prediction mode for decoding is intra mode, the switch 203 may switch to intra. Alternatively, when the prediction mode for decoding is inter mode, the switch 203 may switch to inter.
[0096] The decoding device 200 can obtain a reconstructed residual block by decoding the input bitstream and generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block that becomes the decoding target by adding the reconstructed residual block and the prediction block. The decoding target block can be referred to as the current block.
[0097] The entropy decoding unit 210 may generate symbols by entropy decoding the bitstream according to the probability distribution. The generated symbols may include symbols in the form of quantization levels. In this article, the entropy decoding method may be the inverse process of the above-mentioned entropy encoding method.
[0098] The entropy decoding unit 210 may change the coefficients of the one-dimensional vector shape into coefficients of the two-dimensional block shape through a transform coefficient scanning method to decode the transform coefficient level (quantization level).
[0099] The quantization level may be dequantized in the dequantization unit 220 or inversely transformed in the inverse transform unit 230. The quantization level may be the result of dequantization and / or inverse transformation and may be generated as a reconstructed residual block. In this article, the dequantization unit 220 may apply a quantization matrix to the quantization level. The dequantization unit 220 and the inverse transform unit 230 applied to the decoding device may apply the same technology as the dequantization unit 160 and the inverse transform unit 170 applied to the aforementioned encoding device.
[0100] When intra mode is used, the intra prediction unit 240 can generate a prediction block by performing spatial prediction on the current block using sample values of blocks already decoded around the target block. The intra prediction unit 240 applied to the decoding device can apply the same technology as the intra prediction unit 120 applied to the aforementioned encoding device.
[0101] When using inter-frame mode, the motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block using a motion vector and a reference image stored in the reference picture buffer 270. When the value of the motion vector is not an integer value, the motion compensation unit 250 can generate a prediction block by applying an interpolation filter to a partial area within the reference image. In order to perform motion compensation, it can be determined based on the codec unit whether the motion compensation mode of the prediction unit included in the corresponding codec unit is skip mode, merge mode, AMVP mode, or current picture reference mode, and motion compensation can be performed according to each mode. The motion compensation unit 250 applied to the decoding device can apply the same technology as the motion compensation unit 122 applied to the above-mentioned encoding device.
[0102] The adder 201 can generate a reconstructed block by adding the reconstructed residual block and the prediction block. The filtering unit 260 can apply at least one of an inverse LMCS, a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or the reconstructed image. The filtering unit 260 applied to the decoding device can apply the same filtering technology as the filtering unit 180 applied to the aforementioned encoding device.
[0103] The filtering unit 260 may output a reconstructed image. The reconstructed block or image may be stored in the reference picture buffer 270 and used for inter-frame prediction. The reconstructed block that has passed through the filtering unit 260 may be part of a reference image. In other words, the reference image may be a reconstructed image composed of the reconstructed blocks that have passed through the filtering unit 260. The stored reference image may be later used for inter-frame prediction or motion compensation.
[0104] Figure 3 FIG. 1 is a diagram schematically illustrating a video coding and decoding system to which the present invention is applicable.
[0105] The video coding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in the form of a file or streaming via a digital storage medium or a network.
[0106] The encoding device 10 according to the embodiment may include a video source generating unit 11, an encoding unit 12, and a transmitting unit 13. The decoding device 20 according to the embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitting unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, and the display unit may be configured as a separate device or an external component.
[0107] The video source generation unit 11 can obtain a video / image by capturing, synthesizing, or generating a video / image. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, a smartphone, etc., and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc., in which case the video / image capture process may be replaced by a process for generating relevant data.
[0108] The encoding unit 12 can encode the input video / image. For compression and encoding efficiency, the encoding unit 12 can perform a series of processes such as prediction, transformation and quantization. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bit stream. The detailed configuration of the encoding unit 12 can also be the same as above. Figure 1 The encoding device 100 is configured in the same manner.
[0109] The transmitting unit 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit 21 of the decoding device 20 via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit 13 may include components for generating a media file in a predetermined file format and may also include components for transmission via a broadcast / communication network. The receiving unit 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.
[0110] The decoding unit 22 can decode the video / image by performing a series of processes such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12. The detailed configuration of the decoding unit 22 can also be compared with Figure 2 The above-mentioned decoding device 200 is configured in the same manner.
[0111] The rendering unit 23 may render the decoded video / image. The rendered video / image may be displayed by the display unit.
[0112] This invention describes a method for improving the accuracy of intra-frame prediction by utilizing intra-frame template matching (IntraTMP) and intra-frame block copy (IBC). IntraTMP and IBC are techniques for searching for a prediction block for the current block within an already reconstructed region. According to the proposed method, the accuracy of intra-frame prediction can be improved by generating a final prediction block from a combination of a prediction block derived using IntraTMP and a prediction block derived using IBC.
[0113] Hereinafter, a method of predicting a current block by combining a first prediction block derived from IntraTMP and a second prediction block derived from IBC is described.
[0114] Figure 4 A method for deriving a first prediction block based on intra template matching and a method for deriving a second prediction block based on intra block copying are shown.
[0115] Intra-template matching is a method for deriving a prediction block based on a reference template 412 corresponding to a current template 402 of a current block 400. Figure 4 , a current template 402 having a predetermined form is derived from adjacent reference samples of the current block 400. In addition, a reference template 412 that is most similar to the current template 402 is derived from the reconstructed region of the current picture 404. Next, a prediction block of the current block 400 is determined based on a matching block 410 corresponding to the derived reference template 412.
[0116] exist Figure 4 , the current template 402 has a “┏” shape including both the left reference sample and the upper reference sample of the current block 400. However, according to an embodiment, the current template 402 may be configured to include only the left reference sample of the current block 400 or only the upper reference sample of the current block 400. L1 and L2 of the current template 402 refer to the width of the samples included in the current template 402, and each of L1 and L2 is a positive integer. That is, the current template 402 may include not only the reference samples immediately adjacent to the current block 400, but also the reference samples spaced apart from the current block 400 by a predetermined sample unit.
[0117] Accordingly, when intra template matching is applied to the current block 400, information indicating the application of intra template matching is encoded in the encoder. In addition, the decoder can receive the information and perform intra template matching on the current block in the same way as the encoder, so that the decoder and the encoder can generate the same predicted block.
[0118] The intra block copy is a method of determining a block vector 422 within a predefined search range of a reconstructed area of the current block 400 and deriving a prediction block of the current block 400 according to the block vector 422 .
[0119] In the encoder, a matching block 420 that is most similar to the current block 400 is found within a predefined search range (R1, R2, R3, and R4) of the reconstructed region. In addition, a block vector 422 is derived that represents the displacement between the current block 400 and the matching block 420. In addition, the block vector 422 can be encoded based on block vector information of neighboring blocks of the current block 400.
[0120] The decoder can derive the block vector 422 of the current block 400 from the block vector information of the current block 400 by referring to the block vectors of the neighboring blocks of the current block 400. In addition, the matching block 420 of the current block 400 is determined based on the block vector 422, and the prediction block of the current block 400 is derived from the matching block 420.
[0121] The predefined search range for intra template matching and intra block copy may include the current codec tree block R1, the upper left codec tree block R2, the upper codec tree block R3, and the left codec tree block R4. However, this is merely an example, and the search range may be determined to have any size other than the above search range.
[0122] The prediction block derived through the intra template matching described above is defined as the first prediction block, and the prediction block derived through intra block copying is defined as the second prediction block. To improve the accuracy of intra prediction, the final prediction block of the current block can be generated by a weighted sum of the first prediction block and the second prediction block derived by the two methods. Equation 1 below shows a method for generating the final prediction block by a weighted sum of the first prediction block derived through intra template matching and the second prediction block derived through intra block copying.
[0123] [Equation 1]
[0124] P=W IntraTMP x P IntraTMP +W IBC x P IBC
[0125] In Equation 1, P IntraTMP and P IBC W respectively mean the prediction value of the first prediction block based on intra template matching and the prediction value of the second prediction block based on intra block copying. IntraTMP and W IBC They respectively mean the weight of the first prediction block based on intra template matching and the weight of the second prediction block based on intra block copying. The weights satisfy W IntraTMP +W IBC =1,W IntraTMP ≥0, and W IBC ≥0.
[0126] According to an embodiment, a final prediction block may be generated by weighting the sum of K first blocks derived based on intra-frame template matching and one second prediction block derived based on intra-frame block copying. Herein, the K prediction blocks derived based on intra-frame template matching may be determined in ascending order of template distortion. Herein, K is an arbitrary positive integer.
[0127] According to an embodiment, K first prediction blocks and one second prediction block may have the same weight. Alternatively, the sum of the weights of the K first prediction blocks may be set to be equal to the weight of one second prediction block. In addition, the same weight may be set for each of the K first prediction blocks. Alternatively, different weights may be set for each of the K first prediction blocks according to the distortion of the template. For example, a smaller weight may be given to the first prediction block generated from a template with a larger distortion value. The distortion of the template may be calculated by utilizing various related measurement methods, such as the sum of absolute differences (SAD) or the sum of square errors (SSE).
[0128] According to the embodiment, the weight W of the first prediction block of the intra template matching is IntraTMP and the weight W of the second prediction block according to intra block copy IBC It can be determined as any weight preset at the encoder and decoder.
[0129] Alternatively, the weight W of the first prediction block may be determined among N predefined weight candidates. IntraTMP and the weight W of the second prediction block IBC In this article, the encoder may transmit index information indicating a weight candidate applied to the current block among N weight candidates, and the decoder may determine the weight applied to the current block by parsing the index information. In this article, N is an arbitrary positive integer. In this article, the index information may be configured to indicate the weight W of the first prediction block. IntraTMP Or the weight W of the second prediction block IBC .
[0130] Alternatively, the weight W of the first prediction block can be derived from the neighboring reference blocks of the current block. IntraTMP and the weight W of the second prediction block IBC Hereinafter, a method for determining weights of a first prediction block and a second prediction block from a reference block of a current block will be described.
[0131] Figure 5 Shows the current block and the reference blocks referred to by the current block.
[0132] exist Figure 5 , the sizes of the neighboring reference blocks A1 to A8, L1 to L8, and AL of the current block 500 are the minimum sizes of blocks storing intra prediction mode information. Figure 5Described are the reference blocks A1 to A4 on the upper side of the current block 500 and the reference blocks L1 to L4 on the left side of the current block 500, and a plurality of reference blocks other than four may exist on the upper side and the left side of the current block 500 according to the size of the current block 500. For example, when the block has a minimum size of 4×4 and the current block 500 has a size of 32×16, eight reference blocks may exist on the upper side of the current block 500 and four reference blocks may exist on the left side of the current block 500.
[0133] Table 1 shows a method of determining a weight by using reference blocks A4 and L4 of the current block 500 .
[0134] [Table 1]
[0135]
[0136] As shown in Table 1, when both block A4 and block L4 have the intra prediction mode of intra template matching, W IntraTMP Can be set to be greater than W IBC On the other hand, when both block A4 and block L4 have intra block copy as the type of intra prediction mode, W IBC Can be set to be greater than W IntraTMP When block A4 and block L4 have intra template matching and intra block copy, W IntraTMP and W IBC The weights set in Table 1 are just an example, and any other values may be assigned to W IntraTMP and W IBC The weights in Table 1 are determined with reference to two reference blocks A4 and L4, but the weights may be determined with reference to a reference block at another location. According to an embodiment, W may be determined based on M adjacent reference blocks of the current block 500. IntraTMP and W IBC In this paper, M is an arbitrary positive integer. IntraTMP and W IBC It can be calculated as shown in Equation 2 below.
[0137] [Equation 2]
[0138]
[0139] In Equation 2, N IntraTMP and N IBC=W means the number of reference blocks to which intra template matching is applied and the number of reference blocks to which intra block copy is applied, respectively, among the M neighboring reference blocks of the current block 500. In this article, when the prediction mode of the reference block is neither intra template matching mode nor intra block copy mode, the reference block can be excluded from the determination of the weight. Alternatively, when the prediction mode of the reference block is neither intra template matching mode nor intra block copy mode, the prediction mode of the reference block can be considered as one of the intra template matching mode and the intra block copy mode to calculate W. IntraTMP and W IBC . When the current block 500 is larger than the minimum size of the adjacent blocks, the prediction information of the adjacent blocks can be considered to be applicable to multiple reference blocks corresponding to the adjacent blocks. For example, when the adjacent blocks include all reference blocks A1 to A4 and the adjacent blocks are predicted based on intra template matching, the prediction mode of the reference blocks A1 to A4 can be considered to be the intra template matching mode. When the prediction mode of the reference block is a combination mode of intra template matching and intra block copy, the prediction mode of the reference block can be considered to be one of the intra template matching mode and the intra block copy mode. In addition, when the prediction mode of the reference block is a combination mode of conventional intra prediction, intra template matching and intra block copy, the prediction mode of the reference block can be considered to be one of the conventional intra prediction mode, intra template matching mode and intra block copy mode used.
[0140] According to an embodiment, the weight W of the first prediction block may be determined according to the distortion of the first prediction block and the distortion of the second prediction block. IntraTMP and the weight W of the second prediction block IBC The distortion value of the first prediction block may be determined based on the difference between the current block and the first prediction block. In addition, the distortion value of the second prediction block may be determined based on the difference between the current block and the second prediction block.
[0141] Alternatively, the distortion value of the first prediction block may be determined based on the difference between the current template and the reference template. Figure 4 As shown, the size and shape of the template used to calculate the distortion value of the first prediction block based on intra-frame template matching can be arbitrarily determined. Figure 4 The sizes L1 and L2 of the templates shown in are arbitrary positive integers, and the template may include only left reference samples, only upper reference samples, or both left and upper reference samples. Similarly, the distortion value of the second prediction block may be determined based on the difference between the current template and the adjacent reference template of the matching block 420.
[0142] Herein, Equation 3 shows a method for calculating the weight W of the first prediction block based on the distortion of the first prediction block and the distortion of the second prediction block. IntraTMP and the weight W of the second prediction block IBC method.
[0143] [Equation 3]
[0144]
[0145] In Equation 3, D IntraTMP and D IBC Represent the distortion values of the first prediction block based on intra template matching and the second prediction block based on intra block copying, respectively. The smaller the distortion value, the greater the similarity between the predicted signal and the original signal. Conversely, the larger the distortion value, the less similar the predicted signal is to the original signal. Accordingly, the weight of the prediction block can be determined in proportion to the inverse of the distortion value.
[0146] According to Equation 3, the weight of each prediction block can be calculated using the distortion value between different signals. Specifically, the distortion value D of the second prediction block can be used. IBC To calculate the weight W of the first prediction block according to intra template matching IntraTMP In addition, the distortion value D of the first prediction block can be used IntraTMP To calculate the weight W of the second prediction block according to intra block copy IBC The distortion of each signal can be calculated by utilizing various related measurement methods, such as the sum of absolute differences (SAD) or the sum of square errors (SSE).
[0147] For the calculation of Equation 3, in order to reduce the implementation complexity and computational complexity, all floating-point operations can be replaced with integer operations by using a look-up table (LUT). In this way, the weight W can be derived using only integer multiplication, addition, and shift operations instead of floating-point operations. IntraTMP and W IBC For example, the weights can be derived using the LUT of a cross-component linear model (CCLM).
[0148] According to an embodiment, when the distortion value of the first prediction block according to intra template matching and the distortion value of the second prediction block according to intra block copy are similar to each other, it may not make sense to generate the final prediction block by using the weighted sum of the two prediction blocks. Accordingly, when the distortion values are determined to be similar as described above, the intra prediction mode combining intra template matching and intra block copy may not be applied to the current block. Alternatively, one of the intra template matching and intra block copy may be applied to the current block. In this article, the distortion of each signal can be calculated by using various related measurement methods, such as the sum of absolute differences (SAD) or the sum of squared errors (SSE). The similarity between the distortion values can be determined using Equation 4 below.
[0149] [Equation 4]
[0150] D max -D min ≤Threshold1
[0151] Here, D min means the smaller distortion value between the distortion value of the first prediction block according to intra template matching and the distortion value of the second prediction block according to intra block copying. max This means the larger distortion value between the distortion value of the first prediction block based on intra template matching and the distortion value of the second prediction block based on intra block copying. Furthermore, Threshold 1 (Threshold1) is an arbitrary positive real number. When Equation 4 is satisfied, that is, when the difference in distortion value between the two prediction blocks is equal to or less than an arbitrary threshold, the two prediction blocks can be determined to be similar. Equation 4 is merely an example of determining similarity, and the method for determining similarity is not limited to Equation 4.
[0152] According to the embodiment, when the distortion value D of the first prediction block according to the intra template matching IntraTMP and the distortion value D of the second prediction block according to intra block copy IBC When there is a large difference between the two distortion values, as shown in Equation 5, the final prediction block can be generated according to the weighted sum of the prediction block having the smaller distortion value between the two distortion values and the prediction block generated by the planar mode as the intra prediction mode. In this article, the distortion of each signal can be calculated by using various related measurement methods, such as the sum of absolute differences (SAD) or the sum of squared errors (SSE). Alternatively, even when the distortion value D of the first prediction block is IntraTMP and the distortion value D of the second prediction block according to intra block copy IBC When the larger distortion value between the two is greater than a predetermined threshold, the prediction block generated by the planar mode may be used to replace the prediction block with the larger distortion value. In addition, although Equation 5 uses the planar mode of intra prediction, any intra prediction mode other than the planar mode may also be used.
[0153] [Equation 5]
[0154] P=W min x P min +W Planar x P Planar
[0155] In Equation 5, P min represents a prediction signal having a smaller distortion value between the distortion value of the prediction signal derived based on intra template matching and the distortion value of the prediction signal derived based on intra block copying. Planar W means the prediction signal generated by the planar pattern. min and W planar Represent the weight of the prediction signal with smaller distortion value and the weight of the prediction signal generated by the plane mode respectively. In addition, the two weights satisfy W min +W planar =1,W min ≥0, and W planar ≥0.W min and W planar These can be arbitrary weights set in advance at the encoder and decoder.
[0156] Alternatively, W min and W planar It can be determined from N preset weight candidates. In this article, the encoder can transmit index information indicating the weight candidate applied to the current block among the N weight candidates, and the decoder can determine the weight applied to the current block by parsing the index information. In this article, N is an arbitrary positive integer. In this article, the index information can be configured to indicate the weight W of the prediction signal with a smaller distortion value. min Or according to the weight W of the prediction signal of the plane pattern planar .
[0157] Additionally, the difference between the distortion values may be determined using Equation 6 below.
[0158] [Equation 6]
[0159] D max -D min ≥Threshold2
[0160] Here, D min means the smaller distortion value between the distortion value of the first prediction block according to intra template matching and the distortion value of the second prediction block according to intra block copying. maxThis means the larger distortion value between the distortion value of the first prediction block based on intra template matching and the distortion value of the second prediction block based on intra block copying. Furthermore, Threshold 2 is an arbitrary positive real number. When the condition according to Equation 6 is satisfied, that is, when the difference in distortion value between the two prediction signals is equal to or greater than an arbitrary threshold, it can be determined that omitting the prediction signal with the larger distortion value is advantageous. Equation 6 is merely an example of determining whether the difference in distortion value is large, and the method for determining the difference in distortion value is not limited to Equation 6.
[0161] Hereinafter, a method of predicting a current block by combining a first prediction block derived by intra template matching, a second prediction block derived by intra block copying, and a third prediction block generated from a conventional intra prediction mode is described. The conventional intra prediction mode is an intra prediction mode that uses neighboring reference samples of the current block and may include a planar mode, a DC mode, and a directional intra prediction mode.
[0162] As described above, the first prediction block is generated by the matching block 410 according to intra template matching, and the second prediction block is generated by the matching block 420 according to intra block copying. According to an embodiment, a third prediction block according to conventional intra prediction may also be generated based on reference samples adjacent to the current block 400. In addition, a final prediction block of the current block 400 may be determined based on a weighted sum of the first prediction block, the second prediction block, and the third prediction block. Equation 7 below shows a method for determining the final prediction block of the current block based on a weighted sum of the first prediction block according to intra template matching, the second prediction block according to intra block copying, and the third prediction block according to conventional intra prediction.
[0163] [Equation 7]
[0164] P=W IntraMode x P IntraMode +W IntraTMP x P IntraTMP +W IBC x P IBC
[0165] In Equation 7, P IntraMode 、P IntraTMP and P IBC W respectively mean the prediction signal according to conventional intra prediction, the prediction signal according to intra template matching, and the prediction signal according to intra block copying. IntraMode 、W IntraTMP and W IBC are the weights of the prediction signal based on conventional intra prediction, the weights of the prediction signal based on intra template matching, and the weights of the prediction signal based on intra block copying. IntraMode +W IntraTMP +W IBC=1,W IntraMode ≥0,W IntraTMP ≥0, and W IBC ≥0.
[0166] According to an embodiment, a final prediction block may be generated by weighting the sum of K first blocks derived based on intra template matching, a second prediction block derived based on intra block copying, and a third prediction block generated from a conventional intra prediction mode. Here, the K prediction blocks derived based on intra template matching may be determined in ascending order of template distortion. Here, K is an arbitrary positive integer.
[0167] According to an embodiment, K first prediction blocks, one second prediction block and one third prediction block may have the same weight. Alternatively, the sum of the weights of the K first prediction blocks may be set to be equal to the weight of one second prediction block or the weight of one third prediction block. In addition, the same weight may be set for each of the K first prediction blocks. Alternatively, different weights may be set for each of the K first prediction blocks according to the distortion of the template. For example, a small weight may be given to the first prediction block generated from a template with a large distortion value. The distortion of the template may be calculated by utilizing various related measurement methods, such as the sum of absolute differences (SAD) or the sum of square errors (SSE).
[0168] According to the embodiment, the weight W of the first prediction block of the intra template matching is IntraTMP , the weight W of the second prediction block according to intra block copy IBC and the weight W of the third prediction block according to the conventional intra prediction mode IntraMode It can be determined as any weight preset at the encoder and decoder.
[0169] Alternatively, the weight W of the first prediction block may be determined among N predefined weight candidates. IntraTMP , the weight W of the second prediction block IBC and the weight W of the third prediction block IntraMode In this article, the encoder may transmit index information indicating a weight candidate applied to the current block among N weight candidates, and the decoder may determine the weight applied to the current block by parsing the index information. In this article, N is an arbitrary positive integer. In this article, the index information may include the weight W of the first prediction block. IntraTMP , the weight W of the second prediction block IBC and the weight W of the third prediction block IntraMode .
[0170] Alternatively, the weight W of the first prediction block can be derived from the neighboring reference blocks of the current block. IntraTMP, the weight W of the second prediction block IBC and the weight W of the third prediction block IntraMode Hereinafter, a method for determining weights of a first prediction block, a second prediction block, and a third prediction block from a reference block of a current block will be described.
[0171] Table 2 shows that by using Figure 5 The weights are determined by using the adjacent reference blocks A4 and L4 of the current block.
[0172] [Table 2]
[0173]
[0174]
[0175] As shown in Table 2, a larger weight may be assigned to the prediction block derived according to the intra prediction mode of block A4 and block L4. The weights in Table 2 are merely examples, and different arbitrary weights may be assigned to each prediction block. According to an embodiment, W may be determined based on the M neighboring reference blocks of the current block 500. IntraTMP 、W IBC and W IntraMode In this paper, M is an arbitrary positive integer. IntraTMP 、W IBC and W IntraMode It can be calculated as shown in Equation 8 below.
[0176] [Equation 8]
[0177]
[0178] In Equation 8, N IntraMode 、N IntraTMP and N IBC It means the number of reference blocks to which conventional intra prediction is applied, the number of reference blocks to which intra template matching is applied, and the number of reference blocks to which intra block copy is applied, among the M neighboring reference blocks of the current block 500. Herein, when the prediction mode of the reference block is not the conventional intra prediction mode, the intra template matching mode, or the intra block copy mode, the reference block may be excluded. Alternatively, when the prediction mode of the reference block is not the conventional intra prediction mode, the intra template matching mode, or the intra block copy mode, the prediction mode of the reference block may be considered to be one of the conventional intra prediction mode, the intra template matching mode, and the intra block copy mode to calculate N IntraMode 、N IntraTMP and N IBC. When the current block 500 is larger than the minimum size of the adjacent blocks, the prediction information of the adjacent blocks can be considered to be applicable to multiple reference blocks corresponding to the adjacent blocks. For example, when the adjacent blocks include all reference blocks A1 to A4 and the adjacent blocks are predicted based on intra template matching, the prediction mode of the reference blocks A1 to A4 can be considered to be the intra template matching mode. When the prediction mode of the reference block is a combination mode of intra template matching and intra block copy, the prediction mode of the reference block can be considered to be one of the intra template matching mode and the intra block copy mode. In addition, when the prediction mode of the reference block is a combination mode of conventional intra prediction, intra template matching and intra block copy, the prediction mode of the reference block can be considered to be one of the conventional intra prediction mode, intra template matching mode and intra block copy mode used.
[0179] According to an embodiment, W may be determined based on the distortion of the first prediction block, the distortion of the second prediction block, and the distortion of the third prediction block. IntraTMP 、W IBC and W IntraMode . The distortion value of the first prediction block can be determined based on the difference between the current block and the first prediction block. In addition, the distortion value of the second prediction block can be determined based on the difference between the current block and the second prediction block. In addition, the distortion value of the third prediction block can be determined based on the difference between the current block and the third prediction block. Alternatively, as described above, the distortion value of the first prediction block can be determined based on the difference between the current template and the reference template. In addition, the distortion value of the second prediction block can be determined based on the difference between the current template and the template of the second prediction block. W IntraTMP 、W IBC and W IntraMode Each can be determined to be proportional to the inverse of the distortion value. Accordingly, a larger weight can be given to the prediction block with a smaller distortion value.
[0180] According to an embodiment, when the distortion value of the first prediction block according to intra template matching and the distortion value of the second prediction block according to intra block copy are similar to each other, it may not make sense to generate the final prediction block by using the weighted sum of the two prediction blocks. Accordingly, when the distortion values are determined to be similar as described above, the intra prediction mode that combines conventional intra prediction, intra template matching and intra block copy may not be applied to the current block. Alternatively, one of conventional intra prediction, intra template matching and intra block copy may be applied to the current block. In this article, the distortion of each signal can be calculated by using various related measurement methods, such as the sum of absolute differences (SAD) or the sum of squared errors (SSE). The similarity between the distortion values can be determined using Equation 9 below.
[0181] [Equation 9]
[0182] D max -D min≤Threshold3
[0183] Here, D min means the smaller distortion value between the distortion value of the first prediction block according to intra template matching and the distortion value of the second prediction block according to intra block copying. max This means the larger distortion value between the distortion value of the first prediction block based on intra template matching and the distortion value of the second prediction block based on intra block copying. Furthermore, Threshold 3 is an arbitrary positive real number. When Equation 9 is satisfied, that is, when the difference in distortion value between the two prediction blocks is equal to or less than an arbitrary threshold, the two prediction blocks can be determined to be similar. Equation 9 is merely an example of determining similarity, and the method for determining similarity is not limited to Equation 9.
[0184] According to the embodiment, when the distortion value D of the first prediction block according to the intra template matching IntraTMP and the distortion value D of the second prediction block according to intra block copy IBC When there is a large difference between the two distortion values, the final prediction block can be generated according to the weighted sum of the prediction block generated by the conventional intra prediction, the prediction block with the smaller distortion value between the two distortion values, and the prediction block generated by the planar mode as the intra prediction mode, as shown in Equation 10. In this article, the distortion of each signal can be calculated by using various related measurement methods, such as the sum of absolute differences (SAD) or the sum of squared errors (SSE). Alternatively, even when the distortion value D of the first prediction block is IntraTMP and the distortion value D of the second prediction block according to intra block copy IBC When the larger distortion value between the two is greater than a predetermined threshold, the prediction block generated by the planar mode may be used to replace the prediction block with the larger distortion value. In addition, although Equation 10 uses the planar mode of intra prediction, any intra prediction mode other than the planar mode may also be used.
[0185] [Equation 10]
[0186] P=W IntraMode x P IntraMode +W min x P min +W Planar x P Planar
[0187] In Equation 10, P IntraMode represents the prediction signal generated based on the conventional intra-frame prediction mode. min represents a prediction signal having a smaller distortion value between the distortion value of the prediction signal derived based on intra template matching and the distortion value of the prediction signal derived based on intra block copying. PlanarW means the prediction signal generated by the planar pattern. IntraMode 、W min and W planar They represent the weight of the prediction signal generated from the intra prediction mode, the weight of the prediction signal with a smaller distortion value, and the weight of the prediction signal generated by the planar mode. In addition, the three weights satisfy W IntraMode +W min +W planar =1,W IntraMode ≥0,W min ≥0, and W planar ≥0.W IntraMode 、W min and W planar These can be arbitrary weights set in advance at the encoder and decoder.
[0188] Alternatively, W IntraMode 、W min and W planar It can be determined from N preset weight candidates. In this article, the encoder can transmit index information indicating the weight candidate applied to the current block among the N weight candidates, and the decoder can determine the weight applied to the current block by parsing the index information. In this article, N is an arbitrary positive integer. In this article, the index information can be configured to indicate the weight W of the prediction signal according to the conventional intra prediction mode. IntraMode , the weight W of the prediction signal with smaller distortion value min Or according to the weight W of the prediction signal of the plane pattern planar .
[0189] In the above method, the difference in weight values may be determined according to Equation 11.
[0190] [Equation 11]
[0191] D max -D min ≥Threshold4
[0192] Here, D min means the smaller distortion value between the distortion value of the first prediction block according to intra template matching and the distortion value of the second prediction block according to intra block copying. maxThis means the larger distortion value between the distortion value of the first prediction block based on intra template matching and the distortion value of the second prediction block based on intra block copying. Furthermore, Threshold 4 is an arbitrary positive real number. When the condition according to Equation 11 is satisfied, that is, when the difference in distortion value between the two prediction signals is equal to or greater than an arbitrary threshold, it can be determined that omitting the prediction signal with the larger distortion value is advantageous. Equation 11 is merely an example of determining whether the difference in distortion value is large, and the method for determining the difference in distortion value is not limited to Equation 11.
[0193] Figure 6 A flow chart showing an embodiment of an intra prediction method according to the present invention.
[0194] In step S602 , a first prediction block of the current block is determined according to the intra template matching mode.
[0195] In step S604 , a second prediction block of the current block is determined according to the intra block copy mode.
[0196] In step S606, a final prediction block of the current block is determined based on the weighted sum of the first prediction block and the second prediction block. The following embodiments may be applied to determine the final prediction block of step S606.
[0197] According to the embodiment, a first weight of the first prediction block and a second weight of the second prediction block for calculating a weighted sum of the first prediction block and the second prediction block may be determined. As an example, the first weight and the second weight may be selected from a plurality of weight candidates based on weight information obtained from a bitstream. Weight information indicating the first weight and the second weight among the plurality of weight candidates may be encoded at the encoder, and one of the plurality of weight candidates may be selected at the decoder based on the weight information. When calculating a weighted sum of three or more prediction blocks, the weight information may include not only the first weight and the second weight but also information about additional weights.
[0198] Alternatively, according to an embodiment, the first weight and the second weight may be determined based on a prediction mode applied to a predetermined reference block adjacent to the current block. For example, the predetermined reference block may include an upper block at a predetermined upper position of the current block and a left block at a predetermined left position of the current block, and the first weight and the second weight may be determined based on the prediction mode applied to the upper block and the prediction mode applied to the left block. In this article, when both the upper block and the left block are predicted according to the intra-frame template matching mode, the first weight may have a value greater than the second weight, when both the upper block and the left block are predicted according to the intra-frame block copy mode, the second weight may have a value greater than the first weight, and when one of the upper block and the left block is predicted according to the intra-frame template matching mode and the other of the upper block and the left block is predicted according to the intra-frame block copy mode, the first weight and the second weight may be set to have the same value.
[0199] According to an embodiment, the predetermined reference blocks may include one or more upper blocks adjacent to the upper side of the current block and one or more left blocks adjacent to the left side of the current block, and the first weight and the second weight may be determined according to the prediction mode applied to the one or more upper blocks and the one or more left blocks. According to an embodiment, the first weight and the second weight may be determined to be proportional to the number of intra template matching mode blocks and the number of intra block copy mode blocks applied to the one or more upper blocks and the one or more left blocks, respectively.
[0200] According to an embodiment, when a predetermined reference block is predicted using a prediction mode other than the intra block copy mode or the intra template matching mode, the predetermined reference block may be excluded from the determination of the first weight and the second weight. Alternatively, when a predetermined reference block is predicted using a prediction mode other than the intra block copy mode or the intra template matching mode, during the determination of the first weight and the second weight, the predetermined reference block may be considered to be predicted by one of the intra block copy mode and the intra template matching mode.
[0201] According to an embodiment, a first weight of the first prediction block and a second weight of the second prediction block used to calculate a weighted sum of the first prediction block and the second prediction block can be determined based on a first distortion value of the first prediction block and a second distortion value of the second prediction block. In addition, the first distortion value of the first prediction block can be derived based on a difference between the current block and the first prediction block or a difference between a current template of the current block and a reference template used to derive the first prediction block, and the second distortion value of the second prediction block can be derived based on a difference between the current block and the second prediction block or a difference between the current template of the current block and a template of the second prediction block. In addition, the first weight and the second weight can be determined to be proportional to the inverse of the first distortion value and the inverse of the second distortion value, respectively.
[0202] According to an embodiment, when the difference between the first distortion value and the second distortion value is greater than a predetermined limit value, a third prediction block predicted using a predetermined regular intra prediction mode may be used to determine the final prediction block, rather than a prediction block corresponding to a larger distortion value between the first prediction block and the second prediction block. The predetermined regular intra prediction mode may be one of a planar mode, a DC mode, and a predetermined directional mode.
[0203] According to an embodiment, the method for decoding an image may further include determining a third prediction block of the current block according to a conventional intra prediction mode, and may determine a final prediction block of the current block based on a weighted sum of the first prediction block, the second prediction block, and the third prediction block.
[0204] According to the prediction method performed in steps S602 to S606, the current block can be encoded or decoded. In addition, the bitstream generated by the encoder according to the prediction method performed in steps S602 to S606 can be stored in a recording medium or transmitted to the outside of the encoder.
[0205] Figure 7 A content streaming system to which an embodiment according to the present invention is applicable is exemplarily shown.
[0206] like Figure 7 As shown, the content streaming system to which the embodiment of the present invention is applied may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0207] The encoding server compresses the content received from the multimedia input device such as a smartphone, a camera, a CCTV, etc. into digital data to generate a bitstream and transmits it to the streaming server. As another example, if the multimedia input device such as a smartphone, a camera, a CCTV, etc. directly generates the bitstream, the encoding server can be omitted.
[0208] A bitstream may be generated by the image encoding method and / or the image encoding apparatus to which the embodiments of the present invention are applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0209] The streaming server transmits multimedia data to a user device based on a user request via a network server. The network server can also function as an intermediary, notifying the user of any available services. When a user requests a desired service from the network server, the network server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which can control commands and responses between devices within the content streaming system.
[0210] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a period of time.
[0211] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, HMDs), digital televisions, desktop computers, digital signage, etc.
[0212] Each server in the above-described content streaming system may operate as a distributed server, in which case data received from each server may be distributed and processed.
[0213] The above embodiments may be performed in the same or corresponding manner in the encoding device and the decoding device. In addition, an image may be encoded / decoded using at least one or a combination of at least one of the above embodiments.
[0214] The order of applying the above-described embodiments may be different in the encoding device and the decoding device. Alternatively, the order of applying the above-described embodiments may be the same in the encoding device and the decoding device.
[0215] The above-described embodiments may be performed on each of the luminance signal and the chrominance signal. Alternatively, the above-described embodiments may be performed identically for the luminance signal and the chrominance signal.
[0216] In the above embodiments, the method is described based on a flow chart having a series of steps or units, but the present invention is not limited to the order of the steps. On the contrary, some steps can be performed simultaneously with other steps or in a different order. In addition, it should be understood by those skilled in the art that the steps in the flow chart are not mutually exclusive, and other steps can be added to the flow chart or some steps can be deleted from the flow chart without affecting the scope of the present invention.
[0217] The embodiments may be implemented in the form of program instructions that are executable by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may include independent program instructions, data files, data structures, etc., or a combination thereof. The program instructions recorded in the computer-readable recording medium may be specially designed and constructed for the present invention, or may be well known to those skilled in the art in the field of computer software technology.
[0218] The bit stream generated by the encoding method according to the above embodiment can be stored in a non-volatile computer-readable recording medium. In addition, the bit stream stored in the non-volatile computer-readable recording medium can be decoded by the decoding method according to the above embodiment.
[0219] Examples of computer-readable recording media include: magnetic recording media such as hard disks, floppy disks, and magnetic tapes; optical data storage media such as CD-ROMs or DVD-ROMs; magneto-optical media such as floppy disks; and hardware devices such as read-only memory (ROM), random access memory (RAM), flash memory, etc., which are specifically configured to store and execute program instructions. Examples of program instructions include not only machine language codes formatted by a compiler, but also high-level language codes that can be implemented by a computer using an interpreter. A hardware device can be configured to be operated by one or more software modules or vice versa to perform the process according to the present invention.
[0220] Although the present invention has been described in terms of specific items such as detailed elements and limited embodiments and drawings, these are provided only to help a more comprehensive understanding of the present invention, and the present invention is not limited to the above embodiments. It will be understood by those skilled in the art that various modifications and changes can be made based on the above description.
[0221] Therefore, the spirit of the present invention should not be limited to the above-described embodiments, and the full scope of the appended claims and their equivalents should fall within the scope and spirit of the present invention.
[0222] Industrial Applicability
[0223] The present invention can be used in an apparatus for encoding / decoding an image and a recording medium for storing a bit stream.
Claims
1. A method for decoding an image, the method comprising: Determine a first prediction block of the current block according to an intra-frame template matching mode; determining a second prediction block for the current block according to an intra block copy mode; as well as A final prediction block of the current block is determined based on a weighted sum of the first prediction block and the second prediction block.
2. The method according to claim 1, wherein A first weight of the first prediction block and a second weight of the second prediction block for calculating a weighted sum of the first prediction block and the second prediction block are selected among a plurality of weight candidates based on weight information obtained from a bitstream.
3. The method according to claim 1, wherein A first weight of the first prediction block and a second weight of the second prediction block used to calculate a weighted sum of the first prediction block and the second prediction block are determined according to a prediction mode applied to a predetermined reference block adjacent to the current block.
4. The method according to claim 3, wherein: When the predetermined reference block is predicted using a prediction mode other than the intra block copy mode or the intra template matching mode, the predetermined reference block is excluded from determination of the first weight and the second weight.
5. The method according to claim 3, wherein When a predetermined reference block is predicted using a prediction mode other than the intra block copy mode or the intra template matching mode, during determination of the first weight and the second weight, the predetermined reference block is considered to be predicted by one of the intra block copy mode and the intra template matching mode.
6. The method according to claim 3, wherein: The predetermined reference blocks include an upper block at a predetermined upper position of the current block and a left block at a predetermined left position of the current block, and The first weight and the second weight are determined according to a prediction mode applied to the upper block and a prediction mode applied to the left block.
7. The method according to claim 6, wherein: When both the upper block and the left block are predicted according to the intra template matching mode, the first weight has a larger value than the second weight, wherein, when both the upper block and the left block are predicted according to the intra block copy mode, the second weight has a larger value than the first weight, and Here, when one of the upper block and the left block is predicted according to the intra template matching mode and the other is predicted according to the intra block copy mode, the first weight and the second weight are set to have the same value.
8. The method according to claim 3, wherein: The predetermined reference blocks include one or more upper side blocks adjacent to the upper side of the current block and one or more left side blocks adjacent to the left side of the current block, and The first weight and the second weight are determined according to a prediction mode applied to one or more upper blocks and one or more left blocks.
9. The method according to claim 8, wherein The first weight and the second weight are determined to be proportional to the number of intra template matching mode blocks and the number of intra block copy mode blocks applied to the one or more upper blocks and the one or more left blocks, respectively.
10. The method according to claim 1, wherein A first weight of the first prediction block and a second weight of the second prediction block for calculating a weighted sum of the first prediction block and the second prediction block are determined according to a first distortion value of the first prediction block and a second distortion value of the second prediction block.
11. The method according to claim 10, wherein: deriving a first distortion value of the first prediction block based on a difference between the current block and the first prediction block or a difference between a current template of the current block and a reference template used to derive the first prediction block, and The second distortion value of the second prediction block is derived based on a difference between the current block and the second prediction block or a difference between a current template of the current block and a template of the second prediction block.
12. The method according to claim 10, wherein: The first weight and the second weight are determined to be proportional to the inverse of the first distortion value and the inverse of the second distortion value, respectively.
13. The method according to claim 10, wherein: When the difference between the first distortion value and the second distortion value is greater than a predetermined limit value, a third prediction block predicted using a predetermined conventional intra prediction mode is used to determine a final prediction block, rather than a prediction block corresponding to a larger distortion value between the first prediction block and the second prediction block.
14. The method according to claim 1, further comprising: Determine a third prediction block for the current block according to a conventional intra prediction mode, The final prediction block of the current block is determined based on the weighted sum of the first prediction block, the second prediction block and the third prediction block.
15. A method for encoding an image, the method comprising: Determine a first prediction block of the current block according to an intra-frame template matching mode; determining a second prediction block for the current block according to an intra block copy mode; as well as A final prediction block of the current block is determined based on a weighted sum of the first prediction block and the second prediction block.
16. A computer-readable recording medium for storing a bit stream generated by a method for encoding an image, in, Methods used to encode images include: Determine a first prediction block of the current block according to an intra-frame template matching mode; determining a second prediction block for the current block according to the intra block copy mode; and A final prediction block of the current block is determined based on a weighted sum of the first prediction block and the second prediction block.
17. A method for transmitting a bitstream, the bitstream being generated by a method for encoding an image, the method comprising: encoding the image based on a method for encoding the image; as well as Transmits a bitstream consisting of an encoded image, Among them, the method for encoding the image includes: Determine a first prediction block of the current block according to an intra-frame template matching mode; determining a second prediction block for the current block according to the intra block copy mode; and A final prediction block of the current block is determined based on a weighted sum of the first prediction block and the second prediction block.