Intra prediction mode derivation-based video coding method and apparatus

The image encoding/decoding method determines a prediction method for current blocks based on neighboring blocks and other factors, addressing the inefficiencies of existing video compression technologies and improving image quality.

WO2025116566A1PCT designated stage expired Publication Date: 2025-06-05HYUNDAI MOTOR CO LTD +2

Patent Information

Application Number
PCT/KR2024/019146
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-27
Filing Date
2024-11-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing video compression technologies, such as H.264/AVC, HEVC, and VVC, face challenges in efficiently encoding and decoding high-resolution, high-frame-rate video data, leading to increased data amounts and decreased image quality.

Method used

An image encoding/decoding method and device that determines a prediction method for a current block based on neighboring blocks, current block information, matching costs of restored regions, or signaled information, thereby generating a prediction block and improving image quality.

Benefits of technology

The proposed method efficiently encodes and decodes images, enhancing both objective and subjective image quality of restored images by adaptively combining template regions, feature extraction methods, and intra prediction mode derivation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019146_05062025_PF_FP_ABST
    Figure KR2024019146_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an intra prediction mode derivation-based video coding method and apparatus. According to the disclosed method, a template region for predicting the current block, a template feature, and a method for deriving a directional prediction mode are adaptively determined.
Need to check novelty before this filing date? Find Prior Art

Description

Video coding method and device based on intra prediction mode derivation

[0001] The present disclosure relates to a video coding method and device based on intra prediction mode derivation.

[0002] The content described below merely provides background information related to the present embodiment and does not constitute prior art.

[0003] Since video data has a large amount of data compared to voice data or still image data, it requires a lot of hardware resources, including memory, to store or transmit it without processing for compression.

[0004] Therefore, when storing or transmitting video data, the encoder compresses the video data and stores or transmits it, and the decoder receives the compressed video data, decompresses it, and plays it back. These video compression technologies include H.264 / AVC, HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), which improves encoding efficiency by about 30% compared to HEVC.

[0005] However, as the size, resolution, and frame rate of images are gradually increasing, and the amount of data that needs to be encoded is also increasing, a new compression technology that has better encoding efficiency and better image quality improvement than existing compression technologies is required.

[0006] The present disclosure provides an image encoding / decoding method and device for efficiently encoding / decoding an image and improving the objective / subjective quality of a restored image, and a recording medium for storing a bitstream generated by the image encoding method / device.

[0007] One aspect of the present disclosure provides an image decoding method, comprising: a step of determining a prediction method of a current block based on at least one of prediction methods of neighboring blocks, information of a current block, a matching cost of a restored area, or signaled information, performed by an image decoding device; and a step of generating a prediction block of the current block using the prediction method of the current block.

[0008] One aspect of the present disclosure provides an image encoding method, comprising: a step of determining a prediction method to be applied to a current block among a plurality of prediction methods, performed by an image encoding device; the plurality of prediction methods including at least one of a first prediction method in which the prediction method of the current block is derived based on prediction methods of neighboring blocks, a second prediction method derived based on information of the current block, or a third prediction method derived based on a matching cost of a previously restored region; a step of generating a prediction block of the current block using the prediction method determined to be applied to the current block; and a step of encoding information regarding the prediction method of the current block.

[0009] One aspect of the present disclosure is a method for transmitting data including a bitstream for an image, the method comprising: obtaining a bitstream for the image; and transmitting data including the bitstream. The obtaining step comprises: determining a prediction method to be applied to a current block among a plurality of prediction methods, wherein the plurality of prediction methods include at least one of a first prediction method derived based on prediction methods of neighboring blocks, a second prediction method derived based on information of the current block, or a third prediction method derived based on a matching cost of a previously restored region; generating a prediction block of the current block using the prediction method determined to be applied to the current block; and encoding information regarding the prediction method of the current block.

[0010] According to the present disclosure, a video encoding / decoding device can efficiently encode / decode an image and improve the objective / subjective image quality of a restored image by determining a prediction method of a current block based on at least one of prediction methods of surrounding blocks, information of the current block, matching cost of a restored area, or signaled information for prediction of the current block.

[0011] According to the present disclosure, an image encoding / decoding device can efficiently encode / decode an image and improve the objective / subjective image quality of a restored image by adaptively combining template regions, template feature extraction methods, and intra prediction mode derivation methods for predicting a current block.

[0012] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure.

[0013] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT (QuadTree plus BinaryTree TernaryTree) structure.

[0014] FIGS. 3A and 3B are diagrams illustrating multiple intra prediction modes, including wide-angle intra prediction modes.

[0015] Figure 4 is an example diagram of the surrounding blocks of the current block.

[0016] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure.

[0017] Figure 6 is a diagram for explaining the decoder-side intramode induction mode.

[0018] FIG. 7 is a diagram for explaining prediction of a current block according to one embodiment of the present disclosure.

[0019] FIGS. 8A, 8B, and 8C are diagrams illustrating template areas according to one embodiment of the present disclosure.

[0020] FIG. 9 illustrates an example of a histogram according to one embodiment of the present disclosure.

[0021] FIGS. 10A, 10B, and 10C illustrate examples of preset areas according to one embodiment of the present disclosure.

[0022] FIG. 11 is a diagram for explaining a method for determining a template feature-based prediction method based on information of a surrounding block according to one embodiment of the present disclosure.

[0023] FIG. 12 is a flowchart of an image decoding method according to one embodiment of the present disclosure.

[0024] FIG. 13 is a flowchart of an image encoding method according to one embodiment of the present disclosure.

[0025] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, in describing the present embodiments, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present embodiments.

[0026] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image encoding device and its subcomponents will be described with reference to the illustration in FIG. 1.

[0027] The video encoding device may be configured to include a picture segmentation unit (110), a prediction unit (120), a subtractor (130), a transformation unit (140), a quantization unit (145), a reordering unit (150), an entropy encoding unit (155), an inverse quantization unit (160), an inverse transformation unit (165), an adder (170), a loop filter unit (180), and a memory (190).

[0028] Each component of the video encoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0029] A single image (video) is composed of one or more sequences containing multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile or slice is divided into one or more Coding Tree Units (CTUs). Each CTU is then divided into one or more Coding Units (CUs) by a tree structure. Information applied to each CU is encoded as the syntax of the CU, and information commonly applied to CUs included in a CTU is encoded as the syntax of the CTU. In addition, information commonly applied to all blocks within a single slice is encoded as the syntax of the slice header, and information applied to all blocks constituting one or more pictures is encoded in the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referenced by multiple pictures is encoded in a Sequence Parameter Set (SPS). And, information commonly referenced by one or more SPS is encoded in a Video Parameter Set (VPS). In addition, information commonly applied to one tile or tile group may be encoded as syntax of a tile or tile group header. Syntaxes included in an SPS, PPS, slice header, tile or tile group header may be referred to as high level syntax.

[0030] The picture segmentation unit (110) determines the size of the CTU. Information about the size of the CTU (CTU size) is encoded as the syntax of SPS or PPS and transmitted to the image decoding device.

[0031] The picture segmentation unit (110) divides each picture constituting an image into a plurality of CTUs having a predetermined size, and then recursively divides the CTUs using a tree structure. A leaf node in the tree structure becomes a CU, which is a basic unit of encoding.

[0032] The tree structure may be a QuadTree (QT) in which an upper node (or parent node) is divided into four lower nodes (or child nodes) of the same size, a BinaryTree (BT) in which an upper node is divided into two lower nodes, or a TernaryTree (TT) in which an upper node is divided into three lower nodes in a 1:2:1 ratio, or a structure that mixes two or more of the QT structures, BT structures, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure may be used, or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure may be used. Here, BTTT may be combined and referred to as a MTT (Multiple-Type Tree).

[0033] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT structure.

[0034] As illustrated in FIG. 2, a CTU may first be split into a QT structure. The quadtree splitting may be repeated until the size of the splitting block reaches the minimum block size (MinQTSize) of the leaf node allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoding unit (155) and signaled to the image decoding device. If the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, it may be further split into one or more of the BT structure or the TT structure. There may be multiple splitting directions in the BT structure and / or the TT structure. For example, there may be two directions in which the block of the corresponding node is split horizontally and two directions in which the block is split vertically. As illustrated in FIG. 2, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether nodes have been split, and if splitting has occurred, a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (Binary or Ternary) are encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0035] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes of a lower layer, a CU split flag (split_cu_flag) indicating whether the node is split may be encoded. If the CU split flag (split_cu_flag) value indicates that the node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU (coding unit), which is a basic unit of encoding. If the CU split flag (split_cu_flag) value indicates that the node is split, the video encoding device starts encoding from the first flag in the above-described manner.

[0036] As another example of a tree structure, when QTBT is used, there may be two types: a type that horizontally splits the block of the corresponding node into two blocks of the same size (i.e., symmetric horizontal splitting) and a type that vertically splits it (i.e., symmetric vertical splitting). A split flag (split_flag) indicating whether each node of the BT structure is split into blocks of a lower layer and split type information indicating the type of split are encoded by the entropy encoding unit (155) and transmitted to the image decoding device. Meanwhile, there may additionally be a type that splits the block of the corresponding node into two blocks of an asymmetrical shape. The asymmetric shape may include a shape that splits the block of the corresponding node into two rectangular blocks with a size ratio of 1:3, or a shape that splits the block of the corresponding node in a diagonal direction.

[0037] A CU can have various sizes depending on the QTBT or QTBTTT partitioning from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block." Depending on the QTBTTT partitioning employed, the current block may be rectangular as well as square.

[0038] The prediction unit (120) predicts the current block and generates a prediction block. The prediction unit (120) includes an intra prediction unit (122) and an inter prediction unit (124).

[0039] In general, each current block within a picture can be predictively coded. Prediction of the current block can typically be performed using either intra-prediction (using data from the picture containing the current block) or inter-prediction (using data from a picture coded before the picture containing the current block). Inter-prediction encompasses both unidirectional and bidirectional prediction.

[0040] The intra prediction unit (122) predicts pixels within the current block using pixels (reference pixels) located around the current block within the current picture including the current block. There are multiple intra prediction modes depending on the prediction direction. For example, as shown in Fig. 3a, the multiple intra prediction modes may include two non-directional modes including the Planar mode and the DC mode, and 65 directional modes. The surrounding pixels to be used and the calculation formula are defined differently depending on each prediction mode.

[0041] For efficient directional prediction for a rectangular current block, directional modes (intra prediction modes 67 to 80 and -1 to -14) indicated by dotted arrows in Fig. 3b may be additionally used. These may be referred to as "wide-angle intra-prediction modes." In Fig. 3b, the arrows point to corresponding reference samples used for prediction, and do not indicate the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. Wide-angle intra-prediction modes are modes that perform prediction in the opposite direction of a specific directional mode without additional bit transmission when the current block is rectangular. At this time, among the wide-angle intra-prediction modes, some wide-angle intra-prediction modes available for the current block may be determined based on the ratio of the width and height of the rectangular current block. For example, wide-angle intra prediction modes (intra prediction modes 67 to 80) having an angle less than 45 degrees are available when the current block is a rectangular shape whose height is smaller than its width, and wide-angle intra prediction modes (intra prediction modes -1 to -14) having an angle greater than -135 degrees are available when the current block is a rectangular shape whose width is larger than its height.

[0042] The intra prediction unit (122) can determine the intra prediction mode to be used to encode the current block. In some examples, the intra prediction unit (122) can encode the current block using multiple intra prediction modes and select an appropriate intra prediction mode to be used from the tested modes. For example, the intra prediction unit (122) can calculate bit-rate distortion values ​​using rate-distortion analysis for multiple tested intra prediction modes and select the intra prediction mode with the best bit-rate distortion characteristics among the tested modes.

[0043] The intra prediction unit (122) selects one intra prediction mode from among multiple intra prediction modes and predicts the current block using surrounding pixels (reference pixels) and an operation formula determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0044] The inter prediction unit (124) generates a prediction block for the current block using a motion compensation process. The inter prediction unit (124) searches for a block most similar to the current block within reference pictures that were encoded and decoded before the current picture, and generates a prediction block for the current block using the searched block. Then, a motion vector (MV) corresponding to the displacement between the current block within the current picture and the prediction block within the reference picture is generated. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The motion information including information on the reference picture used to predict the current block and information on the motion vector is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0045] The inter prediction unit (124) may perform interpolation on a reference picture or a reference block to improve prediction accuracy. That is, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including the two integer samples. When a process of searching for a block most similar to the current block is performed on the interpolated reference picture, the motion vector can be expressed up to a precision in decimal units rather than a precision in integer sample units. The precision or resolution of the motion vector can be set differently for each target region to be encoded, such as a slice, tile, CTU, CU, etc. When such adaptive motion vector resolution (AMVR) is applied, information on the motion vector resolution to be applied to each target region must be signaled for each target region. For example, when the target region is a CU, information on the motion vector resolution applied to each CU is signaled. Information on the motion vector resolution may be information indicating the precision of a differential motion vector, which will be described later.

[0046] Meanwhile, the inter prediction unit (124) can perform inter prediction using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors indicating the block position most similar to the current block within each reference picture are used. The inter prediction unit (124) selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively, and searches for a block similar to the current block within each reference picture to generate a first reference block and a second reference block. Then, the first reference block and the second reference block are averaged or weighted averaged to generate a prediction block for the current block. Then, motion information including information on two reference pictures used to predict the current block and information on two motion vectors is transmitted to the entropy encoding unit (155). Here, reference picture list 0 may be composed of pictures that are before the current picture in display order among the restored pictures, and reference picture list 1 may be composed of pictures that are after the current picture in display order among the restored pictures. However, this is not necessarily limited to this, and restored pictures that are after the current picture in display order may be additionally included in reference picture list 0, and conversely, restored pictures that are before the current picture may be additionally included in reference picture list 1.

[0047] Various methods can be used to minimize the number of bits required to encode motion information.

[0048] For example, if the reference picture and motion vector of the current block are identical to those of a neighboring block, the motion information of the current block can be transmitted to the image decoding device by encoding information that can identify the neighboring block. This method is called 'merge mode'.

[0049] In merge mode, the inter prediction unit (124) selects a predetermined number of merge candidate blocks (hereinafter referred to as 'merge candidates') from the surrounding blocks of the current block.

[0050] As the surrounding blocks for deriving merge candidates, all or part of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (B2) adjacent to the current block within the current picture may be used, as illustrated in FIG. 4. In addition, a block located within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as a merge candidate. For example, a block co-located with the current block within the reference picture or blocks adjacent to the block at the co-located block may be additionally used as a merge candidate. If the number of merge candidates selected by the method described above is less than a preset number, a 0 vector is added to the merge candidates.

[0051] The inter prediction unit (124) uses these surrounding blocks to construct a merge list containing a predetermined number of merge candidates. Among the merge candidates included in the merge list, the merge candidate to be used as motion information of the current block is selected and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoding unit (155) and transmitted to the video decoding device.

[0052] Merge Skip mode is a special case of merge mode. After quantization, when all transform coefficients for entropy encoding are close to zero, only neighboring block selection information is transmitted without transmitting residual signals. By utilizing merge skip mode, relatively high encoding efficiency can be achieved for low-motion images, still images, and screen content images.

[0053] Hereinafter, merge mode and merge skip mode are collectively referred to as merge / skip mode.

[0054] Another method for encoding motion information is Advanced Motion Vector Prediction (AMVP) mode.

[0055] In AMVP mode, the inter prediction unit (124) derives predicted motion vector candidates for the motion vector of the current block using neighboring blocks of the current block. As neighboring blocks used to derive predicted motion vector candidates, all or some of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (B2) adjacent to the current block in the current picture as shown in FIG. 4 may be used. In addition, a block located in a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as the neighboring block used to derive predicted motion vector candidates. For example, a block co-located with the current block in the reference picture or blocks adjacent to the block in the co-located block may be used. If the number of motion vector candidates is less than a preset number by the method described above, a 0 vector is added to the motion vector candidates.

[0056] The inter prediction unit (124) derives predicted motion vector candidates using the motion vectors of these surrounding blocks, and determines a predicted motion vector for the motion vector of the current block using the predicted motion vector candidates. Then, the predicted motion vector is subtracted from the motion vector of the current block to produce a differential motion vector.

[0057] The predicted motion vector can be obtained by applying a predefined function (e.g., median, mean, etc.) to the predicted motion vector candidates. In this case, the image decoding device also knows the predefined function. In addition, since the surrounding blocks used to derive the predicted motion vector candidates are blocks that have already been encoded and decoded, the image decoding device also already knows the motion vectors of the surrounding blocks. Therefore, the image encoding device does not need to encode information to identify the predicted motion vector candidates. Therefore, in this case, information about the differential motion vector and information about the reference picture used to predict the current block are encoded.

[0058] Alternatively, the predicted motion vector can be determined by selecting one of the predicted motion vector candidates. In this case, information for identifying the selected predicted motion vector candidate is additionally encoded, along with information about the differential motion vector and the reference picture used to predict the current block.

[0059] The subtractor (130) subtracts the prediction block generated by the intra prediction unit (122) or inter prediction unit (124) from the current block to generate a residual block.

[0060] The transformation unit (140) transforms residual signals within a residual block having pixel values ​​in a spatial domain into transform coefficients in a frequency domain. The transformation unit (140) may transform the residual signals within the residual block using the entire size of the residual block as a transformation unit, or may divide the residual block into a plurality of sub-blocks and use the sub-blocks as transformation units to perform the transformation. Alternatively, the residual signals may be transformed using only the transformation domain sub-block as a transformation unit by dividing the sub-blocks into two sub-blocks, that is, a transformation domain and a non-transform domain. Here, the transformation domain sub-block may be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicating that only a sub-block has been converted, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit (155) and signaled to the image decoding device. In addition, the size of the conversion area sub-block may have a size ratio of 1:3 with respect to the horizontal axis (or vertical axis), and in this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding division is additionally encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0061] Meanwhile, the transformation unit (140) can individually perform transformations on the residual block in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a Multiple Transform Set (MTS). The transformation unit (140) can select one transformation function pair with the best transformation efficiency among the MTS and transform the residual block in the horizontal and vertical directions, respectively. Information (mts_idx) on the transformation function pair selected among the MTS is encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0062] The quantization unit (145) quantizes the transform coefficients output from the transform unit (140) using quantization parameters and outputs the quantized transform coefficients to the entropy encoding unit (155). The quantization unit (145) may directly quantize a related residual block without transformation for a certain block or frame. The quantization unit (145) may also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients within the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions may be encoded and signaled to an image decoding device.

[0063] The rearrangement unit (150) can perform rearrangement of coefficient values ​​for quantized residual values.

[0064] The reordering unit (150) can change a two-dimensional coefficient array into a one-dimensional coefficient sequence by using coefficient scanning. For example, the reordering unit (150) can output a one-dimensional coefficient sequence by scanning from the DC coefficient to the coefficients of the high-frequency region by using a zig-zag scan or a diagonal scan. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficients in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, the scanning method to be used may be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.

[0065] The entropy encoding unit (155) generates a bitstream by encoding a sequence of one-dimensional quantized transform coefficients output from the rearrangement unit (150) using various encoding methods such as CABAC (Context-based Adaptive Binary Arithmetic Code) and Exponential Golomb.

[0066] In addition, the entropy encoding unit (155) encodes information related to block division, such as CTU size, CU division flag, QT division flag, MTT division type, and MTT division direction, so that the image decoding device can divide the block in the same manner as the image encoding device. In addition, the entropy encoding unit (155) encodes information about the prediction type indicating whether the current block is encoded by intra prediction or inter prediction, and encodes intra prediction information (i.e., information about the intra prediction mode) or inter prediction information (information about the encoding mode of motion information (merge mode or AMVP mode), a merge index in the case of the merge mode, and a reference picture index and a differential motion vector in the case of the AMVP mode) according to the prediction type. In addition, the entropy encoding unit (155) encodes information related to quantization, that is, information about quantization parameters and information about a quantization matrix.

[0067] The inverse quantization unit (160) inversely quantizes the quantized transform coefficients output from the quantization unit (145) to generate transform coefficients. The inverse transform unit (165) transforms the transform coefficients output from the inverse quantization unit (160) from the frequency domain to the spatial domain to restore the residual block.

[0068] An adder (170) adds the restored residual block and the predicted block generated by the prediction unit (120) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting the next block.

[0069] The loop filter unit (180) performs filtering on restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transformation / quantization. The loop filter unit (180) may include all or part of a deblocking filter (182), a sample adaptive offset (SAO) filter (184), and an adaptive loop filter (ALF, 186) as an in-loop filter.

[0070] The deblocking filter (182) filters the boundaries between restored blocks to remove blocking artifacts caused by block-based encoding / decoding, and the SAO filter (184) and the ALF (186) perform additional filtering on the deblocking-filtered image. The SAO filter (184) and the ALF (186) are filters used to compensate for the differences between restored pixels and original pixels caused by lossy coding. The SAO filter (184) improves not only subjective image quality but also encoding efficiency by applying an offset in units of CTUs. In contrast, the ALF (186) performs block-based filtering, and compensates for distortion by applying different filters by distinguishing the edges and degrees of variation of the corresponding block. Information on filter coefficients to be used in the ALF can be encoded and signaled to an image decoding device.

[0071] The restored blocks filtered through the deblocking filter (182), SAO filter (184), and ALF (186) are stored in the memory (190). When all blocks within a picture are restored, the restored picture can be used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0072] The video encoding device can store the bitstream of encoded video data on a non-transitory storage medium or transmit it to the video decoding device using a communication network.

[0073] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image decoding device and its subcomponents will be described with reference to FIG. 5.

[0074] The video decoding device may be configured to include an entropy decoding unit (510), a rearrangement unit (515), an inverse quantization unit (520), an inverse transformation unit (530), a prediction unit (540), an adder (550), a loop filter unit (560), and a memory (570).

[0075] Similar to the video encoding device of FIG. 1, each component of the video decoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0076] The entropy decoding unit (510) decodes the bitstream generated by the image encoding device to extract information related to block division, thereby determining the current block to be decoded, and extracts prediction information, information on residual signals, etc. required to restore the current block.

[0077] The entropy decoding unit (510) extracts information about the CTU size from the Sequence Parameter Set (SPS) or the Picture Parameter Set (PPS), determines the size of the CTU, and divides the picture into CTUs of the determined size. Then, the CTU is determined as the top layer of the tree structure, i.e., the root node, and the CTU is divided using the tree structure by extracting division information about the CTU.

[0078] For example, when splitting a CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, the second flag (mtt_split_flag) related to the splitting of MTT and the split direction (vertical / horizontal) and / or split type (binary / ternary) information are extracted, and the corresponding leaf nodes are split into the MTT structure. Accordingly, each node below the leaf nodes of QT are split recursively into the BT or TT structure.

[0079] As another example, when splitting a CTU using the QTBTTT structure, the CU split flag (split_cu_flag) indicating whether the CU is split is first extracted, and if the block is split, the first flag (QT_split_flag) may be extracted. During the splitting process, each node may undergo zero or more repeated QT splits followed by zero or more repeated MTT splits. For example, a CTU may undergo an MTT split right away, or conversely, may undergo only multiple QT splits.

[0080] As another example, when splitting a CTU using the QTBT structure, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Furthermore, for nodes corresponding to leaf nodes of QT, a split flag (split_flag) indicating whether to further split into BTs and splitting direction information are extracted.

[0081] Meanwhile, when the entropy decoding unit (510) determines the current block to be decoded by using the division of the tree structure, it extracts information on the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-prediction, the entropy decoding unit (510) extracts syntax elements for intra-prediction information (intra-prediction mode) of the current block. If the prediction type information indicates inter-prediction, the entropy decoding unit (510) extracts syntax elements for inter-prediction information, i.e., information indicating a motion vector and a reference picture referenced by the motion vector.

[0082] Additionally, the entropy decoding unit (510) extracts information about the quantized transform coefficients of the current block as information related to quantization and information about residual signals.

[0083] The rearrangement unit (515) can change the sequence of one-dimensional quantized transform coefficients entropy-decoded in the entropy decoding unit (510) back into a two-dimensional coefficient array (i.e., block) in the reverse order of the coefficient scanning performed by the image encoding device.

[0084] The inverse quantization unit (520) inversely quantizes the quantized transform coefficients and inversely quantizes the quantized transform coefficients using the quantization parameters. The inverse quantization unit (520) may also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in two dimensions. The inverse quantization unit (520) may perform inverse quantization by applying a matrix of quantized coefficients (scaling values) from an image encoding device to a two-dimensional array of quantized transform coefficients.

[0085] The inverse transform unit (530) inversely transforms the inverse quantized transform coefficients from the frequency domain to the spatial domain to restore residual signals, thereby generating a residual block for the current block.

[0086] In addition, when the inverse transform unit (530) inversely transforms only a portion of a transform block (sub-block), it extracts a flag (cu_sbt_flag) indicating that only a sub-block of the transform block has been transformed, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block, and inversely transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to restore residual signals, and fills “0” values ​​with residual signals for areas that have not been inversely transformed, thereby generating a final residual block for the current block.

[0087] In addition, when MTS is applied, the inverse transform unit (530) determines a transform function or a transform matrix to be applied in the horizontal and vertical directions using MTS information (mts_idx) signaled from the image encoding device, and performs inverse transform on the transform coefficients within the transform block in the horizontal and vertical directions using the determined transform function.

[0088] The prediction unit (540) may include an intra prediction unit (542) and an inter prediction unit (544). The intra prediction unit (542) is activated when the prediction type of the current block is intra prediction, and the inter prediction unit (544) is activated when the prediction type of the current block is inter prediction.

[0089] The intra prediction unit (542) determines the intra prediction mode of the current block among a plurality of intra prediction modes from the syntax elements for the intra prediction mode extracted from the entropy decoding unit (510), and predicts the current block using reference pixels around the current block according to the intra prediction mode.

[0090] The inter prediction unit (544) uses the syntax elements for the inter prediction mode extracted from the entropy decoding unit (510) to determine the motion vector of the current block and the reference picture referenced by the motion vector, and predicts the current block using the motion vector and the reference picture.

[0091] An adder (550) adds the residual block output from the inverse transform unit (530) and the predicted block output from the inter prediction unit (544) or the intra prediction unit (542) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting a block to be decoded later.

[0092] The loop filter unit (560) may include a deblocking filter (562), an SAO filter (564), and an ALF (566) as in-loop filters. The deblocking filter (562) deblocks the boundaries between restored blocks to remove blocking artifacts caused by block-by-block decoding. The SAO filter (564) and the ALF (566) perform additional filtering on restored blocks after deblocking filtering to compensate for differences between restored pixels and original pixels caused by lossy coding. The filter coefficients of the ALF are determined using information about filter coefficients decoded from the non-stream.

[0093] The restored blocks filtered through the deblocking filter (562), SAO filter (564), and ALF (566) are stored in the memory (570). When all blocks within a picture are restored, the restored picture is used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0094] Hereinafter, an improved coding tool performed by the aforementioned video encoding device or video decoding device is disclosed. The coding device refers to at least one of the video encoding device or the video decoding device.

[0095] The decoder-side intra mode derivation (DIMD) technique is a technique in which a decoder derives an intra prediction mode of a current block based on a template region composed of previously reconstructed reference samples, without receiving intra prediction mode information of the current block from an encoder. Here, the template region refers to a region in which encoding or decoding has already been completed. A region having image characteristics similar to those of the current block can be selected as the template region. For example, a region including previously reconstructed samples of at least one neighboring block on the left or top of the current block can be selected as the template region. It refers to a region having image characteristics similar to those of the current block.

[0096] Figure 6 is a diagram for explaining the decoder-side intramode induction mode.

[0097] Referring to FIG. 6, the DIMD mode is a mode that calculates gradients of a template region (620) used for prediction of a current block (610), generates a histogram of gradients (HoG) based on the gradients, selects one or more intra prediction modes from the gradient histogram, and weights and combines prediction blocks according to the selected intra prediction mode to generate a final prediction block of the current block (610).

[0098] The template region (620) may be determined based on a flag that is predefined, parsed from the bitstream, or based on an index indicating the template region (620) within the bitstream. The template region (620) may be included in the current picture including the current block (610), or may be included in a picture different from the current picture. The template region (620) may be adjusted based on the availability of restored reference samples. As shown in FIG. 6, the template region (620) may extend in the upper right direction of the current block by the horizontal length of the current block, or may extend in the lower left direction of the current block by the vertical length of the current block.

[0099] By applying a predetermined filter to the restored reference samples of the template area (620), the gradients of the template area (620) are calculated. The predetermined filter is applied to the reference samples located in the second lines for the current block (610), so that vertical and horizontal gradient values ​​can be calculated in the 3X3 area, the 3X1 area, or the 1X3 area of ​​each reference sample. In Fig. 6, the filter is applied to the positions of the target samples (630) indicated by the bold lines in the template area (620), so that the vertical gradient G at each target sample position ver and horizontal gradient G Hor is calculated. For example, when the size of the filter is 3x3, the filter is applied to a local region of size 3x3 containing one target sample. The size of the local region is equal to the size of the filter.

[0100] A given filter is a filter for calculating the gradient of reference samples within the template area (620). As the filter, an edge detection filter such as a Sobel filter, a Scharr filter, a Prewitt filter, or a Roberts filter may be used. Each filter includes a horizontal filter that outputs a horizontal gradient and a vertical filter that outputs a vertical gradient.

[0101] Table 1 shows examples of filters of various sizes.

[0102] Horizontal FilterVertical FilterFilter 1-1 0 1-2 0 2-1 0 1-1 -2 -10 0 01 2 1Filter 2-1 0 1-1 0 1-1 0 1-1 -1 -10 0 01 1 1Filter 3-1 0 1-101

[0103] Afterwards, a gradient histogram is generated based on the calculated gradients. The gradient histogram represents the gradient accumulation of each directional prediction mode. The histogram can be represented as a graph with the x-axis representing the directional prediction modes and the y-axis representing the gradient magnitude. Specifically, based on the calculated vertical and horizontal gradients, the gradient direction (θ) and the gradient magnitude (I) are derived in each target sample or local region. As an example, the gradient direction and the gradient magnitude can be determined as in Equation 1. In Equation 1, G ver Silver vertical gradient, G Hor represents a horizontal gradient. In other examples, the gradient magnitude may be the sum of the absolute value of the horizontal gradient and the absolute value of the vertical gradient, or the absolute value of the sum of the horizontal / vertical gradients.

[0104]

[0105]

[0106] The directional mode of the intra prediction closest to the gradient direction θ of each target sample in the template area (620) is determined. The gradient magnitude of the target sample is accumulated in the bin of the determined directional prediction mode.

[0107] In this way, a gradient histogram is generated by accumulating the gradient magnitudes of all target samples into bins indicating directional prediction modes.

[0108] Thereafter, up to N directional prediction modes having the largest accumulated gradient size in the gradient histogram are selected, and prediction blocks of the current block (610) are generated based on the N directional prediction modes. Here, N is a positive integer greater than or equal to 1, for example, N may be 5. Additional prediction blocks of the current block (610) may be generated based on a prediction mode selected according to a template cost among non-directional prediction modes such as a planar mode or a block vector-based prediction mode.

[0109] By weighting the predicted blocks for the current block (610), the final predicted block of the current block (610) is derived. Here, if the histogram magnitude of either the left neighboring block or the upper neighboring block is more than twice that of the other, the weights for the predicted blocks can be adjusted depending on the position.

[0110] The DIMD mode can be applied to not only the luminance channel but also the chrominance channel. In the chrominance channel, the directional prediction mode with the largest accumulated gradient magnitude can be used to generate the prediction block of the current block (610).

[0111] Whether to perform DIMD mode may be predefined or may be determined based on flags parsed from the bitstream, the size and shape of the block, or flags of other modes.

[0112] In DIMD mode, the optimal method for predicting the current block may not be selected. For example, a region other than the region adjacent to the current block may be advantageous as a template region. For another example, a filter other than the Sobel filter may be more appropriate for extracting features from the template region. For another example, if the correlation between the luminance and chrominance channels is low, using only the chrominance channel information rather than the luminance channel information may be more appropriate for predicting the chrominance block of the current block.

[0113] The present disclosure relates to a template feature-based prediction method that modifies the DIMD mode to select an optimal method for predicting a current block.

[0114] First, a flag indicating whether a template feature-based prediction method is applied to the current block can be signaled. If the value of the flag information is 1, the flag information can be used to indicate that the intra mode of the current CU is determined based on the template feature-based prediction method. If the value of the flag information is 0, it can indicate that the current CU is not predicted by the template feature-based prediction method and a conventional intra coding technique can be used. The flag information can be encoded at the CU level or the PU level using the CABAC method.

[0115] FIG. 7 is a diagram for explaining prediction of a current block according to one embodiment of the present disclosure.

[0116] Referring to FIG. 7, according to one embodiment of the present disclosure, a template feature-based prediction method is a method of generating a prediction block of a current block based on features of a template region of the current block. The template feature-based prediction method may include a template selection step (S710), a feature extraction step (S720), an intra-prediction mode derivation step (S730), a prediction block candidate generation step (S740), and a final prediction block generation step (S750).

[0117] Template selection step S710 is a step of selecting one template region from among multiple template regions. In template selection step S710, TPL_1 to TPL_L indicate L template regions that can be used for prediction of the current block.

[0118] Feature extraction step S720 is a step for extracting features of a template region by applying a filter to the selected template region. In feature extraction step S720, a feature extraction candidate (FEC) refers to a method for extracting features from the template region. For example, the ith FEC refers to a method for extracting features from the template region using the ith filter.

[0119] Mode derivation step S730 is a step of deriving one or more intra prediction modes based on the extracted features. A mode derivation candidate (MDC) refers to a method of deriving one or more intra prediction modes. In other words, a certain mode derivation candidate may be a method of deriving intra prediction modes based on a histogram of template features, and another mode derivation candidate may be a method of deriving intra prediction modes using a neural network model. Furthermore, there may be other methods of constructing the histogram and other methods of selecting the intra prediction mode. The intra prediction modes may include directional prediction modes and non-directional prediction modes. In the present disclosure, intra prediction modes derived based on template features may be referred to as decoder-derived intra modes (DDIM).

[0120] Prediction block candidate generation step S740 is a step of generating prediction block candidates for the current block using one or more determined intra prediction modes. If one intra prediction mode is determined in step S730, one prediction block candidate is generated for the current block. If multiple intra prediction modes are determined in step S730, multiple prediction block candidates are generated for the current block. In Fig. 7, DDIM_K indicates the Kth prediction block candidate generated by the intra prediction mode corresponding to the index K.

[0121] Final prediction block generation step S750 generates the final prediction block of the current block by weighting one or more prediction block candidates of the current block. The weights applied to each prediction block candidate can be adaptively determined.

[0122] The final prediction block of the current block is used to generate a residual block in an image encoding device and to generate a restoration block in an image decoding device.

[0123] As described above, according to one embodiment of the present disclosure, the image encoding / decoding device can efficiently encode / decode an image and improve the objective / subjective image quality of a restored image by adaptively combining methods for extracting template regions, template features, and deriving intra prediction modes for prediction of a current block.

[0124] The optimal combination can be determined based on signaling according to a preset method, or can be determined based on predetermined conditions in the decoder.

[0125] Below, examples of template areas within the template selection step S710 are described in detail.

[0126] FIGS. 8A, 8B, and 8C are diagrams illustrating template areas according to one embodiment of the present disclosure.

[0127] Referring to Fig. 8a, the current template (813) is an area composed of restored reference samples, and may be an area adjacent to the current block (811) or an area within a preset distance from the current block (811). The area within a preset distance includes an area spaced apart from the current block or luminance block by a preset distance.

[0128] Here, the preset distance may mean the distance from the current block (811) to an arbitrary pixel, or may mean the distance including a neighboring block adjacent to the current block (811) and another neighboring block adjacent to the neighboring block.

[0129] For example, the current template (813) may include three rows of restored reference samples located at the left and top of the current block (811), or may be an area composed of the restored reference samples. When the position coordinate of the upper left sample of the current block (811) is (0, 0), the position coordinate of the upper left reference sample of the current template (813) may be (-3, -3).

[0130] The horizontal length M of the upper reference samples and the vertical length N of the left reference samples can be determined based on the size or shape of the current block (811). For example, the horizontal length M of the upper reference samples can be equal to W+3, and the vertical length N of the left reference samples can be equal to H+3. As another example, the horizontal length M of the upper reference samples can be longer than W+3, and the vertical length N of the left reference samples can be longer than H+3.

[0131] As another example, the current template (813) may be an area of ​​reference samples that are not adjacent to the current block (811). In other words, the current template (813) may be an area of ​​restored samples spaced apart from the current block (811) by a preset distance. The horizontal length M of the upper reference samples may be equal to 3+G+W, and the vertical length N of the left reference samples may be equal to 3+G+H. Here, G indicates a gap between the current block (811) and the current template (813).

[0132] As another example, all or part of the current template (813) may be used as the template area of ​​the current block (811). For example, only the area located above or to the left of the current block (811) within the current template (813) may be used as the template area of ​​the current block (811).

[0133] Referring to FIG. 8b, as a template area, a reference block (831) or a reference template (833) within a reference frame (830) referenced by the current frame (820) can be used.

[0134] In one embodiment, when the current block (821) is predicted based on two or more prediction modes including an inter prediction mode, a reference block (830) or a reference template (833) within a reference frame (830) used in the inter prediction mode can be used as a template region.

[0135] For example, when the Combined Inter-Intra Prediction (CIIP) mode is applied to the current block (821), the current block is predicted from a weighted sum between a prediction block according to the inter prediction mode and a prediction block according to the intra prediction mode. Here, a reference frame (830) is used to inter-predict the current block, and a reference block (831) or a reference template (833) within the reference frame (830) can be used as a template region of the current block (821).

[0136] As another example, when GPM (Geometric Partitioning Mode) is applied to the current block (821), the current block (821) is divided into two or more sub-blocks having various shapes, and the sub-blocks are predicted by a prediction method in which at least one of reference pictures, motion vectors, or prediction directions is different. Here, a reference frame (830) is used to predict the sub-blocks, and a reference block (831) or a reference template (833) within the reference frame (830) can be used as a template region of the current block (821).

[0137] A reference frame (830) is a frame referenced by the current frame (820) including the current block (821) or the slice to which the current block (821) belongs within the current frame (820), and may be referred to as a reference picture or col-picture. Information for identifying the reference frame (830) may be determined based on information signaled between the encoder and the decoder.

[0138] The reference block (831) is an area within the reference frame (830) indicated by the motion vector or block vector of the current block (821). That is, the reference block (831) is an area corresponding to the current block (821) within the reference frame (830). The reference block (831) may have the same size and shape as the current block (821). Alternatively, the reference block (831) may have a different size or shape than the current block (821).

[0139] Alternatively, the reference block (831) may be a collocated block or a collocated block of the current block (821) within the collocated picture. The collocated block refers to a block having the same position as the current block (821) within the collocated picture or a block at the lower right of the block. Specifically, when the sizes of the current frame (820) and the reference frame (830) are the same, the position of the current block (821) and the position of the collocated block are the same. When the sizes of the current frame (820) and the reference frame (830) are different, the relative position of the current block (821) with respect to the current frame (820) is the same as the relative position of the collocated block with respect to the reference frame (830).

[0140] The reference template (833) is a peripheral area of ​​the reference block (831), an area adjacent to the reference block (831) or an area within a preset distance from the reference block (831). For example, the reference template (833) may be composed of three rows of restored reference samples located on the left and top sides of the reference block (831), respectively.

[0141] In another embodiment, the reference frame (830) and the reference block (831) may be referenced by the surrounding blocks or adjacent blocks of the current block (821). If the surrounding blocks of the current block (821) are predicted using the reference block (831) or the reference template (833), the template region of the current block (821) may be determined to be the reference block (831) or the reference template (833). For example, if the surrounding blocks of the current block (821) are inter-predicted based on the reference block (831) or the reference template (833) within the reference frame (830) indicated by the motion vector of the surrounding blocks, the reference block (831) or the reference template (833) may be used as the template region of the current block (821). In another example, if a neighboring block of the current block (821) is predicted based on a block indicated by a block vector of the neighboring block within the current frame (820), the block can be used as a template region of the current block (821).

[0142] As a template area of ​​the current block (821), all or part of the reference block (831) may be used. Alternatively, all or part of the reference template (833) may be used.

[0143] Referring to FIG. 8c, when the current block is a chrominance block (841), at least one of a chrominance template (843), a luminance block (851), or a luminance template (853) can be used as a template area of ​​the chrominance block (841).

[0144] Specifically, the color difference template (843) is an area adjacent to the color difference block (841) or an area within a preset distance from the color difference block (841).

[0145] The luminance block (851) is an area of ​​the luminance channel corresponding to the chrominance block (841) of the chrominance channel. The luminance block (851) may have the same size or shape as the chrominance block (841), or may have a different size or shape. The size relationship between the luminance block and the chrominance block can be defined based on the chrominance sampling rate (or format). The position of the luminance block (851) in the luminance channel can be the same as the position of the chrominance block (841) in the chrominance channel. Alternatively, the position of the luminance block (851) can be determined based on a predefined relationship or condition between the luminance channel and the chrominance channel.

[0146] The luminance template (853) is an area of ​​the luminance channel corresponding to the peripheral area or adjacent area of ​​the chrominance block (841). The position of the luminance template (853) in the luminance channel may be the same as the position of the chrominance template (843) in the chrominance channel. Alternatively, the position of the luminance template (853) may be determined based on a predefined relationship or condition between the luminance channel and the chrominance channel.

[0147] When the size or shape of the chroma block (841) and the luminance block (851) are the same, the size and shape of the luminance template (853) may be the same as the chroma template (843). When the size or shape of the chroma block (841) and the luminance block (851) are different, the size and shape of the luminance template (853) may be different from the chroma template (843). For example, when the size of the luminance block (851) is four times the size of the chroma block (841), the size or number of samples of the luminance template (853) may be four times larger than that of the chroma template (843).

[0148] As the template area of ​​the chrominance block (841), all or part of the chrominance template (843) may be used. Alternatively, all or part of the luminance block (851) may be used. Alternatively, all or part of the luminance template (853) may be used.

[0149] Below, the feature extraction step S720 is described in detail.

[0150] An image encoding / decoding device extracts one or more features from a template region using a filter or a neural network model. Specifically, the image encoding / decoding device can extract template features from previously restored samples within the template region or syntax information for restoring the samples. Furthermore, the image encoding / decoding device can further utilize the position, size, and shape of the current block, the position, size, or shape of the template region, etc., for feature extraction.

[0151] An image encoding / decoding device can extract one or more features by applying one or more filters to a template region. The extracted template features can include directional components of the template region and the magnitudes of each directional component, i.e., gradient directions and gradient magnitudes.

[0152] The extracted feature values ​​may vary depending on the type of filter for feature extraction, the size of the filter, or the padding of the template area. In step S720, one FEC refers to an extraction method defined according to a specific type of filter, a specific size of the filter, or padding.

[0153] As a type of filter, an edge detection filter such as a Sobel filter, a Scharr filter, a Prewitt filter, a Roberts filter, or a Laplacian filter can be used.

[0154] As an example, by applying a 3x3 sized filter to each local region of the template area, the magnitude of the horizontal gradient and the magnitude of the vertical gradient of each local region can be derived.

[0155] As another example, by applying different types of filters to each local region of the template region, two horizontal gradient magnitudes and two vertical gradient magnitudes can be derived for each local region. The arithmetic results, such as the average and weighted sum of the gradient magnitudes, can be determined as characteristics of the local region.

[0156] As another type of filter, a neural network-based filter can be used. Specifically, a convolutional neural network (CNN) filter can be applied to the template region to extract template features. The CNN filter can be pre-trained to extract template features, or can be one of the filters that can be trained during encoding / decoding. For example, by applying a CNN filter to pre-restored reference samples within the template region, one or more directional components and the magnitude of the directional components can be extracted as features of the template region.

[0157] The size of the filter may be equal to or smaller than the template area of ​​the current block. As illustrated in Fig. 6, when the size of the filter is smaller than the template area, the filter is applied to each of the regions of the template area, thereby extracting and aggregating multiple features in the template area.

[0158] Before applying a filter, padding may be performed at the border of the template region. For example, the template region may be expanded by copying samples at the border position of the template region to the surrounding region of the template region. The expanded size of the template region may be determined arbitrarily or according to a predefined definition. As another example, samples adjacent to the template region may be padded with samples of a preset value. The preset value may be 0 or determined based on the bit depth. The filter may be applied to both the template region and the expanded region.

[0159] Below, the intra prediction mode derivation step S730 is described in detail.

[0160] In step S730 of deriving an intra prediction mode, one or more directional prediction modes are derived from features of a template region using a histogram or a neural network, and intra prediction modes are determined based on the directional prediction modes and predetermined non-directional prediction modes. In another embodiment, predetermined non-directional prediction modes may be excluded.

[0161] In one embodiment, a histogram is generated based on features of a template region, and one or more directional prediction modes are derived based on the histogram.

[0162] Here, the derived directional prediction modes may vary depending on the histogram accumulation method and the selection method of the directional prediction modes. In step S730, a single MDC may be defined based on a specific definition of the histogram, a specific histogram accumulation method, and a specific selection method of the directional prediction modes.

[0163] FIG. 9 illustrates an example of a histogram according to one embodiment of the present disclosure.

[0164] Referring to Figure 9, as a definition of a histogram, a bin corresponding to the horizontal axis of the histogram may indicate a directional prediction mode corresponding to a directional component. For example, one bin may indicate one directional prediction mode. The vertical axis of the histogram may indicate the cumulative size for the bin. For example, the vertical axis of the histogram may indicate the accumulation of a specific directional prediction mode.

[0165] Each time a filter is applied to a local region of the template area, the cumulative size of the bins corresponding to the directional components of the local region increases.

[0166] As a method of accumulating a histogram, the accumulated size of a bin can be updated based on the directional component of a local region. For example, when a 3x3 Sobel filter is applied to a 3x3 local region within a template region to derive a gradient direction and a gradient magnitude, the directional prediction mode closest to the gradient direction can be identified, and the gradient magnitude can be accumulated in a bin corresponding to the identified directional prediction mode. The directional prediction mode closest to the gradient direction is bin k When referred to as bin k The gradient size is accumulated.

[0167] As another accumulation method of the histogram, the accumulation size of multiple bins can be updated based on the directional component of a local region. For example, the bin of the directional prediction mode closest to the induced gradient direction and the bins of the directional prediction modes adjacent to the directional prediction mode can be updated. The directional prediction mode closest to the induced gradient direction is bin k When referred to as bin k-1 , bin k , and bin k+1 Each gradient size can be accumulated in the bin of the above directional prediction mode. k The cumulative size of the bins of adjacent directional prediction modes k-1 , bin k+1 It can be larger than the cumulative size. It can be expressed as in mathematical expression 2.

[0168]

[0169] In Equation 2, bin i represents the bin of the i-th directional prediction mode, and F represents the induced gradient magnitude. a i is a weight to reflect the gradient size, and can be predefined as a specific real number or derived from additional information of the current block or template area. a iis a real number greater than or equal to 0 and less than or equal to 1, and its sum can be 1. '+=' is the addition assignment operator.

[0170] As an example, the bin for the directional prediction mode closest to the gradient direction of the local region is bin k When, bin k-1 += 0.25×F, bin k += F, bin k+1 The histogram can be updated like this: += 0.25×F.

[0171] After the histogram is generated, one or more directional prediction modes are derived from the histogram. Intra-prediction modes corresponding to n bins in the histogram can be selected as the directional prediction modes for the current block. Here, n is a positive integer greater than or equal to 1. For example, n can be 5.

[0172] For example, one or more bins can be selected based on the cumulative size of the bins within a histogram, and directional prediction modes corresponding to the selected bins can be derived. For example, n directional prediction modes can be selected in descending order of cumulative size.

[0173] As another example, bins having a cumulative size greater than a threshold value can be selected, and directional prediction modes corresponding to the selected bins can be derived.

[0174] As another example, one or more bins may be selected based on at least one of the mean, mode, or median of the cumulative magnitudes of the bins, and directional prediction modes corresponding to the selected bins may be derived. Here, the mean of the cumulative magnitudes refers to a value obtained by dividing the sum of the cumulative magnitudes of the bins by the number of bins. For example, a bin corresponding to at least one of the mean, mode, or median of the cumulative magnitudes may be selected. As another example, in addition to the bin corresponding to at least one of the mean, mode, or median of the cumulative magnitudes, n-1 bins adjacent to the bin may be selected.

[0175] As another example, in a histogram, one or more bins may be selected based on at least one of bins corresponding to a local maximum or bins corresponding to a local minimum, and directional prediction modes corresponding to the selected bins may be derived. Alternatively, bins having a large cumulative magnitude among bins corresponding to a local maximum, bins having a large cumulative magnitude among bins corresponding to a local minimum, or a combination thereof may be selected. Directional prediction modes corresponding to the selected bins are derived. A bin corresponding to a local maximum refers to a bin having a large cumulative magnitude compared to adjacent bins within a given range. A bin corresponding to a local minimum refers to a bin having a small cumulative magnitude compared to adjacent bins within a given range.

[0176] In another embodiment, one or more directional prediction modes are derived from features of a template region using a neural network.

[0177] Specifically, the image encoding / decoding device can derive one or more intra prediction modes of the current block by applying a pre-trained neural network to features extracted from a template region.

[0178] For training a neural network, information such as the type of filter used to extract features of the template region, coefficients of the filter, extracted features, size, position, shape, channel of the current block, size, position, shape, channel of the template region, or relationship between the current block and the template region can be used.

[0179] In response to receiving the above information, the neural network can be trained to output a probability value or confidence value indicating that the intra prediction mode is the optimal prediction mode.

[0180] For example, a trained neural network, in response to the input information, outputs a probability value or confidence value indicating that each directional prediction mode is the optimal prediction mode for the current block. One or more directional prediction modes with high probability values ​​are included in the intra prediction mode of the current block.

[0181] As another example, in response to the input information, the neural network outputs a probability value or confidence value indicating that each intra prediction mode is the optimal prediction mode for the current block. The intra prediction modes include directional and non-directional prediction modes. One or more intra prediction modes with high probability values ​​are included in the intra prediction modes of the current block.

[0182] Neural networks can be learnable or tunable models during image encoding or decoding.

[0183] As described above, when one or more directional prediction modes are derived from features of a template region using a histogram or a neural network, intra prediction modes of the current block are derived by adding a predetermined non-directional prediction mode to one or more directional prediction modes.

[0184] Here, the predetermined non-directional prediction mode may refer to at least one of a planar mode, a planar vertical mode, a planar horizontal mode, a DC mode, or a block vector-based mode.

[0185] Below, the prediction block candidate generation step S740 and the final prediction block generation step S750 are described.

[0186] In step S740, DDIM_K refers to a prediction block candidate generated by the Kth intra prediction mode.

[0187] If the derived intra prediction mode is a directional prediction mode, the final prediction block of the current block is generated from surrounding reference samples of the current block using the directional prediction mode.

[0188] Alternatively, when the derived intra prediction modes include only multiple directional prediction modes, prediction block candidates for the current block are generated from surrounding reference samples of the current block using the directional prediction modes, and the final prediction block is generated by weighting the prediction block candidates.

[0189] Alternatively, when the derived intra prediction modes include one directional and one or more non-directional prediction modes, the final prediction block of the current block is generated by weighting blocks predicted by the derived intra prediction modes. Specifically, a first prediction block candidate of the current block is generated using the derived directional prediction mode. Second prediction block candidates of the current block are generated based on one or more non-directional prediction modes. The final prediction block of the current block is generated by averaging or weighting the first and second prediction block candidates. When the intra prediction modes are derived from a histogram, the weight of the first prediction block candidate may be predefined or determined based on the cumulative size of a bin in the histogram. When the intra prediction modes are derived using a neural network, the weight of the first prediction block candidate may be predefined or determined based on a probability value which is an output of the neural network. The weight of the second prediction block candidate may be a predefined value or a value obtained by subtracting the weight of the first prediction block candidate from a specific value. At this time, the sum of the weights of the first prediction block candidate and the weights of the second prediction block candidates may be 1.

[0190] Alternatively, when the derived intra prediction modes include multiple directional and non-directional prediction modes, the final prediction block of the current block is generated by weighting blocks predicted by the derived intra prediction modes. Specifically, first prediction block candidates are generated based on the multiple induced directional prediction modes. Furthermore, second prediction block candidates are generated based on the multiple non-directional prediction modes. The final prediction block of the current block is generated by averaging or weighting the first prediction block candidates and the second prediction block candidates. The weight of each first prediction block candidate may be predefined, determined based on a histogram cumulative size, or determined based on a probability value which is an output of a neural network. The larger the cumulative size or probability value, the larger the weight of the first prediction block candidate. The weight of the second prediction block candidate may be a predefined value, or a value obtained by subtracting the weight of the first prediction block candidate from a specific value.

[0191] Below, a method for predicting the current block based on a template feature-based prediction method is described.

[0192] The template feature-based prediction method can be determined through a combination of at least two of the template region selected in step S710, the method for extracting template features in step S720, and the method for deriving intra prediction modes in step S730. The term 'template feature-based prediction method candidate' refers to a specific combination of at least two of the template region, the method for extracting template features, and the method for deriving intra prediction modes.

[0193] In the first embodiment, the prediction method to be applied to the current block is determined based on signaling information between the encoder and decoder.

[0194] First, the encoder and decoder can define and pre-store a list of template feature-based prediction methods. This list includes combinations based on template region, filter type, filter size, etc., along with an index for each combination. The list of template feature-based prediction methods can be exemplified as shown in Table 2.

[0195] ddim_gen_list_index Template area Feature extraction method Mode derivation method Filter type Filter size Feature analysis method Number of DDIM estimates 0 Templates 813 Sobel 3x3 histogram 1 15 2 CNN neural network 2 35x5 5 4 Templates 813+ Templates 853 Sobel 3x3 histogram 5 CNN 5x5 neural network 6 Templates 813+ Block 831+ Templates 833 Sobel 3x3 histogram

[0196] The encoder searches for at least one of template feature-based prediction methods as a prediction method for a current block, and transmits an index of the searched prediction method to the encoder. The decoder can predict the current block using a template feature-based prediction method corresponding to the received index by referring to a pre-stored list. For example, the encoder can transmit syntax ddim_gen_list_index to the decoder. The encoder can transmit ddim_gen_list_index with a value of 2 to the decoder. The decoder predicts the current block using the current template (813), a 3x3 sized Sobel filter, a histogram, and five DDIMs corresponding to ddim_gen_list_index 2.

[0197] Below, the second through fourth embodiments relate to determining a template feature-based prediction method based on predetermined conditions without signaling between the encoder and decoder. The second through fourth embodiments have the advantage of reducing signaling overhead.

[0198] Meanwhile, even if information indicating which of the template feature-based prediction methods is used is not signaled, information indicating which of the second to fourth embodiments is used may be signaled.

[0199] Embodiments 2 through 4 may be applied under specific conditions. Embodiments 2 through 4 may be applied only to luminance blocks. Alternatively, Embodiments 2 through 4 may be applied only to chrominance blocks. Alternatively, Embodiments 2 through 4 may be applied to both luminance blocks and chrominance blocks. Alternatively, Embodiments 1 through 3 may not be applied to blocks smaller than a predetermined size. The predetermined size may refer to 64 samples.

[0200] In detail, in the second embodiment, the prediction method to be applied to the current block can be selected from template feature-based prediction method candidates based on information of the current block.

[0201] For example, a template feature-based prediction method can be determined based on the luminance channel / chrominance channel, width, height, size, shape, aspect ratio, position, and type of frame / slice / tile containing the current block of the current block.

[0202] Here, the types of frames / slices / tiles include the I type that does not reference other frames / slices / tiles, the P type that references frames / slices / tiles in one direction, and the B type that references frames / slices / tiles in at least one direction.

[0203] As an example of the second embodiment, when the current block is a luminance block, the template region described in Fig. 8a can be selected. Furthermore, by applying a 3x3 Sobel filter and a Laplacian filter to a local region of the template region, template features including gradient directions and magnitudes can be extracted. The average of the gradient directions extracted by each filter is calculated, and the magnitudes of the gradients are accumulated in a bin indicating the directional mode closest to the average direction. By repeating this process for other local regions, a histogram is generated. The intra prediction mode corresponding to the bin with the largest accumulated magnitude in the histogram is used to generate the final predicted block of the current block.

[0204] As another example of the second embodiment, when the current block is a chrominance block, the chrominance template (843) and the luminance template (853) described in FIG. 8C can be selected as the template region. By applying a 3x1 Sobel filter to a local region of the template region, template features including gradient direction and magnitude are extracted. The magnitudes of the gradients are accumulated in a bin indicating the directional mode closest to the extracted gradient direction. By repeating this process for other local regions, a histogram is generated. Directional prediction modes corresponding to five bins with large accumulated magnitudes in the histogram are selected. Prediction block candidates of the current block are generated using each of the five directional prediction modes and one non-directional prediction mode. A final prediction block of the current block is generated by weighting the prediction block candidates.

[0205] As another example of the second embodiment, when the current frame is a P-frame or a B-frame, the current template (813) described in FIG. 8A and the reference template (833) described in FIG. 8B can be selected. A 1x3 Sobel filter is applied to a local region of the selected template region, thereby extracting template features including gradient direction and magnitude. The magnitudes of the gradients are accumulated in a bin indicating the directional mode closest to the extracted gradient direction. By repeating this process for other local regions, a histogram is generated. Directional prediction modes corresponding to five bins with large accumulated magnitudes in the histogram are selected. Prediction block candidates for the current block are generated using each of the five directional prediction modes and one non-directional prediction mode. A final prediction block for the current block is generated by weighting the prediction block candidates.

[0206] Meanwhile, in the third embodiment, the prediction method to be applied to the current block is determined based on the matching cost. Among the template feature-based prediction method candidates, the one with the lowest matching cost is selected.

[0207] The matching cost is calculated using a predefined area related to the current block. The predefined area is an area within the template area of ​​the current block that has high spatial similarity to the current block.

[0208] The preset region may be a surrounding region adjacent to the current block, or an region adjacent to the current block within the template region of the current block. Referring to FIGS. 10A, 10B, and 10C, within the template region, the upper region of the current block, the left region of the current block, or both the upper region and the left region of the current block may be used as the preset region. Alternatively, a corresponding region of another channel or a corresponding template may be used as the preset region.

[0209] Since the preset region is included in the template region, it consists of restored reference samples. The restored reference samples within the preset region can be referred to as the restored values ​​of the preset region.

[0210] A predetermined region is predicted using template feature-based prediction method candidates. In other words, a template region for the predetermined region is determined, features of the template region are extracted, a histogram is generated based on the template features, and one or more intra prediction modes are derived from the histogram. Prediction block candidates for the predetermined region are generated using the derived intra prediction modes, and a final prediction block is generated by weighting the prediction block candidates. The final prediction block for the predetermined region is the predicted value of the predetermined region.

[0211] Based on the similarity (or matching cost) between the predicted result of a preset region using each template feature-based prediction method candidate and the restored value, one of the template feature-based prediction method candidates is selected. The similarity can be calculated using a loss function such as SAD (Sum of Absolute Differences), SATD (sum of absolute transformed differences), or MRSAD (mean-removed sum of absolute differences). The higher the similarity between the predicted value of the preset region and the restored value, the lower the matching cost. When the similarity between the predicted value of the preset region and the restored value is the highest, the template feature-based prediction method candidate used to derive the predicted value is selected as the template feature-based prediction method of the current block.

[0212] In the fourth embodiment, the prediction method to be applied to the current block is determined based on information about the surrounding blocks. The prediction method to be applied to the current block can be determined based on the statistics of the prediction methods used to predict the surrounding blocks of the current block.

[0213] Here, the neighboring blocks represent spatial / temporal neighboring blocks of the current block, or adjacent blocks of the current block, or reference blocks of adjacent blocks. The reference block of an adjacent block is an area indicated by the motion vector or block vector of the adjacent block, and refers to an area used to predict the adjacent block. If the adjacent blocks of the current block or the reference blocks of the adjacent blocks are reconstructed using a template feature-based prediction method, the prediction method of the current block can be determined based on the statistics of the template feature-based prediction method.

[0214] Alternatively, when the current block is a chrominance block, the surrounding blocks may include at least one of an adjacent area of ​​the current block, an area within a preset distance from the current block, an adjacent area of ​​a luminance block corresponding to the current block, or an area within a preset distance from the luminance block. The area within a preset distance includes an area spaced apart from the current block or the luminance block by a preset distance.

[0215] FIG. 11 is a diagram for explaining a method for determining a template feature-based prediction method based on information of a surrounding block according to one embodiment of the present disclosure.

[0216] Referring to Figure 11, the upper left block (peripheral block 1), the upper block (peripheral block 3), the left blocks, the lower left block (peripheral block 4), and the upper right block (peripheral block 5) are illustrated. The left blocks include three peripheral blocks, including peripheral block 2.

[0217] Surrounding blocks are predicted by the intra prediction modes: planar mode, 18-prediction mode, DC mode, temp_pred_1 mode, temp_pred_2 mode, and temp_pred_3 mode. temp_pred_x mode refers to a template feature-based prediction method with index x.

[0218] According to one embodiment of the present disclosure, the prediction method for the current block may be determined based on template feature-based prediction methods used to predict the surrounding blocks. If at least one of the surrounding blocks 1 to 4 is predicted using a template feature-based prediction method, the current block may also be predicted using the template feature-based prediction method.

[0219] As an example of the fourth embodiment, if a pre-designated pre-designated block among the pre-designated blocks 1 to 4 is predicted by a specific template feature-based prediction method, the current block can be predicted by the specific template feature-based prediction method. The pre-designated pre-designated pre-designated blocks can be designated arbitrarily or based on user settings.

[0220] As another example of the fourth embodiment, the prediction method for the current block can be determined based on the frequency of template feature-based prediction methods used to predict surrounding blocks. Among the template feature-based prediction methods, at least one template feature-based prediction method frequently used to predict surrounding blocks can be used as the prediction method for the current block.

[0221] The frequency of template feature-based prediction methods used to predict surrounding blocks can be calculated based on the number, area, or number of pixels of surrounding blocks.

[0222] Specifically, the prediction method of the current block may be determined based on the number of neighboring blocks to which each template feature-based prediction method is applied. The prediction method of the current block may be determined to be the same as the template feature-based prediction method applied to the largest number of neighboring blocks. If a specific template feature-based prediction method is used to predict three out of five neighboring blocks, the specific prediction method may be selected as the prediction method of the current block. In Fig. 11, since the first template feature-based prediction method temp_pred_1 is used to predict the two neighboring blocks with the largest number, the prediction mode of the current block may be temp_pred_1 mode.

[0223] Alternatively, the prediction method for the current block can be determined based on the area or pixel count of the surrounding blocks to which each template feature-based prediction method candidate is applied. The prediction method for the current block can be determined to be the same as the template feature-based prediction method applied to the largest number of pixels or the largest area. In Figure 11, since the area of ​​surrounding block 3 is the largest, temp_pred_2 can be determined as the prediction method for the current block.

[0224] Alternatively, the prediction method for the current block may be determined based on whether the current block and the surrounding blocks are the same or the similarity between the current block and the surrounding blocks. When there is a surrounding block among the surrounding blocks that has the same size and shape as the current block, the template feature-based prediction method used to predict the surrounding block may be used as the template feature-based prediction method for the current block. When there is a surrounding block among the surrounding blocks that has a similar size and shape as the current block, the template feature-based prediction method used to predict the surrounding block may be used as the template feature-based prediction method for the current block. In Fig. 11, since the size and shape of surrounding block 3 are most similar to those of the current block, temp_pred_2 used to predict surrounding block 3 may be determined as the prediction method for the current block.

[0225] Meanwhile, when the frequencies of multiple template feature-based prediction methods are the same, a predetermined priority can be used. The priority can include at least one of the following: the distance from a neighboring block to a specific location, the number of neighboring blocks applied, or the area of ​​the neighboring blocks applied.

[0226] When the frequency of multiple template feature-based prediction methods is the same and highest, among the template feature-based prediction methods, the template feature-based prediction method applied to a neighboring block close to a specific location can be used as the prediction method for the current block. For example, among the neighboring blocks to which the two most frequent and identical prediction methods are applied, the prediction method of the neighboring block closest to the neighboring block containing the upper left pixel of the current block (i.e., neighboring block 1) can be used as the prediction method for the current block.

[0227] Alternatively, if there are multiple template feature-based prediction methods with the largest number of applied surrounding blocks, among the multiple template feature-based prediction methods, a template feature-based prediction method with a larger area of ​​applied surrounding blocks can be used as a prediction method for the current block.

[0228] Alternatively, if there are multiple template feature-based prediction methods with the largest area of ​​applied surrounding blocks, among the multiple template feature-based prediction methods, a template feature-based prediction method with a larger number of applied surrounding blocks can be used as a prediction method for the current block.

[0229] Alternatively, if there are multiple surrounding blocks identical or similar to the current block, among the template feature-based prediction methods used to predict the multiple surrounding blocks, one of the methods can be used as a prediction method for the current block based on the distance between each surrounding block and a specific location, the area of ​​each surrounding block, and the number of surrounding blocks.

[0230] The third embodiment described above can be expanded or modified.

[0231] In extended or modified embodiments, the occurrence frequency of intra prediction modes is accumulated by considering not only neighboring blocks to which template feature-based prediction methods are applied, but also neighboring blocks to which general intra prediction modes are applied. For example, if a vertical intra prediction mode is applied to a neighboring block, the frequency of that vertical intra prediction mode increases.

[0232] Additionally, if a neighboring block is predicted using multiple intra prediction modes, all or some of the intra prediction modes used may be reflected in the frequency. For example, if neighboring block 1 is predicted using directional prediction modes 10 and 20, the frequency of each of directional prediction modes 10 and 20 increases by 1. As another example, if a template feature-based prediction method is applied to a neighboring block, the frequency of each of the intra prediction modes derived from the template feature-based prediction method increases.

[0233] Alternatively, when the current block is a chrominance block, all or part of the intra prediction modes used to predict at least one of an adjacent area of ​​the current block, an area within a preset distance from the current block, an adjacent area of ​​a luminance block corresponding to the current block, or an area within a preset distance from the luminance block may be reflected in the frequency of prediction methods of the surrounding blocks.

[0234] In this way, one or more intra prediction modes can be selected for application to the current block based on the occurrence frequency of the prediction modes of the surrounding blocks. In other words, after counting the occurrence frequency for each intra prediction mode, one or more intra prediction modes can be used as modes for predicting the current block based on the occurrence frequency.

[0235] Meanwhile, virtual intra prediction modes of certain blocks are considered only in inter-slices. Here, certain blocks refer to blocks predicted based on at least one of Matrix-based Intra Prediction (MIP), Intra Template Matching Prediction (IntraTMP), Intra Block Copy (IBC), or Extrapolation filter-based intra prediction (EIP).

[0236] In a prediction method based on template features of the current block, the weights of prediction block candidates can be determined based on the frequency of DDIMs.

[0237] FIG. 12 is a flowchart of an image decoding method according to one embodiment of the present disclosure.

[0238] Referring to FIG. 12, the image decoding device determines a prediction method of the current block based on at least one of prediction methods of surrounding blocks, information of the current block, matching cost of the restored area, or signaled information (S1210).

[0239] In one embodiment, an image decoding device determines a prediction method for a current block based on information signaled from an image encoding device. The signaled information includes an index indicating at least one of a plurality of prediction methods. Each index may correspond to a combination of a template region, a template feature extraction method, and a directional prediction mode derivation method for predicting the current block.

[0240] In one embodiment, the image decoding device determines that the prediction method of the current block is the same as the prediction method of the surrounding blocks at the specified location.

[0241] In one embodiment, the image decoding device selects a prediction method for the current block from among prediction methods of surrounding blocks based on the number or area of ​​surrounding blocks predicted by the same prediction method.

[0242] In one embodiment, the image decoding device selects a prediction method for the current block from prediction methods for surrounding blocks based on similarity in size or shape between the current block and surrounding blocks.

[0243] Here, the surrounding blocks refer to an area included in an adjacent area of ​​the current block, or an area included within a preset distance from the current block. Alternatively, the surrounding blocks may refer to an area spaced apart from the current block by a preset distance. Alternatively, when the current block is a chrominance block, the surrounding blocks may include at least one of an adjacent area of ​​the current block, an area within a preset distance from the current block, an adjacent area of ​​a luminance block corresponding to the current block, or an area within a preset distance from the luminance block. The area within a preset distance includes an area spaced apart from the current block or the luminance block by a preset distance.

[0244] In one embodiment, the video decoding device determines a prediction method for the current block based on information about the current block. The information about the current block includes at least one of a channel, width, height, size, aspect ratio, position, type of a frame containing the current block, or type of a slice containing the current block.

[0245] In one embodiment, the image decoding device determines a prediction method of the current block based on a matching cost of a previously restored region. The previously restored region may be a region included in a template region of the current block. Specifically, the image decoding device predicts the previously restored region from a surrounding region of the previously restored region using a plurality of prediction methods. The image decoding device calculates a matching cost representing a difference between a predicted value of the previously restored region by each prediction method and a restored value of the previously restored region. Alternatively, the image decoding device may calculate a similarity between the predicted value and the restored value. The image decoding device determines that the prediction method of the current block is the same as a prediction method that minimizes the matching cost or a prediction method that maximizes the similarity among the plurality of prediction methods.

[0246] Thereafter, the image decoding device generates a prediction block of the current block using the prediction method of the current block (S1220).

[0247] Specifically, the image decoding device determines a template region including previously restored reference samples.

[0248] In one embodiment, the template region may include at least one of an adjacent region of the current block or an region within a predetermined distance from the current block.

[0249] In one embodiment, the template region may include at least one of a reference block referenced by a motion vector or block vector of the current block, or a surrounding area of ​​the reference block.

[0250] In one embodiment, the template region may include at least one of an adjacent region of the current block, a luminance block corresponding to the current block, or an adjacent region of the luminance block when the current block is a chrominance block.

[0251] Thereafter, the image decoding device extracts template features including gradient directions and gradient magnitudes of the template region using a predetermined filter. The image decoding device extracts horizontal gradients and vertical gradients by applying the filter to each of the local regions within the template region. The image decoding device calculates the gradient directions and gradient magnitudes from the horizontal gradients and the vertical gradients. The image decoding device can obtain the gradient directions and magnitudes of the template region by aggregating the gradient directions and magnitudes of each local region.

[0252] Thereafter, the image decoding device derives one or more intra prediction modes of the current block based on the gradient directions and gradient magnitudes.

[0253] In one embodiment, the video decoding device generates histogram data by identifying a directional prediction mode having a closest angle to each gradient direction and accumulating gradient magnitudes in bins representing the identified directional prediction modes. In particular, the video decoding device may accumulate gradient magnitudes not only in the identified directional prediction mode but also in bins representing directional prediction modes adjacent to the identified directional prediction mode. The video decoding device may differentially accumulate gradient magnitudes in bins representing directional prediction modes adjacent to the identified directional prediction mode. The video decoding device accumulates a larger gradient magnitude in a directional prediction mode closer to the gradient direction than in a directional prediction mode farther from the gradient direction.

[0254] An image decoding device derives one or more intra prediction modes of a current block based on histogram data. The image decoding device selects one or more intra prediction modes of the current block based on the cumulative sizes of bins in the histogram data, the mean value, the mode value, the median value, the local maximum value, or the local minimum value of the cumulative sizes of the bins. In detail, the image decoding device may select a directional prediction mode corresponding to at least one of the mean value, the mode value, the median value, the local maximum value, or the local minimum value of the bin sizes. The image decoding device may further select directional prediction modes adjacent to the selected directional prediction mode. The selected directional prediction modes are included in the intra prediction modes of the current block.

[0255] The video decoding device can add a predetermined non-directional prediction mode to the directional prediction modes selected from the histogram. The non-directional prediction mode can indicate at least one of a planar mode, a planar vertical mode, a planar horizontal mode, a DC mode, or a block vector-based mode.

[0256] In one embodiment, an image decoding device derives one or more intra prediction modes of a current block using a pre-trained neural network model. The image decoding device obtains probability values ​​of intra prediction modes by applying the neural network model to information about the current block. The image decoding device determines the intra prediction modes of the current block based on at least one of an average value, a mode, a median value, a local maximum value, and a local minimum value of the probability values ​​of the intra prediction modes. The information about the current block includes features of a template region, a type of a filter used to extract features of the template region, coefficients of the filter, a size, a position, a shape, a channel of the current block, a size, a position, a shape, a channel of the template region, or a relationship between the current block and the template region.

[0257] Thereafter, the video decoding device generates prediction block candidates of the current block using one or more intra prediction modes of the current block, and generates a final prediction block of the current block by weighting the prediction block candidates.

[0258] When intra prediction modes are derived based on a histogram, the weight of each prediction block candidate is determined based on the accumulated gradient magnitude within the histogram. For example, the larger the accumulated gradient magnitude, the greater the weight of the corresponding prediction block candidate.

[0259] When intra prediction modes are derived based on a neural network model, the weights of each prediction block candidate are determined based on the probability values ​​output by the neural network model. For example, a high weight is applied to a prediction block candidate corresponding to an intra prediction mode with a high probability value.

[0260] FIG. 13 is a flowchart of an image encoding method according to one embodiment of the present disclosure.

[0261] Referring to FIG. 13, the image encoding device determines a prediction method to be applied to the current block among a plurality of prediction methods (S1310).

[0262] Here, the plurality of prediction methods include at least one of the first to third prediction methods. The plurality of prediction methods may include at least one of the first prediction method derived based on prediction methods of surrounding blocks, the second prediction method derived based on information about the current block, or the third prediction method derived based on the matching cost of the restored region.

[0263] In one embodiment, the first prediction method may be derived from prediction methods of neighboring blocks based on the number or area of ​​neighboring blocks predicted by the same prediction method. The video encoding device may determine whether the prediction method used to predict the largest number of neighboring blocks is applicable to the current block. Alternatively, the video encoding device may determine whether the prediction method used to predict the largest number of pixels is applicable to the current block.

[0264] In one embodiment, the first prediction method is derived from prediction methods of neighboring blocks based on the similarity in size or shape between the current block and the neighboring blocks. As an example, the image encoding device can determine whether the prediction method of a neighboring block having the same size or shape as the current block is applicable to the current block.

[0265] Here, the surrounding blocks may be included in the adjacent area of ​​the current block or may be included within a preset distance from the current block. Alternatively, when the current block is a chrominance block, the surrounding blocks may include at least one of the adjacent area of ​​the current block, the area within a preset distance from the current block, the adjacent area of ​​a luminance block corresponding to the current block, or the area within a preset distance from the luminance block. The area within a preset distance includes an area spaced apart from the current block or the luminance block by a preset distance.

[0266] In one embodiment, the second prediction method is a method derived based on information about the current block. The video encoding device can determine whether the second prediction method derived based on information about the current block is applied to the current block. Here, the information about the current block includes at least one of the channel, width, height, size, aspect ratio, position, type of frame including the current block, or type of slice including the current block.

[0267] In one embodiment, the third prediction method is a method derived from multiple prediction methods based on the results of predicting the restored region from the surrounding region of the restored region using the multiple prediction methods and the matching cost between the restored region. The image encoding device can determine whether the prediction method that minimizes the matching cost, which represents the difference between the predicted value and the restored value of the restored region, is applied to the current block.

[0268] The video encoding device generates a prediction block of the current block using a prediction method determined to be applied to the current block (S1320).

[0269] Step S1320 corresponds to step S1220 performed by the image decoding device. In other words, the first prediction method, the second prediction method, and the third prediction method include the operations of step S1220. The image encoding device can generate the final prediction block of the current block using a combination of a template region, a template feature extraction method, and a directional prediction mode derivation method according to the determined prediction method.

[0270] The image encoding device encodes information regarding a prediction method of the current block (S1330).

[0271] Information about the prediction method of the current block may include a flag indicating whether the prediction method determined in step 1310 is applied to the current block.

[0272] In one embodiment, the information about the prediction method of the current block may further include an index for identifying a combination of a template region, a template feature extraction method, and a directional prediction mode derivation method within the determined prediction method.

[0273] Although the flowchart / timing diagram of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of ​​one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart / timing diagram without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart / timing diagram is not limited to a chronological order.

[0274] It should be understood that the exemplary embodiments described above can be implemented in many different ways. The functions or methods described in one or more examples can be implemented in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as "units" to further emphasize their implementation independence.

[0275] Meanwhile, the various functions or methods described in this embodiment may be implemented as instructions stored on a non-transitory storage medium that can be read and executed by one or more processors. Non-transitory storage media include, for example, all types of storage devices that store data in a form readable by a computer system. For example, non-transitory storage media include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical drives, magnetic hard drives, and solid-state drives (SSDs).

[0276] The above description is merely an example of the technical idea of ​​the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of ​​the present embodiment, but rather to explain it, and the scope of the technical idea of ​​the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.

Claims

1. In a video decoding method performed by a video decoding device, A step of determining a prediction method of a current block based on at least one of prediction methods of surrounding blocks, information of a current block, matching cost of a restored area, or signaled information; and A step of generating a prediction block of the current block using the prediction method of the current block. A method for decrypting an image, comprising:

2. In paragraph 1, The step of determining the prediction method of the current block is: A step of selecting a prediction method of the current block from the prediction methods of the surrounding blocks based on the number or area of ​​the surrounding blocks predicted by the same prediction method. A method for decrypting an image, comprising:

3. In paragraph 1, The step of determining the prediction method of the current block is: A step of selecting a prediction method of the current block from the prediction methods of the surrounding blocks based on the similarity in size or shape between the current block and the surrounding blocks. A method for decrypting an image, comprising:

4. In paragraph 1, The above surrounding blocks are, An image decoding method that is included in an adjacent area of ​​the current block or within a preset distance from the current block.

5. In paragraph 1, The above surrounding blocks are, An image decoding method, wherein the image decoding method comprises at least one of an adjacent area of ​​the current block, an area within a preset distance from the current block, an adjacent area of ​​a luminance block corresponding to the current block, or an area within a preset distance from the luminance block when the current block is a chrominance block.

6. In paragraph 1, The information of the current block above is: A video decoding method comprising at least one of a channel, a width, a height, a size, an aspect ratio, a position of the current block, a type of a frame including the current block, or a type of a slice including the current block.

7. In paragraph 1, The step of determining the prediction method of the current block is: A step of determining a prediction method of the current block from among the plurality of prediction methods based on the matching cost between the restored region and the results of predicting the restored region from the surrounding region of the restored region using the plurality of prediction methods. A method for decrypting an image, comprising:

8. In paragraph 1, The above signaled information is, An image decoding method comprising an index pointing to one of a plurality of prediction methods.

9. In paragraph 1, The prediction method for the current block above is: A step of determining a template area including the restored reference samples; A step of calculating gradient directions and gradient sizes of the template area using a predetermined filter; A step of deriving one or more intra prediction modes of the current block based on the gradient directions and the gradient magnitudes; and A step of weighting prediction block candidates generated by one or more of the above intra prediction modes. A method for decrypting an image, comprising:

10. In paragraph 9, The above template area is, An image decoding method comprising at least one of an adjacent area of ​​the current block, an area within a predetermined distance from the current block, a reference block referenced by a motion vector or block vector of the current block, or a surrounding area of ​​the reference block.

11. In paragraph 9, The step of deriving one or more intra prediction modes of the current block is: A step of generating histogram data by accumulating gradient magnitudes corresponding to each gradient direction into bins indicating multiple directional prediction modes related to each gradient direction; and A step of determining one or more intra prediction modes of the current block based on the histogram data. A method for decrypting an image, comprising:

12. In paragraph 11, The steps for generating the above histogram data are: A step of differentially accumulating the corresponding gradient magnitudes in bins indicating the directional prediction mode closest to each gradient direction and the directional prediction modes adjacent to the closest directional prediction mode. A method for decrypting an image, comprising:

13. In paragraph 11, The step of determining one or more intra prediction modes of the current block is: A step of selecting one or more intra prediction modes of the current block based on the sizes of the bins, the mean value, the mode value, the median value, the local maximum value, or the local minimum value of the sizes of the bins. A method for decrypting an image, comprising:

14. In a video encoding method performed by a video encoding device, A step of determining a prediction method to be applied to a current block among a plurality of prediction methods, wherein the plurality of prediction methods include at least one of a first prediction method in which the prediction method of the current block is derived based on prediction methods of surrounding blocks, a second prediction method derived based on information of the current block, or a third prediction method derived based on a matching cost of a restored area; A step of generating a prediction block of the current block using a prediction method determined to be applied to the current block; and A step of encoding information about the prediction method of the current block. A method of encoding an image, comprising:

15. A method for transmitting data including a bitstream for an image, The above method, A step of obtaining a bitstream for the above image; and Comprising a step of transmitting data including the above bitstream, The above obtaining steps are: A step of determining a prediction method to be applied to a current block among a plurality of prediction methods, wherein the plurality of prediction methods include at least one of a first prediction method in which the prediction method of the current block is derived based on prediction methods of surrounding blocks, a second prediction method derived based on information of the current block, or a third prediction method derived based on a matching cost of a restored area; A step of generating a prediction block of the current block using a prediction method determined to be applied to the current block; and A step of encoding information about the prediction method of the current block. A method of encoding an image comprising steps.

Citation Information

Patent Citations

  • Chain power transmission mechanism and silent chain

    KR1020190118107A

  • Organoid seed preparing method and organoid seed preparing apparatus for preparing organoid

    KR1020240092558A

  • Measuring apparatus for roller of firing furnace

    KR1020250043878A

  • Body for electiric vehicle

    KR102359710B1

  • KR20200145780A

Cited By

  • Coding unit processing method and coding method

    CN121567858A