Fusion of multiple prediction signals based on template

By employing IntraTMP block vectors and weight matrices for template analysis, the method improves video encoding efficiency and image quality by splitting current templates into sub-templates and combining reference blocks for better prediction.

WO2025146969A1PCT designated stage expired Publication Date: 2025-07-10HYUNDAI MOTOR CO LTD +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019966
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-05
Filing Date
2024-12-06
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing video compression technologies struggle with increasing data sizes due to higher image resolutions and frame rates, requiring improved encoding efficiency and image quality.

Method used

The method involves determining IntraTMP block vectors and weight matrices for template analysis, splitting current templates into sub-templates for enhanced prediction, and applying weighted sums to combine reference blocks for improved prediction accuracy.

Benefits of technology

Enhances coding efficiency by discovering more diverse prediction candidates and improving image quality through better template matching techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019966_10072025_PF_FP_ABST
    Figure KR2024019966_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the fusion of multiple prediction signals derived on the basis of template analysis or template matching. The present disclosure describes technologies for improving the functionality of template analysis or template matching in video coding, thus improving the efficiency of the coding. A template of the current block may be divided into two or more sub-templates, and each of the sub-templates may be used for template analysis or template matching. By using the sub-templates, more different candidates or predictors can be found and fusion can be applied to combine the candidates or predictors for better prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Fusion of multiple template-based prediction signals

[0001] The present disclosure relates to encoding / decoding of video data, and more particularly, to a method for predicting a current block using coding tools based on template analysis or template matching.

[0002] The content described below merely provides background information related to the present invention and does not constitute prior art.

[0003] Since video data has a large amount of data compared to voice data or still image data, it requires a lot of hardware resources, including memory, to store or transmit it without processing for compression.

[0004] Therefore, when storing or transmitting video data, the encoder compresses the video data and stores or transmits it, and the decoder receives the compressed video data, decompresses it, and plays it back. These video compression technologies include H.264 / AVC, HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), which improves encoding efficiency by about 30% compared to HEVC.

[0005] However, as the size, resolution, and frame rate of images are gradually increasing, and the amount of data that needs to be encoded is also increasing, a new compression technology that has better encoding efficiency and better image quality improvement than existing compression technologies is required.

[0006] The present disclosure aims to improve the functionality of template analysis or template matching in video coding and thereby improve coding efficiency.

[0007] One aspect of the present disclosure provides a method for encoding or decoding video data. The method includes the steps of determining a first IntraTMP block vector for a current block, determining a first reference block identified by the first IntraTMP block vector, determining a second IntraTMP block vector for the current block, and determining a second reference block identified by the second IntraTMP block vector. The method further includes the steps of determining a first weight matrix defining weights for each sample position for the first reference block, and determining a second weight matrix defining weights for each sample position for the second reference block. The method further includes the steps of generating a prediction block by fusing sample values ​​of the first reference block and sample values ​​of the second reference block based on the first weight matrix and the second weight matrix, and encoding or decoding the current block based on the prediction block.

[0008] In some embodiments, the first weight matrix and the second weight matrix may be determined based on a mean squared error (MSE) or a sum of absolute difference (SAD) between predicted values ​​of a current template, which is a set of reconstructed samples adjacent to the current block, and reconstructed values ​​of the current template. Here, the predicted values ​​of the current template may be derived as a weighted sum between a first reference template, which is a set of reconstructed samples adjacent to the first reference block, and a second reference template, which is a set of reconstructed samples adjacent to the second reference block.

[0009] In some embodiments, the current template may be split into two sub-templates based on block splitting information of already decrypted neighboring blocks of the current block. The first IntraTMP block vector may be determined based on a template matching cost using a first sub-template among the two sub-templates, and the second IntraTMP block vector may be determined based on a template matching cost using a second sub-template among the two sub-templates.

[0010] In some embodiments, the first IntraTMP block vector and the first IntraTMP block vector may be used to construct a candidate list of IBC block vectors of another block to be encoded subsequent to the current block.

[0011] In some embodiments, the method may further comprise encoding or decoding a candidate index into the bitstream, the candidate index indicating a candidate used for the current block within a candidate list, which is a set of a plurality of available candidates. The candidate index may identify the first IntraTMP block vector and the first IntraTMP block vector.

[0012] In some embodiments, the current block may be divided into two regions by a straight line extending from the partition boundary between the sub-templates. In such a case, the weights in the first weight matrix and the second weight matrix may be derived from a ramp function based on the distance from each sample location to the partition boundary between the two regions within the current block.

[0013] In some embodiments, in a blending region around a segmentation boundary between the two regions within the current block, a weighted sum of sample values ​​of the first reference block and sample values ​​of the second reference block may be used as prediction values ​​of the current block, and in a region other than the blending region, sample values ​​of the first reference block or sample values ​​of the second reference block may be used as prediction values ​​of the current block. The width of the blending region may be adaptively determined according to a size of the current block.

[0014] Another aspect of the present disclosure is a method of providing video data to a video decoding device, comprising the steps of encoding the video data into a bitstream and transmitting the bitstream to the video decoding device. In the step of encoding the video data into a bitstream, the method of encoding the video data described above may be used.

[0015] According to some embodiments of the present disclosure, the weight matrices used in the weighted sum of predictors obtained based on IntraTMP block vectors can be implicitly derived or determined using the current template and reference templates (determined based on template matching using each of the subtemplates of the current template).

[0016] According to some embodiments of the present disclosure, by using sub-templates split from the current template, more different candidates or predictors can be discovered and fusion can be applied to combine them for better prediction.

[0017] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure.

[0018] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT (QuadTree plus BinaryTree TernaryTree) structure.

[0019] FIGS. 3A and 3B are diagrams illustrating multiple intra prediction modes, including wide-angle intra prediction modes.

[0020] Figure 4 is an example diagram of the surrounding blocks of the current block.

[0021] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure.

[0022] Figure 6 is a conceptual diagram illustrating a search area used in IntraTMP (Intra Template Matching Prediction) technology.

[0023] Figure 7 is a conceptual diagram illustrating exemplary templates of the current block that can be used for template matching.

[0024] FIG. 8 is a conceptual diagram illustrating a method for deriving one or more intra prediction modes for a current block based on template analysis in DIMD (Decoder-side Intra Mode Deivation) technology.

[0025] FIG. 9 is a conceptual diagram illustrating a method for deriving one or more intra prediction modes for a current block based on template prediction in TIMD (Template-based Intra Mode Deivation) technology.

[0026] Figure 10 is a conceptual diagram illustrating a method for signaling a combination of a partition mode and an intra prediction mode of SGPM (Spatial Geometric Partitioning Mode).

[0027] Figure 11 is a conceptual diagram illustrating the template of the current block to which the SGPM mode is applied and the weights extended to the template area.

[0028] Figure 12 is a conceptual diagram for explaining the blending area of ​​SGPM and the blending weights applied thereto.

[0029] Figure 13 is a conceptual diagram showing examples of sub-templates divided based on the block boundaries of surrounding blocks that contain the template area of ​​the current block.

[0030] Figure 14 is a conceptual diagram showing another example of sub-templates divided based on the block boundaries of surrounding blocks containing the template of the current block.

[0031] Figures 15a to 15d are conceptual diagrams illustrating a method for deriving TIMD modes and constructing a candidate list when a current template is divided into sub-templates.

[0032] Figures 16a and 16b are conceptual diagrams illustrating how to derive DIMD modes when the current template is divided into sub-templates.

[0033] Figures 17a to 17c are conceptual diagrams for explaining an IntraTMP method using subtemplates according to one embodiment of the present invention.

[0034] FIG. 18a is a conceptual diagram illustrating an intra-template matching prediction (IntraTMP) method using sub-templates according to another embodiment of the present invention.

[0035] Figure 18b is a conceptual diagram illustrating a method for predicting a current template through a weighted sum of reference templates.

[0036] FIG. 19 is a conceptual diagram illustrating a method for deriving a weighted sum when the IntraTMP mode is applied to each of two parts of a current block in the SGPM mode according to one embodiment of the present invention.

[0037] Hereinafter, embodiments of the present invention will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, in describing the present embodiments, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present embodiments.

[0038] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image encoding device and its subcomponents will be described with reference to the illustration in FIG. 1.

[0039] The video encoding device may be configured to include a picture segmentation unit (110), a prediction unit (120), a subtractor (130), a transformation unit (140), a quantization unit (145), a reordering unit (150), an entropy encoding unit (155), an inverse quantization unit (160), an inverse transformation unit (165), an adder (170), a loop filter unit (180), and a memory (190).

[0040] Each component of the video encoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0041] A single image (video) is composed of one or more sequences containing multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile or slice is divided into one or more Coding Tree Units (CTUs). Each CTU is then divided into one or more Coding Units (CUs) by a tree structure. Information applied to each CU is encoded as the syntax of the CU, and information commonly applied to CUs included in a CTU is encoded as the syntax of the CTU. In addition, information commonly applied to all blocks within a single slice is encoded as the syntax of the slice header, and information applied to all blocks constituting one or more pictures is encoded in the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referenced by multiple pictures is encoded in a Sequence Parameter Set (SPS). And, information commonly referenced by one or more SPS is encoded in a Video Parameter Set (VPS). In addition, information commonly applied to one tile or tile group may be encoded as syntax of a tile or tile group header. Syntaxes included in an SPS, PPS, slice header, tile or tile group header may be referred to as high level syntax.

[0042] The picture segmentation unit (110) determines the size of the CTU. Information about the size of the CTU (CTU size) is encoded as the syntax of SPS or PPS and transmitted to the image decoding device.

[0043] The picture segmentation unit (110) divides each picture constituting an image into a plurality of CTUs having a predetermined size, and then recursively divides the CTUs using a tree structure. A leaf node in the tree structure becomes a CU, which is a basic unit of encoding.

[0044] The tree structure may be a QuadTree (QT) in which an upper node (or parent node) is divided into four lower nodes (or child nodes) of the same size, a BinaryTree (BT) in which an upper node is divided into two lower nodes, or a TernaryTree (TT) in which an upper node is divided into three lower nodes in a 1:2:1 ratio, or a structure that mixes two or more of the QT structures, BT structures, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure may be used, or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure may be used. Here, BTTT may be combined and referred to as a MTT (Multiple-Type Tree).

[0045] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT structure.

[0046] As illustrated in FIG. 2, a CTU may first be split into a QT structure. The quadtree splitting may be repeated until the size of the splitting block reaches the minimum block size (MinQTSize) of the leaf node allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoding unit (155) and signaled to the image decoding device. If the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, it may be further split into one or more of the BT structure or the TT structure. There may be multiple splitting directions in the BT structure and / or the TT structure. For example, there may be two directions in which the block of the corresponding node is split horizontally and two directions in which the block is split vertically. As illustrated in FIG. 2, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether nodes have been split, and if splitting has occurred, a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (Binary or Ternary) are encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0047] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes of a lower layer, a CU split flag (split_cu_flag) indicating whether the node is split may be encoded. If the CU split flag (split_cu_flag) value indicates that the node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU (coding unit), which is a basic unit of encoding. If the CU split flag (split_cu_flag) value indicates that the node is split, the video encoding device starts encoding from the first flag in the above-described manner.

[0048] As another example of a tree structure, when QTBT is used, there may be two types: a type that horizontally splits the block of the corresponding node into two blocks of the same size (i.e., symmetric horizontal splitting) and a type that vertically splits it (i.e., symmetric vertical splitting). A split flag (split_flag) indicating whether each node of the BT structure is split into blocks of a lower layer and split type information indicating the type of split are encoded by the entropy encoding unit (155) and transmitted to the image decoding device. Meanwhile, there may additionally be a type that splits the block of the corresponding node into two blocks of an asymmetrical shape. The asymmetric shape may include a shape that splits the block of the corresponding node into two rectangular blocks with a size ratio of 1:3, or a shape that splits the block of the corresponding node in a diagonal direction.

[0049] A CU can have various sizes depending on the QTBT or QTBTTT partitioning from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block." Depending on the QTBTTT partitioning employed, the current block may be rectangular as well as square.

[0050] The prediction unit (120) predicts the current block and generates a prediction block. The prediction unit (120) includes an intra prediction unit (122) and an inter prediction unit (124).

[0051] In general, each current block within a picture can be predictively coded. Prediction of the current block can typically be performed using either intra-prediction (using data from the picture containing the current block) or inter-prediction (using data from a picture coded before the picture containing the current block). Inter-prediction encompasses both unidirectional and bidirectional prediction.

[0052] The intra prediction unit (122) predicts pixels within the current block using pixels (reference pixels) located around the current block within the current picture including the current block. There are multiple intra prediction modes depending on the prediction direction. For example, as shown in Fig. 3a, the multiple intra prediction modes may include two non-directional modes including the Planar mode and the DC mode, and 65 directional modes. The surrounding pixels to be used and the calculation formula are defined differently depending on each prediction mode.

[0053] For efficient directional prediction for a rectangular current block, directional modes (intra prediction modes 67 to 80 and -1 to -14) indicated by dotted arrows in Fig. 3b may be additionally used. These may be referred to as "wide-angle intra-prediction modes." In Fig. 3b, the arrows point to corresponding reference samples used for prediction, and do not indicate the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. Wide-angle intra-prediction modes are modes that perform prediction in the opposite direction of a specific directional mode without additional bit transmission when the current block is rectangular. At this time, among the wide-angle intra-prediction modes, some wide-angle intra-prediction modes available for the current block may be determined based on the ratio of the width and height of the rectangular current block. For example, wide-angle intra prediction modes (intra prediction modes 67 to 80) having an angle less than 45 degrees are available when the current block is a rectangular shape whose height is smaller than its width, and wide-angle intra prediction modes (intra prediction modes -1 to -14) having an angle greater than -135 degrees are available when the current block is a rectangular shape whose width is larger than its height.

[0054] The intra prediction unit (122) can determine an intra prediction mode to be used to encode the current block. In some examples, the intra prediction unit (122) can encode the current block using multiple intra prediction modes and select an appropriate intra prediction mode to be used from the tested modes. For example, the intra prediction unit (122) can calculate bit-rate distortion values ​​using rate-distortion analysis for multiple tested intra prediction modes and select an intra prediction mode with the best bit-rate distortion characteristics among the tested modes.

[0055] The intra prediction unit (122) selects one intra prediction mode from among multiple intra prediction modes and predicts the current block using surrounding pixels (reference pixels) and an operation formula determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0056] The inter prediction unit (124) generates a prediction block for the current block using a motion compensation process. The inter prediction unit (124) searches for a block most similar to the current block within reference pictures that were encoded and decoded before the current picture, and generates a prediction block for the current block using the searched block. Then, a motion vector (MV) corresponding to the displacement between the current block within the current picture and the prediction block within the reference picture is generated. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The motion information including information on the reference picture used to predict the current block and information on the motion vector is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0057] The inter prediction unit (124) may perform interpolation on a reference picture or a reference block to improve prediction accuracy. That is, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including the two integer samples. When a process of searching for a block most similar to the current block is performed on the interpolated reference picture, the motion vector can be expressed up to a precision in decimal units rather than a precision in integer sample units. The precision or resolution of the motion vector can be set differently for each target region to be encoded, such as a slice, tile, CTU, CU, etc. When such adaptive motion vector resolution (AMVR) is applied, information on the motion vector resolution to be applied to each target region must be signaled for each target region. For example, when the target region is a CU, information on the motion vector resolution applied to each CU is signaled. Information on the motion vector resolution may be information indicating the precision of a differential motion vector, which will be described later.

[0058] Meanwhile, the inter prediction unit (124) can perform inter prediction using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors indicating the block position most similar to the current block within each reference picture are used. The inter prediction unit (124) selects a first reference picture and a second reference picture from reference picture list 0 (referred to as RefPicList0 or simply L0) and reference picture list 1 (referred to as RefPicList1 or simply L1), respectively, and searches for a block similar to the current block within each reference picture to generate a first reference block and a second reference block. Then, the first reference block and the second reference block are averaged or weighted averaged to generate a prediction block for the current block. Then, motion information including information on two reference pictures used to predict the current block and information on two motion vectors is transmitted to the entropy encoding unit (155). Here, reference picture list 0 may be composed of pictures that are before the current picture in display order among the restored pictures, and reference picture list 1 may be composed of pictures that are after the current picture in display order among the restored pictures. However, this is not necessarily limited to this, and restored pictures that are after the current picture in display order may be additionally included in reference picture list 0, and conversely, restored pictures that are before the current picture may be additionally included in reference picture list 1.

[0059] Various methods can be used to minimize the number of bits required to encode motion information.

[0060] For example, if the reference picture and motion vector of the current block are identical to those of a neighboring block, the motion information of the current block can be transmitted to the image decoding device by encoding information that can identify the neighboring block. This method is called 'merge mode.'

[0061] In merge mode, the inter prediction unit (124) selects a predetermined number of merge candidate blocks (hereinafter referred to as 'merge candidates') from the surrounding blocks of the current block.

[0062] As the surrounding blocks for deriving merge candidates, all or part of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (B2) adjacent to the current block within the current picture may be used, as illustrated in FIG. 4. In addition, a block located within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as a merge candidate. For example, a block co-located with the current block within the reference picture or blocks adjacent to the block at the co-located block may be additionally used as a merge candidate. If the number of merge candidates selected by the method described above is less than a preset number, a 0 vector is added to the merge candidates.

[0063] The inter prediction unit (124) uses these surrounding blocks to construct a merge list containing a predetermined number of merge candidates. Among the merge candidates included in the merge list, the merge candidate to be used as motion information of the current block is selected and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoding unit (155) and transmitted to the video decoding device.

[0064] Merge Skip mode is a special case of merge mode. After quantization, when all transform coefficients for entropy encoding are close to zero, only neighboring block selection information is transmitted without transmitting residual signals. By utilizing merge skip mode, relatively high encoding efficiency can be achieved for low-motion images, still images, and screen content images.

[0065] Hereinafter, merge mode and merge skip mode are collectively referred to as merge / skip mode.

[0066] Another method for encoding motion information is Advanced Motion Vector Prediction (AMVP) mode.

[0067] In AMVP mode, the inter prediction unit (124) derives predicted motion vector candidates for the motion vector of the current block using neighboring blocks of the current block. As neighboring blocks used to derive predicted motion vector candidates, all or some of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (B2) adjacent to the current block in the current picture as shown in FIG. 4 may be used. In addition, a block located in a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as the neighboring block used to derive predicted motion vector candidates. For example, a block co-located with the current block in the reference picture or blocks adjacent to the block in the co-located block may be used. If the number of motion vector candidates is less than a preset number by the method described above, a 0 vector is added to the motion vector candidates.

[0068] The inter prediction unit (124) derives predicted motion vector candidates using the motion vectors of these surrounding blocks, and determines a predicted motion vector for the motion vector of the current block using the predicted motion vector candidates. Then, the predicted motion vector is subtracted from the motion vector of the current block to produce a differential motion vector.

[0069] The predicted motion vector can be obtained by applying a predefined function (e.g., median, mean, etc.) to the predicted motion vector candidates. In this case, the image decoding device also knows the predefined function. In addition, since the surrounding blocks used to derive the predicted motion vector candidates are blocks that have already been encoded and decoded, the image decoding device also already knows the motion vectors of the surrounding blocks. Therefore, the image encoding device does not need to encode information to identify the predicted motion vector candidates. Therefore, in this case, information about the differential motion vector and information about the reference picture used to predict the current block are encoded.

[0070] Alternatively, the predicted motion vector can be determined by selecting one of the predicted motion vector candidates. In this case, information for identifying the selected predicted motion vector candidate is additionally encoded, along with information about the differential motion vector and the reference picture used to predict the current block.

[0071] The subtractor (130) subtracts the prediction block generated by the intra prediction unit (122) or inter prediction unit (124) from the current block to generate a residual block.

[0072] The transformation unit (140) transforms residual signals within a residual block having pixel values ​​in a spatial domain into transform coefficients in a frequency domain. The transformation unit (140) may transform the residual signals within the residual block using the entire size of the residual block as a transformation unit, or may divide the residual block into a plurality of sub-blocks and use the sub-blocks as transformation units to perform the transformation. Alternatively, the residual signals may be transformed using only the transformation domain sub-block as a transformation unit by dividing the sub-blocks into two sub-blocks, that is, a transformation domain and a non-transform domain. Here, the transformation domain sub-block may be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicating that only a sub-block has been converted, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit (155) and signaled to the image decoding device. In addition, the size of the conversion area sub-block may have a size ratio of 1:3 with respect to the horizontal axis (or vertical axis), and in this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding division is additionally encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0073] Meanwhile, the transformation unit (140) can individually perform transformations on the residual block in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a Multiple Transform Set (MTS). The transformation unit (140) can select one transformation function pair with the best transformation efficiency among the MTS and transform the residual block in the horizontal and vertical directions, respectively. Information (mts_idx) on the transformation function pair selected among the MTS is encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0074] The quantization unit (145) quantizes the transform coefficients output from the transform unit (140) using quantization parameters and outputs the quantized transform coefficients to the entropy encoding unit (155). The quantization unit (145) may directly quantize a related residual block without transformation for a certain block or frame. The quantization unit (145) may also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients within the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions may be encoded and signaled to an image decoding device.

[0075] The rearrangement unit (150) can perform rearrangement of coefficient values ​​for quantized residual values.

[0076] The reordering unit (150) can change a two-dimensional coefficient array into a one-dimensional coefficient sequence by using coefficient scanning. For example, the reordering unit (150) can output a one-dimensional coefficient sequence by scanning from the DC coefficient to the coefficients of the high-frequency region by using a zig-zag scan or a diagonal scan. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficients in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, the scanning method to be used may be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.

[0077] The entropy encoding unit (155) generates a bitstream by encoding a sequence of one-dimensional quantized transform coefficients output from the rearrangement unit (150) using various encoding methods such as CABAC (Context-based Adaptive Binary Arithmetic Code) and Exponential Golomb.

[0078] In addition, the entropy encoding unit (155) encodes information related to block division, such as CTU size, CU division flag, QT division flag, MTT division type, and MTT division direction, so that the image decoding device can divide the block in the same manner as the image encoding device. In addition, the entropy encoding unit (155) encodes information about a prediction type indicating whether the current block is encoded by intra prediction or inter prediction, and encodes intra prediction information (i.e., information about an intra prediction mode) or inter prediction information (information about an encoding mode of motion information (merge mode or AMVP mode), a merge index in the case of a merge mode, and a reference picture index and a differential motion vector in the case of an AMVP mode) according to the prediction type. In addition, the entropy encoding unit (155) encodes information related to quantization, that is, information about a quantization parameter and information about a quantization matrix.

[0079] The inverse quantization unit (160) inversely quantizes the quantized transform coefficients output from the quantization unit (145) to generate transform coefficients. The inverse transform unit (165) transforms the transform coefficients output from the inverse quantization unit (160) from the frequency domain to the spatial domain to restore the residual block.

[0080] An adder (170) adds the restored residual block and the predicted block generated by the prediction unit (120) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting the next block.

[0081] The loop filter unit (180) performs filtering on restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transformation / quantization. The loop filter unit (180) may include all or part of a deblocking filter (182), a sample adaptive offset (SAO) filter (184), and an adaptive loop filter (ALF, 186) as an in-loop filter.

[0082] The deblocking filter (182) filters the boundaries between restored blocks to remove blocking artifacts caused by block-based encoding / decoding, and the SAO filter (184) and the ALF (186) perform additional filtering on the deblocking-filtered image. The SAO filter (184) and the ALF (186) are filters used to compensate for the differences between restored pixels and original pixels caused by lossy coding. The SAO filter (184) improves not only subjective image quality but also encoding efficiency by applying an offset in units of CTUs. In contrast, the ALF (186) performs block-based filtering, and compensates for distortion by applying different filters by distinguishing the edges and degrees of variation of the corresponding block. Information on filter coefficients to be used in the ALF can be encoded and signaled to an image decoding device.

[0083] The restored blocks filtered through the deblocking filter (182), SAO filter (184), and ALF (186) are stored in the memory (190). When all blocks within a picture are restored, the restored picture can be used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0084] The video encoding device can store the bitstream of encoded video data on a non-transitory storage medium or transmit it to the video decoding device using a communication network.

[0085] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image decoding device and its subcomponents will be described with reference to FIG. 5.

[0086] The video decoding device may be configured to include an entropy decoding unit (510), a rearrangement unit (515), an inverse quantization unit (520), an inverse transformation unit (530), a prediction unit (540), an adder (550), a loop filter unit (560), and a memory (570).

[0087] Similar to the video encoding device of FIG. 1, each component of the video decoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0088] The entropy decoding unit (510) decodes the bitstream generated by the image encoding device to extract information related to block division, thereby determining the current block to be decoded, and extracts prediction information, information on residual signals, etc. required to restore the current block.

[0089] The entropy decoding unit (510) extracts information about the CTU size from the Sequence Parameter Set (SPS) or the Picture Parameter Set (PPS), determines the size of the CTU, and divides the picture into CTUs of the determined size. Then, the CTU is determined as the top layer of the tree structure, i.e., the root node, and the CTU is divided using the tree structure by extracting division information about the CTU.

[0090] For example, when splitting a CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, the second flag (mtt_split_flag) related to the splitting of MTT and the split direction (vertical / horizontal) and / or split type (binary / ternary) information are extracted, and the corresponding leaf nodes are split into the MTT structure. Accordingly, each node below the leaf nodes of QT are split recursively into the BT or TT structure.

[0091] As another example, when splitting a CTU using the QTBTTT structure, the CU split flag (split_cu_flag) indicating whether the CU is split is first extracted, and if the block is split, the first flag (QT_split_flag) may be extracted. During the splitting process, each node may undergo zero or more repeated QT splits followed by zero or more repeated MTT splits. For example, a CTU may undergo an MTT split right away, or conversely, may undergo only multiple QT splits.

[0092] As another example, when splitting a CTU using the QTBT structure, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Furthermore, for nodes corresponding to leaf nodes of QT, a split flag (split_flag) indicating whether to further split into BTs and splitting direction information are extracted.

[0093] Meanwhile, when the entropy decoding unit (510) determines the current block to be decoded by using the division of the tree structure, it extracts information on the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-prediction, the entropy decoding unit (510) extracts syntax elements for intra-prediction information (intra-prediction mode) of the current block. If the prediction type information indicates inter-prediction, the entropy decoding unit (510) extracts syntax elements for inter-prediction information, i.e., information indicating a motion vector and a reference picture referenced by the motion vector.

[0094] Additionally, the entropy decoding unit (510) extracts information about the quantized transform coefficients of the current block as information related to quantization and information about residual signals.

[0095] The rearrangement unit (515) can change the sequence of one-dimensional quantized transform coefficients entropy-decoded in the entropy decoding unit (510) back into a two-dimensional coefficient array (i.e., block) in the reverse order of the coefficient scanning performed by the image encoding device.

[0096] The inverse quantization unit (520) inversely quantizes the quantized transform coefficients and inversely quantizes the quantized transform coefficients using the quantization parameters. The inverse quantization unit (520) may also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in two dimensions. The inverse quantization unit (520) may perform inverse quantization by applying a matrix of quantized coefficients (scaling values) from an image encoding device to a two-dimensional array of quantized transform coefficients.

[0097] The inverse transform unit (530) inversely transforms the inverse quantized transform coefficients from the frequency domain to the spatial domain to restore residual signals, thereby generating a residual block for the current block.

[0098] In addition, when the inverse transform unit (530) inversely transforms only a portion of a transform block (sub-block), it extracts a flag (cu_sbt_flag) indicating that only a sub-block of the transform block has been transformed, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block, and inversely transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to restore residual signals, and fills “0” values ​​with residual signals for areas that have not been inversely transformed, thereby generating a final residual block for the current block.

[0099] In addition, when MTS is applied, the inverse transform unit (530) determines a transform function or a transform matrix to be applied in the horizontal and vertical directions using MTS information (mts_idx) signaled from the image encoding device, and performs inverse transform on the transform coefficients within the transform block in the horizontal and vertical directions using the determined transform function.

[0100] The prediction unit (540) may include an intra prediction unit (542) and an inter prediction unit (544). The intra prediction unit (542) is activated when the prediction type of the current block is intra prediction, and the inter prediction unit (544) is activated when the prediction type of the current block is inter prediction.

[0101] The intra prediction unit (542) determines the intra prediction mode of the current block among a plurality of intra prediction modes from the syntax elements for the intra prediction mode extracted from the entropy decoding unit (510), and predicts the current block using reference pixels around the current block according to the intra prediction mode.

[0102] The inter prediction unit (544) uses the syntax elements for the inter prediction mode extracted from the entropy decoding unit (510) to determine the motion vector of the current block and the reference picture referenced by the motion vector, and predicts the current block using the motion vector and the reference picture.

[0103] An adder (550) adds the residual block output from the inverse transform unit (530) and the predicted block output from the inter prediction unit (544) or the intra prediction unit (542) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting a block to be decoded later.

[0104] The loop filter unit (560) may include a deblocking filter (562), an SAO filter (564), and an ALF (566) as in-loop filters. The deblocking filter (562) deblocks the boundaries between restored blocks to remove blocking artifacts caused by block-by-block decoding. The SAO filter (564) and the ALF (566) perform additional filtering on restored blocks after deblocking filtering to compensate for differences between restored pixels and original pixels caused by lossy coding. The filter coefficients of the ALF are determined using information about filter coefficients decoded from the non-stream.

[0105] The restored blocks filtered through the deblocking filter (562), SAO filter (564), and ALF (566) are stored in the memory (570). When all blocks within a picture are restored, the restored picture is used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0106] Hereinafter, improved coding techniques that can be performed by the aforementioned video encoder (e.g., the video encoding device illustrated in FIG. 1) or video decoder (e.g., the video decoding device illustrated in FIG. 5) are disclosed. In particular, the techniques of the present disclosure relate to coding tools that utilize template analysis or template matching (TM). The video encoder and video decoder can encode or decode the current block based on the template analysis and template matching costs.

[0107] A video encoder and a video decoder can analyze a current template (e.g., a set of samples adjacent to a current block within a current picture) to determine one or more intra prediction modes to use in generating a prediction block for the current block.

[0108] In template matching, a video encoder and a video decoder can compare multiple reference templates (e.g., a set of samples from another picture or a set of samples from other blocks within the current picture). The multiple reference templates may be within a search range and may have a specific shape. Based on the comparison between the templates, the video encoder and the video decoder can perform one or more operations to improve video coding efficiency. For example, the video encoder and the video decoder can update an initial motion vector, determine or update prediction samples, and reorder merge lists or intra prediction modes.

[0109] There may be various coding techniques that utilize template analysis or template matching. For example, intra-prediction coding techniques include IntraTMP (Intra template matching prediction), DIMD (Decoder Side Intra Mode Derivation), TIMD (Template-based intra mode derivation), and SGPM (Spatial Geometric partitioning mode). Inter-prediction coding techniques include Inter template matching, ARMC-TM (Adaptive reordering of merge candidates with template matching), and GPM TM (Geometric partitioning mode Template Matching). Video encoders and decoders can use one of these coding techniques to encode or decode the current block.

[0110] Below, several intra-predictive encoding techniques using template analysis or template matching are described.

[0111] Intra-template matching prediction (IntraTMP) is a special intra-prediction method that copies the best predicted block based on template matching from a reconstructed portion of the current frame. For a predefined search region, the video encoder searches for a reference template most similar to the current template in the reconstructed portion of the current frame, and uses the block corresponding to the reference template as the prediction block for the current block. The sum of absolute differences (SAD) can be used as a cost function. Within the search region, the video encoder searches for the reference template with the smallest SAD. The video encoder then signals the use of this IntraTMP mode, and the video decoder performs the same prediction operation.

[0112] Figure 6 is a conceptual diagram illustrating a search area used in IntraTMP technology.

[0113] Referring to the example in Fig. 6, there are six predefined search areas, R1 to R6, which include reconstructed samples from the top and left CTUs and portions of reconstructed samples within the current CTU located at the top, left, bottom left, and top right of the current block. Each area can be determined based on the height (H) and width (W) of the current block. The search order is R4, R5, R6, R1, R3, and R2, and the search is performed in reverse z-scan order from the bottom right to the top left of each area.

[0114] In the current version of the Enhanced Compression Model (ECM), the IntraTMP technique includes several sub-modes, including single predictor mode, fusion of multiple predictors mode, sub-pel precision mode, and linear filter model mode.

[0115] The template search process can proceed in two stages. In the first stage, the video encoder and video decoder perform a search at three-pixel intervals to include block vectors pointing to 30 similar templates in a candidate list. In the second stage, the video encoder and video decoder perform a fine-tuned search at one-pixel intervals in a surrounding 3×3 region among the 30 candidate block vectors in the candidate list, ultimately selecting 19 candidate block vectors.

[0116] In single-predictor mode, an index indicating the optimal candidate selected through a rate-distortion test among multiple candidates in a candidate list can be signaled in the bitstream. In multi-predictor mode, a fusion method that weights and combines reference blocks corresponding to multiple reference templates can be used. For example, in multi-predictor mode, a prediction block of the current block can be generated through a weighted sum of two reference blocks corresponding to the lowest template matching cost and the second lowest template matching cost. Here, the fusion weights can be derived from template matching costs, derived using a Gaussian solver method, or can be fixed values. The number of reference blocks to be fused can be determined by template matching costs and threshold values ​​corresponding to the reference blocks, and a maximum of five reference blocks can be fused.

[0117] In the sub-pixel precision mode of IntraTMP, sub-pixel locations around integer pixel locations obtained through template matching can be activated for template search. A discrete cosine transform-based interpolation filter (DCT-IF) filter can be used for sub-pixel interpolation.

[0118] In linear filter model mode, the current block is predicted from a filtered reference block using a 6-tap linear filter with a bias added to a 5-tap cross-shaped filter. The filter coefficients can be calculated based on the relationship between the current template and the reference template.

[0119] Figure 7 is a conceptual diagram illustrating exemplary templates of the current block that can be used for template matching.

[0120] As illustrated in (a) of FIG. 7, the current block may have an L-shaped template (710) including an above part, an above-left part, and a left part. The above part may have a height of 'n' and the left part may have a width of 'm'. The above-left part may have a height of n and a width of m. Here, m and n may be the same or different depending on the aspect ratio of the current block. For example, if the current block has a square shape, m and n may be the same. If the current block has a horizontally long rectangular shape, n may be larger than m. m and n may be the same regardless of the aspect ratio of the current block.

[0121] As illustrated in (b) of FIG. 7, the current block may have a template (720) that includes only the left part. As illustrated in (c) of FIG. 7, the current block may have a template (730) that includes only the upper part. As illustrated in (d) of FIG. 7, the current block may have a template (740) that includes a left part (740a) and an upper part (740b), but does not include the upper left part. The left part (740a) and the upper part (740b) together form the template (740).

[0122] A video encoder and a video decoder can determine which types of templates are available based on the availability of surrounding blocks. To select a template to use for a current block, the video encoder and the video decoder can determine whether the surrounding blocks to the left of the current block, the surrounding blocks above the current block, and the surrounding blocks to the upper left of the current block are available, respectively. For example, if the current block is not adjacent to the left border of the picture and is not adjacent to the upper border of the picture, an L-shaped template (710) including an upper part, an upper-left part, and a left part for the current block may be used. As another example, if the current block is adjacent to the upper border of the picture but is not adjacent to the left border of the picture, a template (720) including only the left part for the current block may be used. As yet another example, if the current block is adjacent to the left border of the picture but is not adjacent to the upper border of the picture, a template (730) including only the upper part for the current block may be used.

[0123] The video encoder may also signal an indication in the bitstream identifying the template type selected for the current block.

[0124] Additionally, the video encoder and decoder can determine which type of template is available for the current block, depending on the coding tool that utilizes template matching. For example, the templates (710, 720, 730) illustrated in (a), (b), and (c) of FIG. 7 can be used in IntraTMP mode, DIMD, or TIMD, and the template (740) illustrated in (d) of FIG. 7 can be used in SGPM or inter TM (inter template matching).

[0125] DIMD (Decoder-Side Intra Mode Derivation) and TIMD (Template-based Intra Mode Derivation) are techniques for deriving one or more intra prediction modes based on template analysis.

[0126] FIG. 8 is a conceptual diagram illustrating a method for deriving one or more intra prediction modes for a current block based on template analysis in DIMD technology.

[0127] The DIMD method derives directional intra prediction modes by performing texture gradient analysis on a template, which is a set of reconstructed pixels adjacent to the current block. This technique extracts orientation and magnitude information by analyzing vertical and horizontal gradients within the template. To this end, the encoder and decoder can calculate the horizontal gradient Gx and the vertical gradient Gy, respectively, by applying a 3×3 horizontal Sobel filter and a vertical Sobel filter to the pixel position at the center of the template, as shown in Fig. 8 (a). The orientation of the corresponding pixel position is calculated using atan(Gy / Gx), and the sum of the absolute values ​​of Gx and Gy can be calculated as the magnitude of the orientation. The size of the template may be equal to or larger than the size of the Sobel filter. The template may also include adjacent samples at the lower left and upper right of the current block.

[0128] The angle and magnitude calculated for each pixel position in the center of the template area are used to generate a Histogram of Gradient (HoG) as shown in (b) of Fig. 8, and the HoG is configured by accumulating the angle along the x-axis and the magnitude along the y-axis. The encoder and decoder can derive or determine the corresponding intra prediction modes by mapping several angles (for example, up to 5) with relatively high accumulated magnitude values ​​on the HoG to directional modes. The intra prediction modes determined by the DIMD method may also be referred to as DIMD modes and can be utilized for subsequent MPM generation.

[0129] The predictors of the intra prediction modes determined in this way can be weighted and combined with non-directional predictors (based on planar or block vectors) to form the final predicted block.

[0130] FIG. 9 is a conceptual diagram illustrating a method for deriving one or more intra prediction modes for a current block based on template prediction in TIMD technology.

[0131] The Template-based Intra Mode Derivation (TIMD) method derives the intra prediction mode of the current block in the video decoder by testing various intra prediction modes on the template of the current block (i.e., the current template) without signaling the intra prediction mode, and selecting the optimal mode based on the template cost. For example, the video decoder constructs a list of MPM candidates, and uses each MPM candidate to derive prediction samples for the current template from surrounding reference samples of the current template, and then calculates the Sum of Absolute Transformed Difference (SATD) cost between the prediction samples of the current template and the reconstructed samples of the current template, and then selects one, two, or three intra prediction modes with the smallest SATD cost for the prediction of the current block.

[0132] When two or more prediction modes are selected, the prediction blocks for the selected prediction modes can be blended using a weighted sum or weighted average method to generate the final prediction block of the current block. For example, when the following mathematical equation is satisfied, the prediction mode with the smallest SATD and the prediction mode with the second-smallest SATD can be used for weighted intra prediction for the current block. When the following mathematical equation is not satisfied, the mode with the smallest SATD can be used as the intra prediction mode for the current block.

[0133]

[0134] Here, costMode 1 represents the SATD cost of the prediction mode with the lowest SATD cost, and costMode 2 represents the SATD cost of the prediction mode with the second-lowest SATD cost. Furthermore, the weights used to blend the two prediction blocks can be calculated using the mathematical formula below. A higher weight is assigned to the mode with a relatively low SATD cost.

[0135]

[0136] Spatial geometric partitioning mode (SGPM) is an intra mode conceptually similar to GPM, which is applied to inter prediction in the VVC standard. In this mode, a CU is partitioned into two parts that can use different intra prediction modes. Because there are many possible combinations of partitions and intra prediction modes, SGPM uses a different signaling mechanism than GPM. To more efficiently represent the selected partition and prediction information for the current block in the bitstream, a candidate list is used, and the selected candidate index is signaled in the bitstream. Each candidate in the candidate list indicates a combination of one partition mode and two intra prediction modes.

[0137] Figure 10 is a conceptual diagram illustrating a method for signaling a combination of partition mode and intra prediction mode of SGPM.

[0138] As illustrated in (a) of Fig. 10, a current block (CU; 1011) can be partitioned into two parts (1011a, 1011b) having different intra prediction modes ("intra_pred_mode0" and "intra_pred_mode1") using a partition mode. This information can be expressed as a syntax structure having one candidate index ("sgpm_cand_idx") indicating a combination of one partition mode ("partition mode idx") and two intra prediction modes ("intra_pred_mode0_idx" and "intra_pred_mode0_idx"), as illustrated in (b) of Fig. 10.

[0139] Figure 11 is a conceptual diagram illustrating the template of the current block to which the SGPM mode is applied and the weights extended to the template area.

[0140] Templates can be used to generate a candidate list. FIG. 11 illustrates one template (1120) represented by a current block (1110) and two template parts (1120a, 1120b), wherein the size (width) of the template (1120) is set to 4. In the case of SGPM adopted in the current version of ECM, the size of the template is set to 1.

[0141] For each possible combination of one partition mode and two intra prediction modes, a prediction for a template (1120) is generated, and blending weights for the current block are extended to the template (1120) as illustrated in FIG. 11 . Alternatively, each template region can use a single value as a blending weight based on the partition boundary. For example, a value of 0 or 1 can be used as the weight for each template region. These combinations are ranked in ascending order according to the sum of absolute difference (SAD) cost between the template prediction and reconstruction. The size of the candidate list can be set to, for example, 16, and these candidates are considered the most probable SGPM combinations for the current block. The video encoder and the video decoder construct the same candidate list using the template (1120).

[0142] To reduce the complexity of candidate list generation, both the number of available partition modes and the number of available intra prediction modes can be limited. For example, in the current version of SGPM adopted for ECM, 26 partition modes and 9 intra prediction modes are used to form combinations. The 9 intra prediction modes include some of the typical intra prediction modes illustrated in Figure 3a (e.g., a directional mode parallel to the partition boundary, a directional mode perpendicular to the partition boundary, and a PLANAR mode), and may also include prediction modes based on block vectors obtained from neighboring blocks (e.g., an IntraTMP mode or an IBC mode).

[0143] Figure 12 is a conceptual diagram for explaining the blending area of ​​SGPM and the blending weights applied thereto.

[0144] As illustrated in Fig. 12, the adaptive SGPM blending scheme can be used to obtain better predictions for pixels in the blending region near the partitioning boundary between two prediction parts, where the weighted average (or weighted sum) of the two prediction parts is used in the blending region around the partitioning boundary. The width (τ) of the blending region is also called the blending depth. The blending depth is adaptively determined based on the size of the current block, and the weights are derived from a predefined ramp function based on the distance (d) from the sample location (x, y) to be predicted to the partitioning boundary. Therefore, adaptive blending does not require signaling.

[0145] The present disclosure describes techniques for enhancing the functionality of template analysis or template matching in video coding, thereby improving coding efficiency. A template of a current block can be divided into two or more sub-templates, each of which can be used for template analysis or template matching. Using these sub-templates allows for the discovery of more diverse candidates or predictors, and fusion can be applied to combine them for better prediction.

[0146] According to one aspect of the present disclosure, once a template of a current block is determined, a video encoder and a video decoder can divide the template of the current block into a plurality of sub-templates before using the template to encode the current block.

[0147] The segmentation of the template of the current block may be based on segmentation information, prediction information, etc. used for encoding and decoding of surrounding blocks that contain at least some of the samples in the template of the current block.

[0148] FIG. 13 is a conceptual diagram showing an example of sub-templates divided based on block boundaries of surrounding blocks including a template area of ​​a current block according to one embodiment of the present invention.

[0149] The template (1320) of the current block is composed of already decoded samples of already decoded neighboring blocks. The video encoder and video decoder can define multiple subtemplates based on the block division boundaries of the neighboring blocks. The block division boundaries of the neighboring blocks can be identified from the block division information of the neighboring blocks.

[0150] Subtemplates belonging to different blocks are likely to have different textures. Therefore, performing template analysis or template matching on a subtemplate-by-subtemplate basis may be advantageous in discovering more different candidates or predictors favorable for predicting the current block than performing template analysis or template matching on the entire region of the current block's template (1320).

[0151] The following restrictions may be imposed on the division of a template (1320). A minimum area (P) allowed for a sub-template may be defined. The minimum area may be a predefined fixed value and may be determined based on the size and / or aspect ratio of the current block. A maximum number (Q) of sub-templates that a template (1320) may have may be defined. The maximum number may be a predefined fixed value (e.g., 2 or 3) and may be determined based on the size and / or aspect ratio of the current block.

[0152] When dividing a template of a current block into sub-templates according to the block division boundaries of surrounding blocks, a sub-template having a smaller area (P) than the minimum area allowed for the sub-template may be generated. Fig. 14 is a conceptual diagram showing another example of sub-templates divided based on the block boundaries of surrounding blocks including the template of the current block according to an embodiment of the present invention. As illustrated in Fig. 14, a sub-template having a smaller area (P) can be merged into an adjacent sub-template, so that two sub-templates can form a single larger sub-template.

[0153] Two or more sub-templates may be adjacent to a sub-template having a smaller area than the minimum area (P). In this case, one or more sub-templates into which the smaller sub-template will be merged may be selected according to a predefined rule. For example, among two or more adjacent sub-templates, a sub-template having a relatively larger x value or a sub-template having a relatively larger y value may be selected based on the upper-left sample location (x, y) of each sub-template. In another example, among two or more adjacent sub-templates, a sub-template using the same prediction mode as a sub-template having a smaller area than the minimum area (P) may be selected. In yet another example, among two or more adjacent sub-templates, a sub-template having a relatively larger area may be selected.

[0154] When dividing the template of the current block into subtemplates according to the block boundaries of surrounding blocks, more subtemplates may be generated than the maximum number of subtemplates (Q). In this case, adjacent subtemplates may be merged to form a single larger subtemplate so that the template of the current block has subtemplates less than or equal to the maximum number (Q). For example, subtemplates selected in the order of smaller areas may be merged into adjacent subtemplates until the number of subtemplates reaches the maximum number (Q). Alternatively, two or more adjacent subtemplates having the same prediction mode may be merged to form a single larger subtemplate.

[0155] When dividing the template of the current block into subtemplates according to the block boundaries of the surrounding blocks, two constraints may be imposed, including a minimum area (P) allowed for the subtemplates and a maximum number (Q) of subtemplates. The video encoder and video decoder can first merge subtemplates with a smaller area (P) into adjacent subtemplates, and then determine whether the template of the current block has subtemplates with a smaller number (Q) or less. If the number of remaining subtemplates is greater than the maximum number (Q), the subtemplates selected in the order of the smaller area can be merged into the adjacent subtemplates.

[0156] The video encoder and the video decoder may compare an unfiltered current template with an unfiltered reference template to determine a template matching cost. In some embodiments, the video encoder and the video decoder may filter at least one of the current template or the reference template to generate a filtered current template or a filtered reference template. For example, a smoothing filter having a coefficient of [1 / 4, 2 / 4, 1 / 4] may be used for the filtering. To determine the template matching cost, the video encoder and the video decoder may compare the filtered current template with the reference template, compare the current template with the filtered reference template, or compare the filtered current template with the filtered reference template.

[0157] Below, we describe how subtemplates can be used in coding techniques for encoding or decoding the current block based on the aforementioned template analysis or template matching cost.

[0158] FIGS. 15A to 15D are conceptual diagrams illustrating a method for deriving TIMD modes and constructing a candidate list when a current template is divided into sub-templates according to one embodiment of the present invention.

[0159] According to one embodiment of the present invention, when a template-based intra-mode derivation (TIMD) method is used for a current block, sub-templates derived from templates of the current block may be used. A video encoder and a video decoder may divide the current block into a plurality of regions by extension of boundaries between sub-templates, and derive one or more intra-prediction modes for corresponding regions of the current block using each sub-template. The derived intra-prediction modes may be referred to as TIMD modes.

[0160] For example, as illustrated in FIGS. 15a to 15d, one or more TIMD modes for region A' of the current block can be derived by applying the TIMD method to sub-template A adjacent to region A', and one or more TIMD modes for region B' of the current block can be derived by applying the TIMD method to sub-template B adjacent to region B'.

[0161] The video encoder and the video decoder may construct a candidate list, use each candidate to generate prediction samples for the subtemplate from surrounding reference samples of the subtemplate, calculate a SATD cost between the prediction samples of the subtemplate and the reconstructed samples of the subtemplate, and then select one or two intra prediction modes with the smallest SATD cost as the TIMD mode for the corresponding prediction region of the current block.

[0162] Here, the candidate list may be constructed based on prediction information of surrounding blocks of the current block (e.g., left, top, bottom left, top right, top left) and / or previously restored blocks that are not adjacent to the current block, as illustrated in FIGS. 15a and 15b, rather than being constructed by sub-template or by prediction region of the current block.

[0163] Alternatively, the candidate lists can be organized by sub-template or by region of the current block. For example, as illustrated in FIGS. 15c and 15d, the candidate lists to be used for generating prediction samples for sub-template A can be organized based on prediction information of neighboring blocks of region A' of the current block, and the candidate lists to be used for generating prediction samples for sub-template B can be organized based on prediction information of neighboring blocks of region B' of the current block.

[0164] FIGS. 16A and 16B are conceptual diagrams illustrating a method for deriving DIMD modes when a current template is divided into sub-templates according to one embodiment of the present invention.

[0165] According to one embodiment of the present invention, when the DIMD (Decoder-Side Intra Mode Derivation) method is used for the current block, subtemplates derived for the template of the current block can be used.

[0166] A video encoder and a video decoder may divide a current block into multiple regions by extending boundaries between sub-templates, and perform gradient analysis on each sub-template to derive one or more intra prediction modes for corresponding regions of the current block. The derived intra prediction modes may be referred to as DIMD modes.

[0167] That is, the video encoder and video decoder can perform gradient analysis on each sub-template to generate a gradient histogram (HoG) for each sub-template and map one or more angles with relatively high cumulative magnitude values ​​on each HoG to directional modes, thereby deriving one or more DIMD modes for the region of the current block adjacent to each sub-template.

[0168] For example, as illustrated in FIG. 16a, one or more DIMD modes for region A' of the current block can be derived by performing gradient analysis on subtemplate A adjacent to region A', and one or more TIMD modes for region B' of the current block can be derived by performing gradient analysis on subtemplate B adjacent to region B'.

[0169] The video encoder and video decoder can generate prediction samples for each region of the current block by weighting predictors from one or more intra prediction modes derived for each region with non-directional predictors, such as from a PLANAR mode.

[0170] In some embodiments, as illustrated in FIG. 16b, the current block is not divided into multiple regions, and DIMD modes selected through gradient analysis for each sub-template may be used for the current block. That is, one or more DIMD modes for the current block may be selected from the gradient analysis for sub-template A, and one or more DIMD modes for the current block may be selected from the gradient analysis for sub-template B. The video encoder and the video decoder may also predict the current block through a weighted sum of multiple predictors obtained using the DIMD modes selected from the two gradient analyses.

[0171] In some embodiments, the video encoder and the video decoder may derive one or more DIMD modes for the current block by performing gradient analysis only on the largest size subtemplate among the subtemplates derived for the template of the current block.

[0172] According to some embodiments of the present invention, when the intra-template matching prediction (IntraTMP) mode is applied to the current block, sub-templates derived for the template of the current block may be used.

[0173] Figures 17a to 17c are conceptual diagrams for explaining an IntraTMP method using subtemplates according to one embodiment of the present invention.

[0174] In one embodiment, when the IntraTMP mode is applied to the current block, the video encoder and the video decoder can divide the current block into a plurality of regions by straight lines extending from boundaries between sub-templates, and generate a prediction signal for a corresponding region of the current block based on template matching using each sub-template.

[0175] In the example of Fig. 17a, the template of the current block is divided into two sub-templates (A, B). The current block is divided into two regions (A', B') by a straight line extending from the boundary between the two sub-templates. The prediction signal of each region is generated based on template matching using adjacent sub-templates. The prediction signal of region A' is generated based on template matching using the sub-template A adjacent to the left, and the prediction signal of region B' is generated based on template matching using the sub-template B adjacent to the top.

[0176] As described above, the Intra TMP technique includes several sub-modes, and each of the two regions (A', B') of the current block can be applied with one of the single predictor mode, the fusion of multiple predictors mode, the sub-pel precision mode, and the linear filter model mode.

[0177] Fig. 17b illustrates a case where two regions (A', B') of the current block are each predicted in a single predictor mode. Referring to Fig. 17b, in template matching using sub-template A, a first reference template (1721) having a region (1721a) most similar to sub-template A and a first reference block (1720: 1720a, 1720b) ​​corresponding thereto can be determined. Regions (1721a) and (1721b) together form the first reference template (1721), and regions (1720a) and (1720b) ​​together form the first reference block (1720). The block vector (1711) identifies the location of the first reference block (1720a, 1720b) ​​based on the location of the current block within the current picture.

[0178] The video encoder and the video decoder can perform template matching using the sub-template B to determine a second reference template (1731) having a region (1731b) most similar to the sub-template B and a second reference block (1730a, 1730b) corresponding thereto. The region (1731a) and the region (1731b) together form the second reference template (1731), and the region (1730a) and the region (1730b) together form the second reference block (1730). The block vector (1712) identifies the location of the second reference block (1730a, 1730b) based on the location of the current block within the current picture.

[0179] Reconstructed samples (or filtered samples thereof) of the region (1720a) of the first reference block (1720a, 1720b) ​​can be used as prediction samples for the region A' of the current block, and reconstructed samples (or filtered samples thereof) of the region (1730b) of the second reference block (1730: 1730a, 1730b) can be used as prediction samples for the region B' of the current block.

[0180] For each of the subtemplates of the current block, the video encoder and the video decoder can search for regions similar to the subtemplate (which may be referred to as 'reference subtemplates' or 'subtemplates of the reference templates') in the predefined search regions (R1 to R6) illustrated in FIG. 6, and construct a candidate list including candidate block vectors pointing to each of the similar regions. A candidate list may be constructed first for a subtemplate having the largest area among the subtemplates of the current block. A search for constructing candidate lists for the remaining subtemplates may be performed only in regions (R1 to R6) that include each of the candidates included in the candidate list constructed for the subtemplate having the largest area.

[0181] When the current block is divided into multiple regions by straight lines extending from each of the boundaries between multiple sub-templates, regions that are not adjacent to any sub-template may occur. For example, referring to the example shown in Fig. 17c, the upper part of the template of the current block is vertically divided, and at the same time, the left part of the template of the current block is horizontally divided. The current block is divided into four regions along straight lines extending from each of the boundaries between the sub-templates, and the lower right region is not adjacent to any sub-template.

[0182] In this case, a prediction signal for the lower right region can be generated based on template matching using a sub-template with a larger area among sub-templates A located on the left side of the lower right region and C located on the upper side of the lower right region. Alternatively, the prediction signal for the lower right region can be generated through a weighted sum of a prediction signal generated based on template matching using sub-template A and a prediction signal generated based on template matching using sub-template C. At this time, the weights can be a fixed ratio such as 1:1, or determined in proportion to the area of ​​the template, or determined based on a template matching cost value of each sub-template.

[0183] In some embodiments, when IntraTMP mode is applied to the current block, splitting the template of the current block into more than two subtemplates may not be allowed.

[0184] FIG. 18a is a conceptual diagram illustrating an intra-template matching prediction (IntraTMP) method using sub-templates according to another embodiment of the present invention. FIG. 18b is a diagram illustrating a method for predicting a current template through a weighted sum of reference templates according to another embodiment of the present invention.

[0185] In one embodiment, when the IntraTMP mode is applied to the current block, the video encoder and the video decoder can generate predictors for the current block based on template matching using subtemplates, and weight the predictors corresponding to each subtemplate to generate a final prediction signal for the current block.

[0186] Referring to Fig. 18a, the template of the current block is divided into two sub-templates (A, B). The video encoder and the video decoder can perform template matching using sub-template A to determine a first reference template (1821) having a region (1821a) most similar to sub-template A and a first reference block (1820) corresponding thereto. The region (1821a) and the region (1821b) together form the first reference template (1821). The block vector (1811) identifies the location of the first reference block (1820) based on the location of the current block within the current picture.

[0187] The video encoder and the video decoder can perform template matching using subtemplate B to determine a second reference template (1831) having a region (1831b) most similar to subtemplate B and a second reference block (1830) corresponding thereto. The region (1831a) and the region (1831b) together form the second reference template (1831). The block vector (1812) identifies the location of the second reference block (1830) based on the location of the current block within the current picture.

[0188] The first reference block (1820) (or a filtered version thereof) is the first predictor (P) for the current block. a ) is determined, and the second reference block (1830) (or its filtered version) is the second predictor (P) for the current block. b) can be determined. As shown in the following mathematical formula, the prediction block of the current block can be generated using the weighted sum by pixel position between each predictor, and the weight matrices (W0, W1) of dimension W×H can be used for the weighted sum by sample position for the current block of size W×H.

[0189]

[0190] In some embodiments, the weight matrices can be implicitly derived or determined using the current template and reference templates (determined based on template matching using each subtemplate). The weight matrices can be modeled as affine linear functions of the sample locations (x, y), for example, as follows:

[0191]

[0192] Here, parameters a, b, and c can be determined to minimize a loss function based on mean squared error (MSE) or sum of absolute difference (SAD) between the reconstructed samples of the template of the current block and the predicted samples of the template of the current block. As illustrated in Fig. 18b, the predicted samples of the template of the current block are obtained through a weighted sum of the first reference template (1821: 1821a, 1821b) and the second reference template (1831: 1831a, 1831b) using the weight matrix of Equation 4.

[0193] In some other embodiments, the weight matrices may be implicitly derived or determined based on the size, aspect ratio, etc. of the current block. For example, the weights may be derived from a ramp function based on the distance (d) from the sample location (x, y) to be predicted to the segmentation boundary within the current block. That is, the weights of the first weight matrix may be set to gradually increase in value and the weights of the second weight matrix may be set to gradually decrease in value as the distance from the segmentation boundary within the current block defined by the extension of the boundary between two sub-templates in the first normal direction. The weights of the first weight matrix may be set to gradually decrease in value and the weights of the second weight matrix may be set to gradually increase in value as the distance from the segmentation boundary within the current block in the second normal direction opposite to the first normal direction.

[0194] In some other embodiments, the weight matrices may be implicitly derived or determined using the subtemplates of the current block. For example, the weight values, the weighted sum width, and / or the weighted sum area may be determined based on the (average) difference in brightness values ​​between the reconstructed samples adjacent to the subtemplate boundaries of the two subtemplates. If the difference in brightness values ​​is greater than a predefined threshold, the weighted sum width may be set narrower than otherwise.

[0195] In some other embodiments, the video encoder can signal an index indicating a weighted sum width, and the video decoder can determine the weighted sum width corresponding to the index from a predefined table. The values ​​of the weights within the weighted sum region can be derived using a predefined mathematical formula, or can be derived using the (sub)template of the current block and the reference (sub)template, as in the methods described above.

[0196] As mentioned above, the IntraTMP block vector of a block coded in the IntraTMP mode can be stored together with the block and used to construct a candidate list of IBC block vectors for a subsequent block coded in the IBC mode. If the aforementioned subtemplates were used when a given block was coded in the IntraTMP mode, multiple IntraTMP block vectors may be determined for the given block based on template matching using each subtemplate. In this case, all of the multiple IntraTMP block vectors may be used as IBC block vector candidates for the subsequent block, or only the block vector determined based on template matching using the subtemplate with the largest area may be used as the IBC block vector candidate for the subsequent block, or an average value of the multiple block vectors may be used as the IBC block vector candidate for the subsequent block.

[0197] Meanwhile, the method for deriving a weight matrix described with reference to FIGS. 18A and 18B can also be used for SGPM. As described above, when SGPM is applied to a current block, a partition mode defining the partition boundary of the current block and two prediction modes to be applied to each part of the current block are specified by candidate indices signaled in the bitstream. The two prediction modes can be the conventional intra modes illustrated in FIG. 3A, or a block vector-based prediction mode derived from neighboring blocks coded in IntraTMP mode or IBC mode. Therefore, when SGPM is applied to a current block, the two parts of the current block may each be predicted in IntraTMP mode, and in such a case, the method for deriving a weight matrix described with reference to FIGS. 18A and 18B can also be used.

[0198] For example, as illustrated in FIG. 19, two IntraTMP block vectors (1911, 1912) to be applied to parts (1900a, 1900b) of the current block can be determined by the SGPM candidate index. In such a case, a prediction block (1940) of the current block can be generated through a weighted sum between a first reference block (1920) identified by a first block vector (1911) determined for a first part (1900a) of the current block and a second reference block (1930) identified by a second block vector (1912) determined for a second part (1900b).

[0199] The weight matrices can be implicitly derived or determined using the template (1901) of the current block and the reference templates (1921, 1931). The template (1901) of the current block includes a left part (1901a) and an upper part (1901b), the reference template (1921) includes a left part (1921a) and an upper part (1921b), and the reference template (1931) includes a left part (1931a) and an upper part (1931b). The weight matrices (W0, W1) for the weighted sum can be modeled as affine linear functions of the sample locations (x, y) within the current block as in Equation 4.

[0200] Parameters a, b, and c may be determined to minimize an MSE or SAD-based loss function between the reconstructed samples of the template (1901) of the current block and the predicted samples of the template (1901) of the current block. Here, the predicted samples of the template of the current block may be defined as a weighted sum of pixel positions of the reference template (1921) of the first reference block (1920) and the reference template (1931) of the second reference block (1930).

[0201] The weight matrices derived by Equation 4 no longer depend on the partition boundary between two regions within the current block in the example of Fig. 18a or the SGPM partition boundary within the current block in the example of Fig. 19. That is, when predicting the current block by a weighted sum by pixel positions for two reference blocks using the weight matrices derived by Equation 4, the partition boundary between two regions within the current block or the SGPM partition boundary within the current block no longer affects the generation of prediction samples of the current block. Therefore, when predicting the current block in the example of Fig. 18a, there is no need to consider the partition boundary between two regions within the current block, and in the example of Fig. 18a, there is no need to specify the partition mode defining the partition boundary of the current block by the candidate index signaled in the bitstream, and only two prediction modes (two IntraTMP block vectors in the above example) to be applied to each part of the current block are sufficient.

[0202] FIG. 20 is a flowchart illustrating a method of encoding or decoding video data according to one embodiment of the present invention.

[0203] A video encoder or a video decoder may determine a first IntraTMP block vector for a current block and determine a first reference block identified by the first IntraTMP block vector (S2010). The video encoder or a video decoder may determine a second IntraTMP block vector for the current block and determine a second reference block identified by the second IntraTMP block vector (S2020).

[0204] In some embodiments, a current template, which is a set of reconstructed samples adjacent to a current block, may be split into two sub-templates based on block splitting information of already decoded neighboring blocks of the current block. The first IntraTMP block vector may be determined based on a template matching cost using the first sub-template among the two sub-templates, and the second IntraTMP block vector may be determined based on a template matching cost using the second sub-template among the two sub-templates.

[0205] In some embodiments, the video encoder may encode a candidate index into the bitstream, the candidate index indicating a candidate to be used for the current block within a candidate list, which is a set of multiple available candidates. The candidate index may identify a first IntraTMP block vector and a first IntraTMP block vector. The video decoder may decode the candidate index from the bitstream to determine the first IntraTMP block vector and the first IntraTMP block vector to be used for the current block.

[0206] A video encoder or video decoder may determine a first weight matrix defining weights per sample position for a first reference block, and may determine a second weight matrix defining weights per sample position for a second reference block (S2030).

[0207] In some embodiments, the first weight matrix and the second weight matrix may be determined based on a mean squared error (MSE) or a sum of absolute differences (SAD) between the predicted values ​​of the current template and the reconstructed values ​​of the current template. Here, the predicted values ​​of the current template may be derived as a weighted sum between the first reference template, which is a set of reconstructed samples adjacent to the first reference block, and the second reference template, which is a set of reconstructed samples adjacent to the second reference block.

[0208] In some embodiments, a video encoder or a video decoder may partition a current block into two regions by a straight line extending from a partitioning boundary between sub-templates. Weights in the first weight matrix and the second weight matrix may be derived from a ramp function based on a distance from each sample position to the partitioning boundary between the two regions within the current block. A blending region may be defined around the partitioning boundary between the two regions within the current block, and a weighted sum of sample values ​​of a first reference block and sample values ​​of a second reference block in the blending region of the current block may be used as prediction values ​​of the current block, and in the remaining region of the current block, sample values ​​of the first reference block or sample values ​​of the second reference block may be used as prediction values ​​of the current block. The width of the blending region may be adaptively determined according to a size of the current block.

[0209] A video encoder or video decoder can generate a prediction block by fusing sample values ​​of a first reference block and sample values ​​of a second reference block based on a first weight matrix and a second weight matrix (S2040).

[0210] A video encoder may encode a current block based on a prediction block (S2050). For example, the video encoder may encode residual data indicating the difference between the current block and the prediction block into a bitstream. A video decoder may decode the current block based on the prediction block (S2050). For example, the video decoder may decode residual data indicating the difference between the current block and the prediction block from a bitstream, and add the prediction block to the residual data to reconstruct the current block.

[0211] Although the flowchart / timing diagram of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of ​​one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart / timing diagram without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart / timing diagram is not limited to a chronological order.

[0212] It should be understood that the exemplary embodiments described above can be implemented in many different ways. The functions or methods described in one or more examples can be implemented in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as "units" to further emphasize their implementation independence.

[0213] Meanwhile, the various functions or methods described in this embodiment may also be implemented as instructions stored on a non-transitory storage medium that can be read and executed by one or more processors. Non-transitory storage media include, for example, all types of storage devices that store data in a form readable by a computer system. For example, non-transitory storage media include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical drives, magnetic hard drives, and solid-state drives (SSDs).

[0214] The above description is merely an example of the technical idea of ​​the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of ​​the present embodiment, but rather to explain it, and the scope of the technical idea of ​​the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.

[0215]

[0216]

[0217] CROSS-REFERENCE TO RELATED APPLICATION

[0218] This patent application claims priority to Korean patent application No. 10-2024-0001102, filed in Korea on January 3, 2024, and Korean patent application No. 10-2024-0179702, filed in Korea on December 5, 2024, the entire contents of which are incorporated herein by reference.

Claims

1. A method for encoding video data, A step of determining a first IntraTMP block vector for the current block; A step of determining a first reference block identified by the first IntraTMP block vector A step of determining a second IntraTMP block vector for the current block; A step of determining a second reference block identified by the second IntraTMP block vector A step of determining a first weight matrix defining weights for each sample location for the first reference block; A step of determining a second weight matrix defining weights for each sample location for the second reference block; A step of generating a prediction block by fusing sample values ​​of the first reference block and sample values ​​of the second reference block based on the first weight matrix and the second weight matrix; and A step of encoding the current block based on the above prediction block. A method comprising:

2. In paragraph 1, The first weight matrix and the second weight matrix are determined based on the mean squared error (MSE) or the sum of absolute difference (SAD) between the predicted values ​​of the current template, which is a set of reconstructed samples adjacent to the current block, and the reconstructed values ​​of the current template, A method, characterized in that the prediction values ​​of the current template are derived as a weighted sum between a first reference template, which is a set of reconstructed samples adjacent to the first reference block, and a second reference template, which is a set of reconstructed samples adjacent to the second reference block.

3. In paragraph 2, Further comprising a step of dividing the current template into two sub-templates based on block division information of already decrypted surrounding blocks of the current block, A method characterized in that, wherein the first IntraTMP block vector is determined based on a template matching cost using a first sub-template among the two sub-templates, and the second IntraTMP block vector is determined based on a template matching cost using a second sub-template among the two sub-templates.

4. In paragraph 3, A method characterized in that the first IntraTMP block vector and the second IntraTMP block vector are used to construct an IBC block vector candidate list of another block to be encoded subsequent to the current block.

5. In paragraph 2, A method further comprising the step of encoding into a bitstream a candidate index indicating a candidate used for said current block within a candidate list which is a set of multiple available candidates, wherein said candidate index is characterized in that it identifies said first IntraTMP block vector and said second IntraTMP block vector.

6. In paragraph 1, A step of dividing a current template, which is a set of reconstructed samples adjacent to the current block, into two sub-templates based on block division information of already decrypted neighboring blocks of the current block; and A step of dividing the current block into two regions by a straight line extending from the dividing boundary between the above sub-templates. Including more, A method, characterized in that the weights in the first weight matrix and the second weight matrix are derived from a ramp function based on the distance from each sample location to the segmentation boundary between the two regions within the current block.

7. In paragraph 6, A method characterized in that, in a blending region around a division boundary between the two regions within the current block, a weighted sum of sample values ​​of the first reference block and sample values ​​of the second reference block is used as prediction values ​​of the current block, and in an region other than the blending region, sample values ​​of the first reference block or sample values ​​of the second reference block are used as prediction values ​​of the current block.

8. In paragraph 7, A method, characterized in that the width of the blending region is adaptively determined according to the size of the current block.

9. A method for decrypting video data, A step of determining a first IntraTMP block vector for the current block; A step of determining a first reference block identified by the first IntraTMP block vector A step of determining a second IntraTMP block vector for the current block; A step of determining a second reference block identified by the second IntraTMP block vector A step of determining a first weight matrix defining weights for each sample location for the first reference block; A step of determining a second weight matrix defining weights for each sample location for the second reference block; A step of generating a prediction block by fusing sample values ​​of the first reference block and sample values ​​of the second reference block based on the first weight matrix and the second weight matrix; and A step of decrypting the current block based on the above predicted block. A method comprising:

10. In paragraph 9, The first weight matrix and the second weight matrix are determined based on the mean squared error (MSE) or the sum of absolute difference (SAD) between the predicted values ​​of the current template, which is a set of reconstructed samples adjacent to the current block, and the reconstructed values ​​of the current template, A method, characterized in that the prediction values ​​of the current template are derived as a weighted sum between a first reference template, which is a set of reconstructed samples adjacent to the first reference block, and a second reference template, which is a set of reconstructed samples adjacent to the second reference block.

11. In clause 10, Further comprising a step of dividing the current template into two sub-templates based on block division information of already decrypted surrounding blocks of the current block, A method characterized in that, wherein the first IntraTMP block vector is determined based on a template matching cost using a first sub-template among the two sub-templates, and the second IntraTMP block vector is determined based on a template matching cost using a second sub-template among the two sub-templates.

12. In paragraph 11, A method, characterized in that the first IntraTMP block vector and the second IntraTMP block vector are used to construct an IBC block vector candidate list of another block to be decrypted subsequent to the current block.

13. In paragraph 10, A method further comprising the step of encoding into a bitstream a candidate index indicating a candidate used for said current block within a candidate list which is a set of multiple available candidates, wherein said candidate index is characterized in that it identifies said first IntraTMP block vector and said second IntraTMP block vector.

14. In paragraph 9, A step of dividing a current template, which is a set of reconstructed samples adjacent to the current block, into two sub-templates based on block division information of already decrypted neighboring blocks of the current block; and A step of dividing the current block into two regions by a straight line extending from the dividing boundary between the above sub-templates. Including more, A method, characterized in that the weights in the first weight matrix and the second weight matrix are derived from a ramp function based on the distance from each sample location to the segmentation boundary between the two regions within the current block.

15. In paragraph 14, A method characterized in that, in a blending region around a division boundary between the two regions within the current block, a weighted sum of sample values ​​of the first reference block and sample values ​​of the second reference block is used as prediction values ​​of the current block, and in an region other than the blending region, sample values ​​of the first reference block or sample values ​​of the second reference block are used as prediction values ​​of the current block.

16. In paragraph 15, A method, characterized in that the width of the blending region is adaptively determined according to the size of the current block.

17. A method for providing video data to a video decoding device. A step of encoding the above video data into a bitstream; and A step of transmitting the bitstream to the video decoding device, and a step of encoding the video data into a bitstream, A step of determining a first IntraTMP block vector for the current block; A step of determining a first reference block identified by the first IntraTMP block vector A step of determining a second IntraTMP block vector for the current block; A step of determining a second reference block identified by the second IntraTMP block vector A step of determining a first weight matrix defining weights for each sample location for the first reference block; A step of determining a second weight matrix defining weights for each sample location for the second reference block; A step of generating a prediction block by fusing sample values ​​of the first reference block and sample values ​​of the second reference block based on the first weight matrix and the second weight matrix; and A step of encoding the current block based on the above prediction block. A method comprising:

Citation Information

Patent Citations

  • Method and apparatus for compressing video using template matching and motion prediction

    KR1020120049435A

  • Vibration-based pipe flow rate measurement method and system therefor

    KR1020250024316A

  • Entrance order device for flood prevention

    KR1020250071407A

  • Mumtiple perforations pipe construction for basement waterproof

    KR102670109B1

  • Method and device for video coding using intra prediction based on template matching

    WO2023090613A1