Inter prediction signal-based intra prediction mode derivation and fusion prediction

By deriving an intra prediction mode based on an inter prediction signal and weighting it with the intra prediction signal, the video coding method enhances encoding efficiency and video quality for high-resolution video data.

WO2025116309A1PCT designated stage expired Publication Date: 2025-06-05HYUNDAI MOTOR CO LTD +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016589
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-24
Filing Date
2024-10-29
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently encoding and decoding high-resolution, high-frame-rate video data, leading to increased data sizes and reduced encoding efficiency.

Method used

A video coding method and device that derive an intra prediction mode based on an inter prediction signal when generating a final prediction signal for a current block by weighting the inter prediction signal and the intra prediction signal.

Benefits of technology

This approach improves video encoding efficiency and enhances video quality by effectively utilizing inter and intra prediction signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016589_05062025_PF_FP_ABST
    Figure KR2024016589_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method for deriving an inter prediction signal-based intra prediction mode. In the present embodiment, an image decoding device determines a motion vector of the current block from a motion vector candidate list of the current block, and generates an inter prediction signal of the current block on the basis of a reference block indicated by the motion vector of the current block. The image decoding device derives an intra prediction mode of the current block on the basis of the reference block and a template of the reference block, and generates an intra prediction signal of the current block on the basis of the intra prediction mode. The image decoding device generates final prediction signals of the current block on the basis of the inter prediction signal and the intra prediction signal.
Need to check novelty before this filing date? Find Prior Art

Description

Intra prediction mode derivation and fusion prediction based on inter prediction signal

[0001] The present disclosure relates to a video coding method and device using intra prediction mode derivation and fusion prediction based on an inter prediction signal.

[0002] The content described below merely provides background information related to the present invention and does not constitute prior art.

[0003] Since video data has a large amount of data compared to voice data or still image data, it requires a lot of hardware resources, including memory, to store or transmit it without processing for compression.

[0004] Therefore, when storing or transmitting video data, an encoder is used to compress the video data and store or transmit it, and a decoder receives the compressed video data, decompresses it, and plays it back. Examples of such video compression technologies include H.264 / AVC, HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), which improves encoding efficiency by about 30% compared to HEVC.

[0005] However, as the size, resolution, and frame rate of images are gradually increasing, and the amount of data that needs to be encoded is also increasing, a new compression technology that has better encoding efficiency and better image quality improvement than existing compression technologies is required.

[0006] The Combined Inter and Intra Prediction (CIIP) mode generates a final prediction signal by weighting the inter prediction signal and the intra prediction signal for the current block. When the CIIP mode is applied in VVC, the inter prediction signal is generated based on the signaling / parsing of the merge index according to the regular merge mode, and the intra prediction signal is generated based on the planar mode. Therefore, to improve video encoding efficiency and enhance video quality, various methods for generating inter prediction signals and intra prediction signals can be considered when the CIIP mode is applied.

[0007] The present disclosure aims to provide a video coding method and device for deriving an intra prediction mode based on an inter prediction signal when generating a final prediction signal of a current block by weighting an inter prediction signal and an intra prediction signal.

[0008] According to an embodiment of the present disclosure, a method for restoring a current block, performed by a video decoding device, is provided, comprising: determining a motion vector of the current block from a motion vector candidate list based on information indicating a motion vector of the current block; generating an inter prediction signal of the current block based on a reference block that exists in a reference picture and is indicated by the motion vector of the current block; deriving an intra prediction mode of the current block based on the reference block and a template of the reference block; generating an intra prediction signal of the current block based on the intra prediction mode; and generating final prediction signals of the current block based on the inter prediction signal and the intra prediction signal.

[0009] According to another embodiment of the present disclosure, a method for encoding a current block, performed by a video encoding apparatus, is provided, comprising: obtaining information indicating a motion vector of the current block; determining a motion vector of the current block from a motion vector candidate list based on the information indicating the motion vector; generating an inter prediction signal of the current block based on a reference block that exists in a reference picture and is indicated by the motion vector of the current block; deriving an intra prediction mode of the current block based on the reference block and a template of the reference block; generating an intra prediction signal of the current block based on the intra prediction mode; and generating final prediction signals of the current block based on the inter prediction signal and the intra prediction signal.

[0010] According to another embodiment of the present disclosure, a method for providing video data to a video decoding device is provided, comprising: encoding the video data into a bitstream; and transmitting the bitstream to the video decoding device, wherein the encoding the video data comprises: obtaining information indicating a motion vector of a current block; determining a motion vector of the current block from a motion vector candidate list based on the information indicating the motion vector; generating an inter prediction signal of the current block based on a reference block that exists in a reference picture and is indicated by the motion vector of the current block; deriving an intra prediction mode of the current block based on the reference block and a template of the reference block; generating an intra prediction signal of the current block based on the intra prediction mode; and generating final prediction signals of the current block based on the inter prediction signal and the intra prediction signal.

[0011] As described above, according to the present embodiment, when generating a final prediction signal of a current block by weighting an inter prediction signal and an intra prediction signal, a video coding method and device for deriving an intra prediction mode based on an inter prediction signal are provided, thereby improving video encoding efficiency and enhancing video quality.

[0012] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure.

[0013] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT (QuadTree plus BinaryTree TernaryTree) structure.

[0014] FIGS. 3A and 3B are diagrams illustrating multiple intra prediction modes, including wide-angle intra prediction modes.

[0015] Figure 4 is an example diagram of the surrounding blocks of the current block.

[0016] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure.

[0017] Figure 6 is an example diagram showing surrounding blocks used in calculating weights.

[0018] Figures 7a and 7b are exemplary diagrams showing sub-blocks according to division of the current block.

[0019] FIG. 8 is a block diagram illustrating in detail a portion of an image decoding device according to one embodiment of the present disclosure.

[0020] FIG. 9a and FIG. 9b are exemplary diagrams showing the surrounding locations of the current block according to one embodiment of the present disclosure.

[0021] FIG. 10 is an exemplary diagram showing an area where template-based motion compensation is performed according to one embodiment of the present disclosure.

[0022] FIG. 11 is an exemplary diagram showing a template of a current block according to one embodiment of the present disclosure.

[0023] FIG. 12 is an exemplary diagram illustrating determination of an intra prediction mode candidate according to the location of a sample according to one embodiment of the present disclosure.

[0024] FIG. 13 is an exemplary diagram showing a reference line of a template according to one embodiment of the present disclosure.

[0025] FIG. 14 is an exemplary diagram showing a restored template surrounding a current block according to one embodiment of the present disclosure.

[0026] FIG. 15 is an exemplary diagram showing the calculation of a slope in a portion of a template according to one embodiment of the present disclosure.

[0027] FIG. 16 is an exemplary diagram showing the calculation of a slope in a portion of a template according to another embodiment of the present disclosure.

[0028] FIG. 17 is an exemplary diagram showing a reference block within a reference picture according to another embodiment of the present disclosure.

[0029] FIG. 18 is an exemplary diagram showing an area used for deriving an intra prediction mode according to one embodiment of the present disclosure.

[0030] FIG. 19a and FIG. 19b are exemplary diagrams showing sub-blocks according to division of a current block according to one embodiment of the present disclosure.

[0031] FIG. 20 is an exemplary diagram showing the division of directional prediction modes according to one embodiment of the present disclosure.

[0032] FIGS. 21A to 21C are exemplary diagrams showing sub-blocks according to division of a current block according to one embodiment of the present disclosure.

[0033] FIG. 22 is a flowchart illustrating a method for an image encoding device to encode a current block according to an embodiment of the present disclosure.

[0034] FIG. 23 is a flowchart illustrating a method for an image decoding device to restore a current block according to one embodiment of the present disclosure.

[0035] Hereinafter, embodiments of the present invention will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, in describing the present embodiments, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present embodiments.

[0036] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image encoding device and its subcomponents will be described with reference to the illustration in FIG. 1.

[0037] The video encoding device may be configured to include a picture segmentation unit (110), a prediction unit (120), a subtractor (130), a transformation unit (140), a quantization unit (145), a reordering unit (150), an entropy encoding unit (155), an inverse quantization unit (160), an inverse transformation unit (165), an adder (170), a loop filter unit (180), and a memory (190).

[0038] Each component of the video encoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0039] A single image (video) is composed of one or more sequences containing multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile or slice is divided into one or more Coding Tree Units (CTUs). Each CTU is then divided into one or more Coding Units (CUs) by a tree structure. Information applied to each CU is encoded as the syntax of the CU, and information commonly applied to CUs included in a CTU is encoded as the syntax of the CTU. In addition, information commonly applied to all blocks within a single slice is encoded as the syntax of the slice header, and information applied to all blocks constituting one or more pictures is encoded in the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referenced by multiple pictures is encoded in a Sequence Parameter Set (SPS). And, information commonly referenced by one or more SPS is encoded in a Video Parameter Set (VPS). In addition, information commonly applied to one tile or tile group may be encoded as syntax of a tile or tile group header. Syntaxes included in an SPS, PPS, slice header, tile or tile group header may be referred to as high level syntax.

[0040] The picture segmentation unit (110) determines the size of the CTU. Information about the size of the CTU (CTU size) is encoded as the syntax of SPS or PPS and transmitted to the image decoding device.

[0041] The picture segmentation unit (110) divides each picture constituting an image into a plurality of CTUs having a predetermined size, and then recursively divides the CTUs using a tree structure. A leaf node in the tree structure becomes a CU, which is a basic unit of encoding.

[0042] The tree structure may be a QuadTree (QT) in which an upper node (or parent node) is divided into four lower nodes (or child nodes) of the same size, a BinaryTree (BT) in which an upper node is divided into two lower nodes, or a TernaryTree (TT) in which an upper node is divided into three lower nodes in a 1:2:1 ratio, or a structure that mixes two or more of the QT structures, BT structures, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure may be used, or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure may be used. Here, BTTT may be combined and referred to as a MTT (Multiple-Type Tree).

[0043] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT structure.

[0044] As illustrated in FIG. 2, a CTU may first be split into a QT structure. The quadtree splitting may be repeated until the size of the splitting block reaches the minimum block size (MinQTSize) of the leaf node allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoding unit (155) and signaled to the image decoding device. If the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, it may be further split into one or more of the BT structure or the TT structure. There may be multiple splitting directions in the BT structure and / or the TT structure. For example, there may be two directions in which the block of the corresponding node is split horizontally and two directions in which the block is split vertically. As illustrated in FIG. 2, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether nodes have been split, and if splitting has occurred, a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (Binary or Ternary) are encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0045] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes of a lower layer, a CU split flag (split_cu_flag) indicating whether the node is split may be encoded. If the CU split flag (split_cu_flag) value indicates that the node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU (coding unit), which is a basic unit of encoding. If the CU split flag (split_cu_flag) value indicates that the node is split, the video encoding device starts encoding from the first flag in the above-described manner.

[0046] As another example of a tree structure, when QTBT is used, there may be two types: a type that horizontally splits the block of the corresponding node into two blocks of the same size (i.e., symmetric horizontal splitting) and a type that vertically splits it (i.e., symmetric vertical splitting). A split flag (split_flag) indicating whether each node of the BT structure is split into blocks of a lower layer and split type information indicating the type of split are encoded by the entropy encoding unit (155) and transmitted to the image decoding device. Meanwhile, there may additionally be a type that splits the block of the corresponding node into two blocks of an asymmetrical shape. The asymmetric shape may include a shape that splits the block of the corresponding node into two rectangular blocks with a size ratio of 1:3, or a shape that splits the block of the corresponding node in a diagonal direction.

[0047] A CU can have various sizes depending on the QTBT or QTBTTT partitioning from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block." Depending on the QTBTTT partitioning employed, the current block may be rectangular as well as square.

[0048] The prediction unit (120) predicts the current block and generates a prediction block. The prediction unit (120) includes an intra prediction unit (122) and an inter prediction unit (124).

[0049] In general, each current block within a picture can be predictively coded. Prediction of the current block can typically be performed using either intra-prediction (using data from the picture containing the current block) or inter-prediction (using data from a picture coded before the picture containing the current block). Inter-prediction encompasses both unidirectional and bidirectional prediction.

[0050] The intra prediction unit (122) predicts pixels within the current block using pixels (reference pixels) located around the current block within the current picture including the current block. There are multiple intra prediction modes depending on the prediction direction. For example, as shown in Fig. 3a, the multiple intra prediction modes may include two non-directional modes including the Planar mode and the DC mode, and 65 directional modes. The surrounding pixels to be used and the calculation formula are defined differently depending on each prediction mode.

[0051] For efficient directional prediction for a rectangular current block, directional modes (intra prediction modes 67 to 80 and -1 to -14) indicated by dotted arrows in Fig. 3b may be additionally used. These may be referred to as "wide-angle intra-prediction modes." In Fig. 3b, the arrows point to corresponding reference samples used for prediction, and do not indicate the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. Wide-angle intra-prediction modes are modes that perform prediction in the opposite direction of a specific directional mode without additional bit transmission when the current block is rectangular. At this time, among the wide-angle intra-prediction modes, some wide-angle intra-prediction modes available for the current block may be determined based on the ratio of the width and height of the rectangular current block. For example, wide-angle intra prediction modes (intra prediction modes 67 to 80) having an angle less than 45 degrees are available when the current block is a rectangular shape whose height is smaller than its width, and wide-angle intra prediction modes (intra prediction modes -1 to -14) having an angle greater than -135 degrees are available when the current block is a rectangular shape whose width is larger than its height.

[0052] The intra prediction unit (122) can determine an intra prediction mode to be used to encode the current block. In some examples, the intra prediction unit (122) can encode the current block using multiple intra prediction modes and select an appropriate intra prediction mode to be used from the tested modes. For example, the intra prediction unit (122) can calculate bit-rate distortion values ​​using rate-distortion analysis for multiple tested intra prediction modes and select an intra prediction mode with the best bit-rate distortion characteristics among the tested modes.

[0053] The intra prediction unit (122) selects one intra prediction mode from among multiple intra prediction modes and predicts the current block using surrounding pixels (reference pixels) and an operation formula determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0054] The inter prediction unit (124) generates a prediction block for the current block using a motion compensation process. The inter prediction unit (124) searches for a block most similar to the current block within reference pictures that were encoded and decoded before the current picture, and generates a prediction block for the current block using the searched block. Then, a motion vector (MV) corresponding to the displacement between the current block within the current picture and the prediction block within the reference picture is generated. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The motion information including information on the reference picture used to predict the current block and information on the motion vector is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0055] The inter prediction unit (124) may perform interpolation on a reference picture or a reference block to improve prediction accuracy. That is, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including the two integer samples. When a process of searching for a block most similar to the current block is performed on the interpolated reference picture, the motion vector can be expressed up to a precision in decimal units rather than a precision in integer sample units. The precision or resolution of the motion vector can be set differently for each target region to be encoded, such as a slice, tile, CTU, CU, etc. When such adaptive motion vector resolution (AMVR) is applied, information on the motion vector resolution to be applied to each target region must be signaled for each target region. For example, when the target region is a CU, information on the motion vector resolution applied to each CU is signaled. Information on the motion vector resolution may be information indicating the precision of a differential motion vector, which will be described later.

[0056] Meanwhile, the inter prediction unit (124) can perform inter prediction using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors indicating the block position most similar to the current block within each reference picture are used. The inter prediction unit (124) selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively, and searches for a block similar to the current block within each reference picture to generate a first reference block and a second reference block. Then, the first reference block and the second reference block are averaged or weighted averaged to generate a prediction block for the current block. Then, motion information including information on two reference pictures used to predict the current block and information on two motion vectors is transmitted to the entropy encoding unit (155). Here, reference picture list 0 may be composed of pictures that are before the current picture in display order among the restored pictures, and reference picture list 1 may be composed of pictures that are after the current picture in display order among the restored pictures. However, this is not necessarily limited to this, and restored pictures that are after the current picture in display order may be additionally included in reference picture list 0, and conversely, restored pictures that are before the current picture may be additionally included in reference picture list 1.

[0057] Various methods can be used to minimize the number of bits required to encode motion information.

[0058] For example, if the reference picture and motion vector of the current block are identical to those of a neighboring block, the motion information of the current block can be transmitted to the image decoding device by encoding information that can identify the neighboring block. This method is called 'merge mode.'

[0059] In merge mode, the inter prediction unit (124) selects a predetermined number of merge candidate blocks (hereinafter referred to as 'merge candidates') from the surrounding blocks of the current block.

[0060] As the surrounding blocks for deriving merge candidates, all or part of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (B2) adjacent to the current block within the current picture may be used, as illustrated in FIG. 4. In addition, a block located within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as a merge candidate. For example, a block co-located with the current block within the reference picture or blocks adjacent to the block at the co-located block may be additionally used as a merge candidate. If the number of merge candidates selected by the method described above is less than a preset number, a 0 vector is added to the merge candidates.

[0061] The inter prediction unit (124) uses these surrounding blocks to construct a merge list containing a predetermined number of merge candidates. Among the merge candidates included in the merge list, the merge candidate to be used as motion information of the current block is selected and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoding unit (155) and transmitted to the video decoding device.

[0062] Merge Skip mode is a special case of merge mode. After quantization, when all transform coefficients for entropy encoding are close to zero, only neighboring block selection information is transmitted without transmitting residual signals. By utilizing merge skip mode, relatively high encoding efficiency can be achieved for low-motion images, still images, and screen content images.

[0063] Hereinafter, merge mode and merge skip mode are collectively referred to as merge / skip mode.

[0064] Another method for encoding motion information is Advanced Motion Vector Prediction (AMVP) mode.

[0065] In AMVP mode, the inter prediction unit (124) derives predicted motion vector candidates for the motion vector of the current block using neighboring blocks of the current block. As neighboring blocks used to derive predicted motion vector candidates, all or some of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (B2) adjacent to the current block in the current picture as shown in FIG. 4 may be used. In addition, a block located in a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as the neighboring block used to derive predicted motion vector candidates. For example, a block co-located with the current block in the reference picture or blocks adjacent to the block in the co-located block may be used. If the number of motion vector candidates is less than a preset number by the method described above, a 0 vector is added to the motion vector candidates.

[0066] The inter prediction unit (124) derives predicted motion vector candidates using the motion vectors of these surrounding blocks, and determines a predicted motion vector for the motion vector of the current block using the predicted motion vector candidates. Then, the predicted motion vector is subtracted from the motion vector of the current block to produce a differential motion vector.

[0067] The predicted motion vector can be obtained by applying a predefined function (e.g., median, mean, etc.) to the predicted motion vector candidates. In this case, the image decoding device also knows the predefined function. In addition, since the surrounding blocks used to derive the predicted motion vector candidates are blocks that have already been encoded and decoded, the image decoding device also already knows the motion vectors of the surrounding blocks. Therefore, the image encoding device does not need to encode information to identify the predicted motion vector candidates. Therefore, in this case, information about the differential motion vector and information about the reference picture used to predict the current block are encoded.

[0068] Alternatively, the predicted motion vector can be determined by selecting one of the predicted motion vector candidates. In this case, information for identifying the selected predicted motion vector candidate is additionally encoded, along with information about the differential motion vector and the reference picture used to predict the current block.

[0069] The subtractor (130) subtracts the prediction block generated by the intra prediction unit (122) or inter prediction unit (124) from the current block to generate a residual block.

[0070] The transformation unit (140) transforms residual signals within a residual block having pixel values ​​in a spatial domain into transform coefficients in a frequency domain. The transformation unit (140) may transform the residual signals within the residual block using the entire size of the residual block as a transformation unit, or may divide the residual block into a plurality of sub-blocks and use the sub-blocks as transformation units to perform the transformation. Alternatively, the residual signals may be transformed using only the transformation domain sub-block as a transformation unit by dividing the sub-blocks into two sub-blocks, that is, a transformation domain and a non-transform domain. Here, the transformation domain sub-block may be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicating that only a sub-block has been converted, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit (155) and signaled to the image decoding device. In addition, the size of the conversion area sub-block may have a size ratio of 1:3 with respect to the horizontal axis (or vertical axis), and in this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding division is additionally encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0071] Meanwhile, the transformation unit (140) can individually perform transformations on the residual block in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a Multiple Transform Set (MTS). The transformation unit (140) can select one transformation function pair with the best transformation efficiency among the MTS and transform the residual block in the horizontal and vertical directions, respectively. Information (mts_idx) on the transformation function pair selected among the MTS is encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0072] The quantization unit (145) quantizes the transform coefficients output from the transform unit (140) using quantization parameters and outputs the quantized transform coefficients to the entropy encoding unit (155). The quantization unit (145) may directly quantize a related residual block without transformation for a certain block or frame. The quantization unit (145) may also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients within the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions may be encoded and signaled to an image decoding device.

[0073] The rearrangement unit (150) can perform rearrangement of coefficient values ​​for quantized residual values.

[0074] The reordering unit (150) can change a two-dimensional coefficient array into a one-dimensional coefficient sequence by using coefficient scanning. For example, the reordering unit (150) can output a one-dimensional coefficient sequence by scanning from the DC coefficient to the coefficients of the high-frequency region by using a zig-zag scan or a diagonal scan. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficients in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, the scanning method to be used may be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.

[0075] The entropy encoding unit (155) generates a bitstream by encoding a sequence of one-dimensional quantized transform coefficients output from the rearrangement unit (150) using various encoding methods such as CABAC (Context-based Adaptive Binary Arithmetic Code) and Exponential Golomb.

[0076] In addition, the entropy encoding unit (155) encodes information related to block division, such as CTU size, CU division flag, QT division flag, MTT division type, and MTT division direction, so that the image decoding device can divide the block in the same manner as the image encoding device. In addition, the entropy encoding unit (155) encodes information about a prediction type indicating whether the current block is encoded by intra prediction or inter prediction, and encodes intra prediction information (i.e., information about an intra prediction mode) or inter prediction information (information about an encoding mode of motion information (merge mode or AMVP mode), a merge index in the case of a merge mode, and a reference picture index and a differential motion vector in the case of an AMVP mode) according to the prediction type. In addition, the entropy encoding unit (155) encodes information related to quantization, that is, information about a quantization parameter and information about a quantization matrix.

[0077] The inverse quantization unit (160) inversely quantizes the quantized transform coefficients output from the quantization unit (145) to generate transform coefficients. The inverse transform unit (165) transforms the transform coefficients output from the inverse quantization unit (160) from the frequency domain to the spatial domain to restore the residual block.

[0078] An adder (170) adds the restored residual block and the predicted block generated by the prediction unit (120) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting the next block.

[0079] The loop filter unit (180) performs filtering on restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transformation / quantization. The loop filter unit (180) may include all or part of a deblocking filter (182), a sample adaptive offset (SAO) filter (184), and an adaptive loop filter (ALF, 186) as an in-loop filter.

[0080] The deblocking filter (182) filters the boundaries between restored blocks to remove blocking artifacts caused by block-based encoding / decoding, and the SAO filter (184) and the ALF (186) perform additional filtering on the deblocking-filtered image. The SAO filter (184) and the ALF (186) are filters used to compensate for the differences between restored pixels and original pixels caused by lossy coding. The SAO filter (184) improves not only subjective image quality but also encoding efficiency by applying an offset in units of CTUs. In contrast, the ALF (186) performs block-based filtering, and compensates for distortion by applying different filters by distinguishing the edges and degrees of variation of the corresponding block. Information on filter coefficients to be used in the ALF can be encoded and signaled to an image decoding device.

[0081] The restored blocks filtered through the deblocking filter (182), SAO filter (184), and ALF (186) are stored in the memory (190). When all blocks within a picture are restored, the restored picture can be used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0082] The video encoding device can store the bitstream of encoded video data on a non-transitory storage medium or transmit it to the video decoding device using a communication network.

[0083] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image decoding device and its subcomponents will be described with reference to FIG. 5.

[0084] The video decoding device may be configured to include an entropy decoding unit (510), a rearrangement unit (515), an inverse quantization unit (520), an inverse transformation unit (530), a prediction unit (540), an adder (550), a loop filter unit (560), and a memory (570).

[0085] Similar to the video encoding device of FIG. 1, each component of the video decoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0086] The entropy decoding unit (510) decodes the bitstream generated by the image encoding device to extract information related to block division, thereby determining the current block to be decoded, and extracts prediction information, information on residual signals, etc. required to restore the current block.

[0087] The entropy decoding unit (510) extracts information about the CTU size from the Sequence Parameter Set (SPS) or the Picture Parameter Set (PPS), determines the size of the CTU, and divides the picture into CTUs of the determined size. Then, the CTU is determined as the top layer of the tree structure, i.e., the root node, and the CTU is divided using the tree structure by extracting division information about the CTU.

[0088] For example, when splitting a CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, the second flag (mtt_split_flag) related to the splitting of MTT and the split direction (vertical / horizontal) and / or split type (binary / ternary) information are extracted, and the corresponding leaf nodes are split into the MTT structure. Accordingly, each node below the leaf nodes of QT are split recursively into the BT or TT structure.

[0089] As another example, when splitting a CTU using the QTBTTT structure, the CU split flag (split_cu_flag) indicating whether the CU is split is first extracted, and if the block is split, the first flag (QT_split_flag) may be extracted. During the splitting process, each node may undergo zero or more repeated QT splits followed by zero or more repeated MTT splits. For example, a CTU may undergo an MTT split right away, or conversely, may undergo only multiple QT splits.

[0090] As another example, when splitting a CTU using the QTBT structure, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Furthermore, for nodes corresponding to leaf nodes of QT, a split flag (split_flag) indicating whether to further split into BTs and splitting direction information are extracted.

[0091] Meanwhile, when the entropy decoding unit (510) determines the current block to be decoded by using the division of the tree structure, it extracts information on the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-prediction, the entropy decoding unit (510) extracts syntax elements for intra-prediction information (intra-prediction mode) of the current block. If the prediction type information indicates inter-prediction, the entropy decoding unit (510) extracts syntax elements for inter-prediction information, i.e., information indicating a motion vector and a reference picture referenced by the motion vector.

[0092] Additionally, the entropy decoding unit (510) extracts information about the quantized transform coefficients of the current block as information related to quantization and information about residual signals.

[0093] The rearrangement unit (515) can change the sequence of one-dimensional quantized transform coefficients entropy-decoded in the entropy decoding unit (510) back into a two-dimensional coefficient array (i.e., block) in the reverse order of the coefficient scanning performed by the image encoding device.

[0094] The inverse quantization unit (520) inversely quantizes the quantized transform coefficients and inversely quantizes the quantized transform coefficients using the quantization parameters. The inverse quantization unit (520) may also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in two dimensions. The inverse quantization unit (520) may perform inverse quantization by applying a matrix of quantized coefficients (scaling values) from an image encoding device to a two-dimensional array of quantized transform coefficients.

[0095] The inverse transform unit (530) inversely transforms the inverse quantized transform coefficients from the frequency domain to the spatial domain to restore residual signals, thereby generating a residual block for the current block.

[0096] In addition, when the inverse transform unit (530) inversely transforms only a portion of a transform block (sub-block), it extracts a flag (cu_sbt_flag) indicating that only a sub-block of the transform block has been transformed, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block, and inversely transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to restore residual signals, and fills “0” values ​​with residual signals for areas that have not been inversely transformed, thereby generating a final residual block for the current block.

[0097] In addition, when MTS is applied, the inverse transform unit (530) determines a transform function or a transform matrix to be applied in the horizontal and vertical directions using MTS information (mts_idx) signaled from the image encoding device, and performs inverse transform on the transform coefficients within the transform block in the horizontal and vertical directions using the determined transform function.

[0098] The prediction unit (540) may include an intra prediction unit (542) and an inter prediction unit (544). The intra prediction unit (542) is activated when the prediction type of the current block is intra prediction, and the inter prediction unit (544) is activated when the prediction type of the current block is inter prediction.

[0099] The intra prediction unit (542) determines the intra prediction mode of the current block among a plurality of intra prediction modes from the syntax elements for the intra prediction mode extracted from the entropy decoding unit (510), and predicts the current block using reference pixels around the current block according to the intra prediction mode.

[0100] The inter prediction unit (544) uses the syntax elements for the inter prediction mode extracted from the entropy decoding unit (510) to determine the motion vector of the current block and the reference picture referenced by the motion vector, and predicts the current block using the motion vector and the reference picture.

[0101] An adder (550) adds the residual block output from the inverse transform unit (530) and the predicted block output from the inter prediction unit (544) or the intra prediction unit (542) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting a block to be decoded later.

[0102] The loop filter unit (560) may include a deblocking filter (562), an SAO filter (564), and an ALF (566) as in-loop filters. The deblocking filter (562) deblocks the boundaries between restored blocks to remove blocking artifacts caused by block-by-block decoding. The SAO filter (564) and the ALF (566) perform additional filtering on restored blocks after deblocking filtering to compensate for differences between restored pixels and original pixels caused by lossy coding. The filter coefficients of the ALF are determined using information about filter coefficients decoded from the non-stream.

[0103] The restored blocks filtered through the deblocking filter (562), SAO filter (564), and ALF (566) are stored in the memory (570). When all blocks within a picture are restored, the restored picture is used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0104] The present embodiment relates to encoding and decoding of images (videos) as described above. More specifically, the present invention provides a video coding method and device for deriving an intra prediction mode based on an inter prediction signal when generating a final prediction signal of a current block by weighting an inter prediction signal and an intra prediction signal.

[0105] The following embodiments may be performed by a prediction unit (120) within a video encoding apparatus. Additionally, they may be performed by a prediction unit (540) within a video decoding apparatus.

[0106] The video encoding device can generate signaling information related to the present embodiment in terms of rate distortion optimization in encoding the current block. The video encoding device can encode the signaling information using the entropy encoding unit (155) and then transmit it to the video decoding device. The video decoding device can decode the signaling information related to the decoding of the current block from the bitstream using the entropy decoding unit (510).

[0107] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU). Alternatively, the term "target block" may also refer to a portion of a coding unit.

[0108] Also, a value of a flag being true indicates that the flag is set to 1. Also, a value of a flag being false indicates that the flag is set to 0.

[0109] I-1. Merge / Skip Mode and MMVD in Inter Prediction

[0110] Hereinafter, a method for constructing a merge candidate list of motion information in the merge / skip mode of inter prediction is described. To support the merge / skip mode, a video encoding device can construct a merge candidate list by selecting a preset number of merge candidates (e.g., 6).

[0111] The video encoding device searches for spatial merge candidates. The video encoding device searches for spatial merge candidates from surrounding blocks, as illustrated in FIG. 4. Up to four spatial merge candidates can be selected.

[0112] A video encoding device searches for temporal merge candidates. The video encoding device may add a co-located block as a temporal merge candidate, which is a block located in the same location as the current block within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the target block is located. Only one temporal merge candidate may be selected.

[0113] A video encoding device searches for HMVP (History-based Motion Vector Predictor) candidates. The video encoding device can store the motion vectors of the previous h CUs (where h is a natural number) in a table and use them as merge candidates. The table size is 6, and the motion vectors of the previous CUs are stored in a First-in First Out (FIFO) manner. This indicates that up to 6 HMVP candidates are stored in the table. The video encoding device can set the most recent motion vectors among the HMVP candidates stored in the table as merge candidates.

[0114] The video encoding device searches for PAMVP (Pairwise Average MVP) candidates. The video encoding device can set the motion vector average of the first and second candidates in the merge candidate list as the merge candidate.

[0115] If the merge candidate list cannot be filled even after performing all of the above-described search processes (i.e., if the preset number cannot be filled), the video encoding device adds a zero motion vector as a merge candidate.

[0116] In terms of optimizing encoding efficiency, a video encoding device can determine a merge index that indicates one candidate within a merge candidate list. The video encoding device can then derive a motion vector predictor (MVP) from the merge candidate list using the merge index, and then determine the MVP as the motion vector of the current block. Furthermore, the video encoding device can signal the merge index to a video decoding device.

[0117] The video encoding device uses the same motion vector transmission method as the merge mode in the skip mode, but does not transmit the residual block corresponding to the difference between the current block and the predicted block.

[0118] The method for constructing the aforementioned merge candidate list can be performed in the same manner by a video decoding device. The video decoding device can decode the merge index. The video decoding device can then derive an MVP from the merge candidate list using the merge index, and then determine the MVP as the motion vector of the current block.

[0119] Meanwhile, when using MMVD (Merge mode with Motion Vector Difference) technology, the video encoding device can derive an MVP from the merge candidate list using a merge index. For example, the first or second candidate in the merge candidate list can be used as the MVP. In addition, in terms of optimizing encoding efficiency, the video encoding device determines a distance index and a direction index. The video encoding device can derive a motion vector difference (MVD) using the distance index and the direction index, and then reconstruct the motion vector of the current block by adding the MVD and the MVP. In addition, the video encoding device can signal the merge index, the direction index, and the direction index to the video decoding device.

[0120] The aforementioned MMVD technique can be performed in the same manner by the inter prediction unit (544) within the video decoding device. The video decoding device can decode a merge index, a distance index, and a direction index. The video decoding device can construct a merge candidate list and then derive an MVP from the merge candidate list using the merge index. The video decoding device can derive an MVD using the distance index and the direction index, and then reconstruct the motion vector of the current block by adding the MVD and the MVP.

[0121] I-2. CIIP (Combined Inter and Intra Prediction)

[0122] CIIP mode generates a final prediction signal by weighting the inter prediction signal and the intra prediction signal for the current block with a size of W×H.

[0123] When the CIIP mode is applied in VVC, an inter prediction signal is generated based on signaling / parsing of the merge index according to the regular merge mode, and an intra prediction signal is generated based on the Planar mode. As shown in Fig. 6, weights used for the weighted sum of the intra prediction signal and the inter prediction signal are determined based on the prediction techniques of the previously restored blocks on the left (L) and top (T) of the current block. When intra prediction is applied to both the left and top blocks, the weight ratio between the intra prediction signal and the inter prediction signal is determined as 3:1. When intra prediction is applied to one of the two blocks, the weight ratio is determined as 1:1, and when inter prediction is applied to both blocks, the weight ratio is determined as 1:3. The CIIP mode generates the final prediction signal based on the determined weights.

[0124] When the CIIP mode is applied in the ECM (Enhanced Compression Model) corresponding to Beyond VVC, the inter prediction signal is generated in the following order. The CIIP mode constructs a general merge candidate list of the current block, and the candidate list size is 2. For the candidates in the constructed list, the CIIP mode compensates for motion based on template matching. The CIIP mode reorders the two candidates by calculating the cost between the template of the compensated candidate and the template of the current block. The CIIP mode generates the inter prediction signal using the candidate with the lowest cost. The intra prediction signal is generated in the following order. The CIIP mode derives the intra prediction mode by applying the Template-based Intra Mode Derivation (TIMD) method to the Most Probable Mode (MPM) candidates using the template of the current block. The CIIP mode generates the intra prediction signal of the current block using the derived intra prediction mode.

[0125] After generating the inter prediction signal and the intra prediction signal according to the above-described process, the CIIP mode calculates weights based on the directionality of the derived intra prediction mode. For example, if the derived intra prediction mode is a horizontal directionality mode, the current block may be vertically divided into subblocks as in Fig. 7a, and if it is a vertical directionality mode, the current block may be horizontally divided into subblocks as in Fig. 7b. In Figs. 7a and 7b, numbers in the subblocks indicate subblock indices. The CIIP mode may determine the weight ratio of the intra prediction signal and the inter prediction signal for each subblock as in Table 1. The CIIP mode generates the final prediction signal based on the determined weights.

[0126]

[0127] w in Table 1 intra and w interrepresent the weight of the intra prediction signal and the weight of the inter prediction signal, respectively.

[0128] The following embodiments are described with a focus on an image decoding device, but can be implemented identically or similarly in an image encoding device.

[0129] II. Embodiments according to the present disclosure

[0130] FIG. 8 is a block diagram illustrating in detail a portion of an image decoding device according to one embodiment of the present disclosure.

[0131] The video decoding device according to the present embodiment determines prediction and transformation units, and performs prediction and inverse transformation on the current block corresponding to the determined unit using the determined prediction technique and prediction mode, thereby finally generating a restoration block of the current block. The example illustrated in FIG. 8 may be performed by the inverse transformation unit (530), the prediction unit (540), and the adder (550) of the video decoding device. Meanwhile, the same operations as the example illustrated in FIG. 8 may be performed by the inverse transformation unit (165), the picture division unit (110), the prediction unit (120), and the adder (170) of the video encoding device. At this time, the video decoding device uses encoding information parsed from the bitstream, but the video encoding device may use encoding information set from a higher level in terms of minimizing rate distortion. Hereinafter, for convenience, the present embodiment will be described with reference to the video decoding device.

[0132] As in the example of FIG. 5, the prediction unit (540) includes an intra prediction unit (542) and an inter prediction unit (544) depending on the prediction technology, but as illustrated in FIG. 8, the prediction unit (540) may include all or part of the prediction unit determination unit (802), the prediction technology determination unit (804), the prediction mode determination unit (806), and the prediction execution unit (808).

[0133] If the color format of the input video is a YUV format (such as YUV420, YUV411, YUV422, YUV444), the video decoding device can perform prediction and restoration of the chroma component after performing prediction and restoration of the luma component. That is, the luma component and the chroma component can be sequentially restored by the components illustrated in FIG. 8. Meanwhile, if the color format of the input video is RGB, the video encoding device can perform color format conversion from RGB to YUV and then encode the converted video. Here, in the case of the YUV format, the color format represents the correspondence between the pixels of the luma component and the pixels of the chroma component.

[0134] The prediction unit determination unit (802) determines a prediction unit (PU). The prediction technique determination unit (804) determines a prediction technique (e.g., intra prediction, inter prediction, IBC (Intra Block Copy) mode, palette mode, etc.) for the prediction unit. The prediction mode determination unit (806) determines a detailed prediction mode for the prediction technique. The prediction execution unit (808) generates a prediction block of the current block according to the determined prediction mode.

[0135] The inverse transform unit (530) includes a transform unit determination unit (810) and an inverse transform execution unit (812). The transform unit determination unit (810) determines a transform unit (TU) for the inverse quantization signals of the current block, and the inverse transform execution unit (812) inversely transforms the transform unit expressed as the inverse quantization signals to generate residual signals.

[0136] An adder (550) adds the prediction block and residual signals to generate a restoration block. The restoration block is stored in memory and can be used to predict other blocks.

[0137] The prediction unit determined by the prediction unit determination unit (802) may be the current block or one of the sub-blocks into which the current block is divided. At this time, the prediction unit of the chroma component may have a size corresponding to the prediction unit of the luma component according to the color format. Alternatively, after the prediction units of the luma component and the chroma component are determined separately, prediction may be performed on the prediction unit of the chroma component.

[0138] The prediction technology determination unit (804) determines the prediction technology for each prediction unit. As described above, the prediction technology may be one of inter-prediction, intra-prediction, IBC mode, and palette mode. In this case, the prediction technology for the chroma component can be determined in the same manner as the prediction technology for the corresponding luma component, without separate signaling or parsing of information.

[0139] For example, if the prediction technique of the current block is not intra prediction, the video decoding device parses 1-bit flag information. For example, if the parsed flag indicates Skip mode, the video decoding device determines the prediction mode of the current block as the merge mode of inter prediction or the IBC merge mode. In the case of Skip mode, the video decoding device can use the prediction signals as restored signals without performing the inverse transformation process (i.e., without parsing the residual signals).

[0140] On the other hand, if the parsed flag does not indicate a Skip mode for the current block, the prediction technique determination unit (804) can parse a series of 1-bit flags to determine the prediction technique of the current block as one of techniques such as inter prediction, intra prediction, IBC mode, palette mode, etc.

[0141] For example, if Skip is not applied to the current block and the prediction technique is determined to be Inter-Prediction or IBC mode, the video decoding device parses a 1-bit flag. Depending on the parsed flag, the prediction mode of the current block can be determined as either General Merge mode or Advanced Motion Vector Prediction (AMVP) mode.

[0142] The prediction mode determination unit (806) determines a detailed prediction mode for the prediction technology.

[0143] For example, if the prediction technology of the current block is inter prediction, the prediction mode determining unit (806) can determine the general merge mode or AMVP mode as the prediction mode of the current block. In the general merge mode or AMVP mode, the image decoding device generates prediction blocks according to one or more motion compensations based on parsed motion information, and weights and combines the generated multiple prediction blocks to generate final prediction signals of the current block.

[0144] As another example, if the prediction technique of the current block is inter prediction, the prediction mode determining unit (806) may determine Geometric Partitioning Mode (GPM) as the prediction mode of the current block. In GPM, the image decoding device divides the current block into two or more partition blocks according to geometric partitioning, generates prediction blocks according to one or more motion compensations based on motion information of the parsed current block, and weights and combines the generated multiple prediction blocks to generate final prediction signals of the current block. For example, as described above, the current block may be divided into two partition blocks.

[0145] The prediction execution unit (808) generates a prediction block of the current block according to the determined prediction technology and prediction mode.

[0146] As an example, the prediction performing unit (808) generates a prediction block of the current block according to the prediction mode, and the adder (550) adds the prediction block of the current block and residual signals to generate a restoration block.

[0147] Meanwhile, the entropy decoding unit (510) can restore the quantized second-order transform coefficients when the second-order transform is applied. On the other hand, when the second-order transform is not applied, the entropy decoding unit (510) can restore the quantized first-order transform coefficients. The inverse quantization unit (520) can apply inverse quantization to the restored transform coefficients based on the quantization parameter to generate inverse quantized transform coefficients.

[0148] The transformation unit determination unit (810) of the inverse transformation unit (530) determines a transformation unit (TU) for the inverse quantized transformation coefficients. At this time, if one TU is divided into multiple sub-blocks, one sub-block can be used as a TU.

[0149] The inverse transform unit (530) can determine a non-separable second-order inverse transform kernel and a separable first-order inverse transform kernel when a second-order transform is applied. On the other hand, when the second-order transform is not applied, the inverse transform unit (530) can determine a separable first-order inverse transform kernel or a non-separable first-order inverse transform kernel. The inverse transform unit (530) can perform an inverse transform on the inverse quantized transform coefficients using the inverse transform kernel. The inverse transform unit (530) can determine whether to perform a non-separable first-order inverse transform and a non-separable second-order inverse transform based on signaling / parsing, or can implicitly determine the same based on information such as the size of the current TU, aspect ratio, etc.

[0150] Below, the operation of the prediction mode determination unit (806) is described in detail with respect to intra prediction mode derivation.

[0151] For example, if the prediction technique of the current block is inter prediction and an intra prediction signal is used to generate final prediction signals of the current block, an image decoding device can define a surrounding pre-reconstructed region of the current block as a template and derive an intra prediction mode using the template. The template may also include a region that is not adjacent to the current block. The image decoding device can generate final prediction signals of the current block based on the derived intra prediction mode. For example, the image decoding device can generate final prediction signals by weighting an intra prediction signal determined according to derivation and / or parsing and an inter prediction signal determined according to derivation and / or parsing. At this time, the weights used in the weighted sum may be different depending on each sample position. For example, the weights may be adaptively determined based on prediction information of the current block and information of the surrounding pre-reconstructed region, an aspect ratio, etc.

[0152] As an example, if the prediction technique of the current block is intra prediction, the prediction mode of the current block can be one directional prediction mode, a planar mode (Horizontal Planar or Vertical Planar or Regular Planar), or a DC mode.

[0153] As another example, if the prediction technique of the current block is intra prediction, the image decoding device may determine an intra geometric segmentation-based prediction mode as the prediction mode of the current block. The current block may be divided into one or more sub-regions according to the geometric segmentation, and a prediction mode within the intra prediction technique, including different directional prediction modes, planar modes, or DC modes, may be used for each region. The image decoding device may generate final prediction signals of the current block by weighting and adding the prediction signals of each region.

[0154] As another example, if the prediction technique of the current block is intra prediction, the video decoding device may determine a matrix-based intra prediction mode as the prediction mode of the current block. According to an agreement between the video encoding device and the video decoding device, the matrix index is signaled / parsed based on a predefined matrix, or the matrix itself is signaled / parsed. The video decoding device may generate prediction signals of the current block based on the matrix index or matrix.

[0155] As another example, if the prediction technique for the current block is intra prediction, the image decoding device may determine the intra template matching prediction mode as the prediction mode for the current block. The image decoding device may define a restoration area surrounding the current block as a template and perform template matching in the defined area, thereby generating prediction signals for the current block.

[0156] As another example, if the prediction technique for the current block is intra prediction, the image decoding device can define the surrounding restoration region of the current block as a template and derive an intra prediction mode using the template. The derived mode can be used to generate the final prediction signals for the current block. The template can also include regions not adjacent to the current block.

[0157] There are various ways to derive the intra prediction mode of the current block. Among the numerous methods, flags and / or indices can be signaled / parsed to indicate one or more methods. Alternatively, a specific prediction mode can be used in a fixed manner.

[0158] Below, the operation of the prediction execution unit (808) is described in detail with respect to the generation of final prediction signals.

[0159] As an example, a case is described where the prediction mode of the current block is determined to be a mode that generates final prediction signals based on a weighted sum of intra-prediction signals and inter-prediction signals. The aforementioned mode can be determined based on the signaling / parsing of a 1-bit flag. The intra-prediction signal and inter-prediction signal of the current block can be generated from derived information or based on the signaling / parsing of some information. The weights used in the weighted sum can be fixed values. The weights can be implicitly determined using information such as each prediction mode, the prediction signal, information on a restored area around the current block, the size and / or aspect ratio of the current block, etc.

[0160] As an example, an inter prediction signal can be generated as follows. For example, a video decoding device can construct a motion vector candidate list at predetermined locations around and not adjacent to a current block. The video decoding device can construct the candidate list using adjacent spatial motion vector candidates (A0, A1, B0, B1, B2 in FIG. 9a), non-adjacent spatial motion vector candidates (1 to 18 in FIG. 9b), adjacent temporal motion vector candidates (C0, C1 in FIG. 9a), a Pairwise Motion Vector Predictor (PAMVP), a History-based Motion Vector Predictor (HMVP), and a zero motion vector. When constructing the candidate list, an order determined according to an agreement between a video encoding device and a video decoding device can be utilized. Here, according to the agreement between a video encoding device and a video decoding device, the maximum number of candidates included in the candidate list can be n (n is an integer greater than or equal to 1). In Fig. 9b, w and h represent the width and height of the current block, respectively.

[0161] FIG. 10 is an exemplary diagram showing an area where template-based motion compensation is performed according to one embodiment of the present disclosure.

[0162] As an example, a 1-bit flag indicating the use of template matching may be signaled / parsed. When the aforementioned flag is 1, the video decoding device may perform template-based motion compensation on each candidate in the motion vector candidate list, as shown in FIG. 10. In FIG. 10, -b to +b on the x-axis and -a to +a on the y-axis are search ranges for template-based motion compensation. a and b may be fixed values ​​and / or values ​​determined according to the aspect ratio of the current block. For template-based motion compensation, a cost function such as mean squared error (MSE), sum of absolute error (SAE), etc. between the template of the current block and the template of the reference block may be used. The video decoding device may perform motion compensation in a direction that reduces the cost function. Motion compensation may be performed at the location of a sampled pixel within the search range.

[0163] As an example, a video decoding device may, for each candidate subjected to motion compensation, calculate a cost between the template of the predicted block and the template of the current block based on the predicted block according to motion compensation. The video decoding device may rearrange the candidates in the candidate list in ascending or descending order based on the calculated cost.

[0164] As an example, the video decoding device may generate an inter prediction signal using the 0th motion vector candidate from the rearranged candidate list. Alternatively, the video decoding device may generate an inter prediction signal using a motion vector candidate determined by signaling / parsing an index or flag from the rearranged candidate list. In this case, as described above, template-based motion compensation is applied in advance to the motion vector candidate in the candidate list.

[0165] Alternatively, if the 1-bit flag indicating the use of template matching is 1, the video decoding device can generate an inter prediction signal using a motion vector candidate determined by signaling / parsing an index or flag from a motion vector candidate list constructed based on surrounding information of the current block. At this time, template-based motion compensation is applied in advance to the motion vector candidates in the candidate list, as described above.

[0166] As an example, an intra prediction signal can be generated as follows. For example, a video decoding device can generate an intra prediction signal using an intra prediction mode determined according to signaling / parsing of prediction mode information. The video decoding device can construct a most probable mode (MPM) list based on information of a pre-restored area surrounding a current block. According to an embodiment, the MPM list can include a PMPM (Primary MPM) list and an SMPM (Secondary MPM) list. The video decoding device can construct the MPM list based on the prediction mode of a pre-restored area surrounding a current block. A 1-bit flag indicating whether to use a prediction mode in the MPM list can be signaled / parsed. If the MPM list consists of the PMPM list and the SMPM list and the MPM flag is 1, a 1-bit flag may be signaled / parsed to indicate whether to signal / parse an index for a mode in the PMPM list or to signal / parse an index for a mode in the SMPM list. If the MPM flag is 0, the intra prediction mode of the current block may be determined by signaling / parsing an index indicating one of the intra prediction modes other than the intra prediction modes in the MPM list.

[0167] As an example, a video decoding device can derive the intra prediction mode of the current block. At this time, one or more of various methods can be used as the derivation method, or a fixed method can be used. Alternatively, the derivation method can be determined using signaling / parsing.

[0168] Below, the TIMD (Template-based Intra Mode Derivation) method is described.

[0169] An image decoding device can derive an intra prediction mode of a current block using a surrounding restored template of the current block and reference lines of the template. As shown in Fig. 11, the template of the current block is the restored regions on the top and left, and the reference lines of the template can be defined as one or more lines. In Fig. 11, b, c, d, e, and f can be integers greater than or equal to 1, can be fixed values, and can vary depending on the size and aspect ratio of the current block.

[0170] FIG. 12 is an exemplary diagram illustrating determination of an intra prediction mode candidate according to the location of a sample according to one embodiment of the present disclosure.

[0171] As an example, the candidates for the prediction mode to be used in TIMD may be all available intra directional prediction modes of the current block, DC mode, Planar mode, etc. Alternatively, some intra prediction modes may be candidates for the prediction mode based on the prediction modes of the reconstructed surrounding regions of the current block (e.g., the left region and the upper region). For example, the intra prediction mode candidates may be determined according to the positions of the available samples of the reconstructed region of the current block. As shown in Fig. 12, the availability of samples may be checked at positions A, B, C, and D based on the current block, and candidate modes may be added according to the aspect ratio of the current block. At this time, the values ​​of a and b may be determined according to the size of the current block.

[0172] As an example, for a current block of size W×H, if W / H=1 and A samples are available, some modes (e.g., 67 or more) with an angle smaller than the intra directional prediction mode with an angle of 45° (e.g., mode 66) can be added as candidates. If B samples are available, modes with an angle up to the available minimum angle (e.g., mode 80) can be added as candidates. The above can also be applied to C and D samples. In addition, if the value of W / H is not 1, the application of candidates for wide angles (e.g., intra directional prediction modes with an angle smaller than 45° or greater than 225°) can be restricted depending on the availability of A, B, C, and D position samples.

[0173] After determining candidates for intra prediction modes, the video decoding device can predict templates for each candidate prediction mode using reference lines of the templates, and determine the order of the prediction modes based on the cost between the predicted signal and the template. The reference lines may be multiple lines within the reference line area. For example, a group of one reference line and one prediction mode may be set as one candidate. As shown in Fig. 13, if the available reference line is the reference line of (1, 5) and the number of intra prediction mode candidates is 3, six candidate groups may be configured. The video decoding device can predict templates for the six candidate groups and rearrange the prediction modes in ascending or descending order based on the cost between the predicted signal and the template. The video decoding device can determine the intra prediction mode based on the signaling / parsing of the index. Alternatively, the video decoding device can determine the candidate and reference line of the 0th index among the rearranged candidates as the intra prediction mode of the current block.

[0174] Below, the DIMD (Decoder-side Intra Mode Derivation) method is described.

[0175] FIG. 14 is an exemplary diagram showing a restored template surrounding a current block according to one embodiment of the present disclosure.

[0176] An image decoding device can derive an intra prediction mode of a current block using a surrounding pre-reconstructed template of the current block. As shown in Fig. 14, a surrounding pre-reconstructed region of a current block having a size of W×H can be defined as a template. In Fig. 14, m and n can be integers greater than or equal to 1. If the upper right region of the current block is available (③ of Fig. 14), a template can be defined including the upper right region, and if the lower left region of the current block is available (② of Fig. 14), a template can be defined including the lower left region. If both the upper right and lower left regions are available (④ of Fig. 14), a template can be defined including the upper right and lower left regions, and if neither region is available (① of Fig. 14), a template can be defined excluding both regions.

[0177] In Fig. 14, a and b can be integers greater than or equal to 1. a can range from 0 to min (the length of the maximum possible used area, W or 2W), and b can range from 0 to min (the length of the maximum possible used area, H or 2H). If part or all of the left area of ​​the current block in the restored surrounding area of ​​the current block is unusable, only the upper template can be used. Alternatively, if part or all of the upper area of ​​the restored surrounding area of ​​the current block is unusable, only the left template can be used. As another example, templates can be defined excluding the upper left position of the current block.

[0178] When the template of the current block is defined as in the process described above, the image decoding device can derive an intra prediction mode using the directionality of the template. For example, in FIG. 14, when m and n are 3 or more, the image decoding device can calculate the gradient of p×q blocks in all areas of the template, and derive the intra prediction modes corresponding to the most dominant gradient order based on the calculated gradient. p and q can be integers 2 or more. In FIG. 14, when m is less than p or n is less than q, the image decoding device can perform the process described above after copying the outermost pixel values ​​in the template so that m and n become p and q.

[0179] For example, for a template as described above, an image decoding device may calculate a gradient. The image decoding device may calculate the gradient after filtering the template. Alternatively, the image decoding device may calculate the gradient for a portion of the template. For example, a smoothing filter having coefficients of [1 / 4, 2 / 4, 1 / 4] may be utilized for filtering the template.

[0180] When calculating the gradient for a portion of a template, as shown in Fig. 15, the image decoding device can calculate the gradient of a p×p block (p=3 in Fig. 15) containing each u pixel (u=2 in Fig. 15). u and / or p can be adaptively determined based on information such as the size of the current block, the resolution of the current image, the aspect ratio of the current block, etc.

[0181] For example, when calculating directionality (gradient) for the location of a pixel sampled from a surrounding restored template on a block-by-block basis, the sampling location can be determined by various methods. As in Fig. 15, the sampling location can be set at the same pixel distance, so that the gradient can be calculated at the sampled location. As in Fig. 16, the sampling location can be set using segmentation and prediction information of the surrounding restored template of the current block. For example, with respect to a boundary where block division is performed, an image decoding device can calculate the gradient for a non-adjacent area to the block division boundary without calculating directionality for a certain pixel.

[0182] As an example, an image decoding device may calculate a gradient for a block unit for a surrounding restored template, and then store the calculated gradient in a table. The image decoding device may map a directional mode available for prediction of the current block to each calculated gradient, and store the strength of the gradient and the corresponding prediction mode together in a table (hereinafter, “mapping table”). The strength of the gradient may be an absolute value of the gradient. The strength of the gradient may be the sum of the absolute value of the gradient in the x direction and the absolute value of the gradient in the y direction. The directional mode available for the current block may vary depending on the aspect ratio of the current block, the range of the available area of ​​the restored template, etc. As an example, the image decoding device first configures a table (hereinafter, “directional table”) for storing the calculated gradient. Thereafter, the image decoding device may determine the intra prediction mode of the current block by corresponding (mapping) the most dominant gradient in the directional table to a directional mode available for prediction of the current block.

[0183] For example, for a surrounding restored template, the available modes of the current block may be determined differently depending on the range of available templates. Furthermore, the available modes of the current block may be determined by considering both the aspect ratio of the current block and / or the available area of ​​the templates in the current block.

[0184] When constructing a table by calculating slopes block-by-block for the surrounding restored templates, the image decoding device defines the available directional prediction modes for the current block. Thereafter, the image decoding device can restrictively add some of the available directional prediction modes to the table based on the position of each block relative to the current block.

[0185] When deriving a directional prediction mode for the current block based on a method of accumulating information (e.g., slope) acquired for each block relative to the surrounding restored template, the directionality (e.g., slope) that can be accumulated may vary depending on the location of each block. In this case, the definition of the location of each block and the different directionality depending on the location can be determined using one of various examples.

[0186] When constructing a table by calculating the slope for each block of the surrounding restored templates, a template region is defined based on the position of the surrounding restored templates of the current block. Thereafter, for each region, the directionality (slope) that can be included in the table or the directional intra prediction mode corresponding to the directionality can be determined according to an agreement between the video encoding device and the video decoding device. The method of constructing the table may vary. For example, the video decoding device may construct a separate table for each template region, perform normalization for each template region, then combine the normalized tables and derive the intra prediction mode using the combined table. Alternatively, the video decoding device may construct a single table for the entire template region and derive the intra prediction mode using the constructed table.

[0187] For example, when calculating the slope of a template, if the slope in the x direction is 0 or below a certain threshold, a vertical directional prediction mode may be added to the table. Similarly, when calculating the slope of a template, if the slope in the y direction is 0 or below a certain threshold, a horizontal directional prediction mode may be added to the table. At this time, the threshold may be a fixed value or an intermediate value of the brightness values ​​of the surrounding restored templates.

[0188] In the present disclosure, an image decoding device can derive an intra prediction mode of a current block by using 1) a reference block of an inter prediction signal of a current block, and a template of the reference block of the inter prediction signal, 2) a reference block in a reference picture of the current block, and a template of the reference block in the reference picture, 3) a reference block of an inter prediction signal of the current block, a template of the reference block of the inter prediction signal, and a template of the current block, 4) a reference block in a reference picture of the current block, a template of the reference block in the reference picture, and a template of the current block. At this time, the reference block in the reference picture can be obtained by using a merge index.

[0189] As an example, as shown in FIG. 17, a video decoding device can derive an intra prediction mode of a current block using a reference block in a reference picture. At this time, a reference MV (Motion Vector) indicating a reference block may be a final motion vector and / or an initial motion vector of an inter prediction signal of the current block. The reference MV may be determined by signaling / parsing an index after a motion vector candidate list is constructed from a surrounding restored area of ​​the current block. For example, a final motion vector may be generated by applying template-based motion compensation as described above to the initial motion vector.

[0190] As an example, the intra prediction mode of the current block can be derived using the reference block of the current block and the template of the reference block.

[0191] An image decoding device generates a prediction signal based on each intra prediction mode candidate of the current block using reference lines within the template of the reference block. The image decoding device can calculate a cost between the reference block and each prediction signal, and determine the reference line and intra prediction mode corresponding to the minimum cost as the reference line and intra prediction mode of the current block. The intra prediction mode candidates of the current block can be configured, as in the TIMD method described above, based on information about the surrounding restored regions of the current block, the size of the current block, the aspect ratio of the current block, etc. The information about the reference line can be identical to the information about the reference line available for the current block. For example, the nearest line can be used as the reference line.

[0192] When generating a prediction signal based on each intra prediction mode candidate of the current block using a reference line in the template of the reference block, prediction can be performed on a part of the reference block (an area excluding the lower right (Wb) × (Ha) in the reference block of FIG. 18) rather than the entire reference block, as shown in FIG. 18. The image decoding device can calculate a cost between a reference block of the same area and each prediction signal, and determine the reference line and intra prediction mode corresponding to the minimum cost as the intra prediction mode of the current block. Some areas can be defined as shown in FIG. 18. The values ​​of a and b can be integers less than or equal to W and H, respectively, and greater than 1, can be values ​​predefined according to an agreement between the image encoding device and the image decoding device, and can be values ​​determined according to the aspect ratio of the current block.

[0193] The video decoding device can use a reference block and a template of the reference block to calculate the directionality of a p×q block for integers p and q that are 2 or greater, and can construct a mapping table or a directionality table based on the calculated directionality, as in the aforementioned DIMD method. The video decoding device can derive the intra directional prediction mode with the largest accumulation in the mapping table as the intra prediction mode of the current block. Alternatively, the video decoding device can search for the directionality with the largest accumulation in the directionality table, and then derive the intra directional prediction mode corresponding to (mapped to) the searched directionality as the intra prediction mode of the current block. At this time, when calculating a region-by-region table for a table constructed with a template of a reference block, a region-by-region table can be constructed by setting a region according to the same division for the reference block as well.

[0194] When calculating the directionality of a p×q block for integers p and q greater than or equal to 2 using a reference block and a template of the reference block, the image decoding device may calculate the directionality not in all areas within the reference block, but in some areas, as shown in FIG. 18.

[0195] As another example, the intra prediction mode of the current block can be derived using the template of the reference block, the reference block, and the template of the current block.

[0196] The video decoding device can configure n intra prediction mode candidate groups (n is an integer greater than or equal to 1) based on the reference lines and intra prediction mode candidates of the template of the current block, as in the TIMD method described above. Here, each intra prediction mode candidate group includes a reference line candidate and an intra prediction mode candidate. The video decoding device generates a prediction signal based on the intra prediction mode of each candidate group from the template of the reference block. The video decoding device can calculate a cost between the reference block and each prediction signal, and determine the intra prediction mode of the current block based on the candidate group corresponding to the minimum cost. That is, the video decoding device can determine the reference line candidate and the intra prediction mode candidate of the candidate group corresponding to the minimum cost as the reference line and intra prediction mode of the current block. The video decoding device can generate a prediction signal for a partial region of the reference block, as shown in FIG. 18, without predicting the entire reference block, and calculate the cost for each candidate.

[0197] The video decoding device constructs a table based on the directionality generated according to the aforementioned DIMD method for the template of the current block, and also constructs a table for the reference block and the template of the reference block. The video decoding device can combine the two tables into a single mapping table or directional table to sum the same directional values. When combining, the weights for the accumulated values ​​of the directionality or prediction modes in the table constructed from the template of the current block may be different from the weights for the accumulated values ​​of the same directionality or prediction modes in the table constructed from the template of the reference block and the template of the reference block. The video decoding device can derive the intra directional prediction mode with the largest accumulation in the newly constructed mapping table as the intra prediction mode of the current block. Alternatively, the video decoding device can search for the directionality with the largest accumulation in the newly constructed directional table, and then determine the intra directional prediction mode corresponding to (mapped to) the searched directionality as the intra prediction mode of the current block.

[0198] Below, the process of determining the weights used to generate the final prediction signals is described.

[0199] The video decoding device can generate inter prediction signals and intra prediction signals, respectively, and then weight the two signals to generate final prediction signals.

[0200] For example, the weights of the inter-prediction signal and the intra-prediction signal may be determined based on the directionality of the prediction mode if the intra-prediction signal is generated according to a directional prediction mode. If the intra-prediction signal is not generated according to a directional prediction mode, the weights may be determined based on the prediction technique and prediction mode of the restored area surrounding the current block.

[0201] As an example, based on the intra directional prediction mode at a 135° angle (mode 34 in FIG. 3a), the directional modes can be divided into a first region and a second region. Here, the first region includes directional prediction modes greater than or equal to mode 34, and the second region includes directional prediction modes less than or equal to mode 34. If the intra prediction mode of the current block is included in the first region, the image decoding device may horizontally divide the current block into n pieces (where n is an integer greater than or equal to 1), as shown in FIG. 19a, and determine weights of the inter prediction block and the intra prediction block for each region. If the intra prediction mode of the current block is included in the second region, the image decoding device may vertically divide the current block into n pieces (where n is an integer greater than or equal to 1), as shown in FIG. 19b, and determine weights of the inter prediction block and the intra prediction block for each region. At this time, the first and second regions can be divided while including the wide angle (67~80, -14~-1 mode in Fig. 3b). n is 2 for an integer k greater than or equal to 1. k It can be determined differently depending on the size and aspect ratio of the current block. In Figures 19a and 19b, the numbers within the subblocks indicate the subblock index.

[0202] For example, if n is 4, the current block can be divided horizontally or vertically into 4 subblocks, and a weight can be determined for each subblock as shown in Table 2.

[0203]

[0204] Additionally, when n is 8, the current block is horizontally or vertically divided into 8 subblocks, and a weight can be determined for each subblock as shown in Table 3.

[0205]

[0206] According to Table 2 or Table 3, the closer a subblock is to the reference line, the greater the weight is assigned to the intra prediction signal.

[0207] As another example, as shown in FIG. 20, directional modes can be divided into regions 0 to 4. In FIG. 20, region 0 includes modes 56 and above, region 1 includes modes 45 to 55, region 2 includes modes 24 to 44, region 3 includes modes 13 to 23, and region 4 includes modes 12 and below. According to each region, the current block can be divided into sub-blocks, and weights of inter-prediction signals and intra-prediction signals of the current block can be determined. At this time, the current block can be divided into n (where n is an integer greater than or equal to 1). n is 2 for an integer greater than or equal to 1, k It can be different depending on the size and aspect ratio of the current block.

[0208] For example, if the intra prediction mode of the current block is included in the first region, the image decoding device may horizontally divide the current block as shown in Fig. 19a and determine weights for each sub-block as shown in Table 2 or Table 3. If the intra prediction mode of the current block is included in the third region, the image decoding device may vertically divide the current block as shown in Fig. 19b and determine weights for each sub-block as shown in Table 2 or Table 3.

[0209] If the intra prediction mode of the current block is included in the 0th region, the image decoding device may divide the current block as shown in FIG. 21a and determine a weight for each subblock as shown in Table 2 or Table 3. If the intra prediction mode of the current block is included in the 2nd region, the image decoding device may divide the current block as shown in FIG. 21b and determine a weight for each subblock as shown in Table 2 or Table 3. If the intra prediction mode of the current block is included in the 4th region, the image decoding device may divide the current block as shown in FIG. 21c and determine a weight for each subblock as shown in Table 2 or Table 3. In FIGS. 21a to 21c, numbers in the subblocks represent subblock indexes. According to Table 2 or Table 3, as described above, the closer a subblock is to the reference line, the greater the weight is assigned to the intra prediction signal.

[0210] The video decoding device can generate final prediction signals of the current block by weighting the inter prediction signal and the intra prediction signal using the weights determined as described above.

[0211] Hereinafter, a method for deriving an intra prediction mode based on an inter prediction signal is described using the cities of FIGS. 22 and 23.

[0212] FIG. 22 is a flowchart illustrating a method for an image encoding device to encode a current block according to an embodiment of the present disclosure.

[0213] The video encoding device obtains information indicating the motion vector of the current block (S2200). For example, the video encoding device may obtain information indicating the motion vector from a higher level. Alternatively, the video encoding device may determine information indicating the motion vector from the perspective of rate distortion optimization. The information indicating the motion vector may be a flag or an index.

[0214] The video encoding device constructs a list of motion vector candidates at predefined locations around and not adjacent to the current block.

[0215] The video encoding device can obtain a flag indicating the use of template matching from a higher level.

[0216] If the flag indicating the use of template matching is true, the video encoding device generates a reference block candidate indicated by each motion vector candidate in the motion vector candidate list. The video encoding device can perform template-based motion compensation on each candidate motion vector based on the template of the current block and the template of the reference block candidate. The video encoding device can rearrange candidate motion vectors in the motion vector candidate list by applying template matching to the template of the prediction block according to the motion compensation and the template of the current block.

[0217] In the future, the video encoding device may encode a flag indicating the use of template matching.

[0218] The video encoding device determines the motion vector of the current block from the motion vector candidate list based on information indicating the motion vector (S2202).

[0219] The video encoding device generates an inter prediction signal of the current block based on a reference block that exists in a reference picture and is indicated by a motion vector of the current block (S2204).

[0220] The video encoding device derives an intra prediction mode of the current block based on a reference block and a template of the reference block (S2206).

[0221] As an example, the intra prediction mode of the current block can be derived using the reference block of the current block and the template of the reference block.

[0222] The video encoding device can configure intra prediction mode candidates of the current block based on information of the surrounding restored area of ​​the current block, the size of the current block, or the aspect ratio of the current block, as in the aforementioned TIMD method. The video encoding device can generate a prediction signal based on each intra prediction mode candidate of the current block using a reference line in the template of the reference block. The video encoding device can calculate a cost between the reference block and the generated prediction signal, and determine the intra prediction mode candidate corresponding to the minimum cost as the intra prediction mode of the current block.

[0223] An image encoding device can use a reference block and a template of the reference block to calculate the directionality of a block having a preset size, as in the aforementioned DIMD method, and can configure a mapping table or a directionality table based on the calculated directionality. The image encoding device can derive the intra directional prediction mode with the largest accumulation in the mapping table as the intra prediction mode of the current block. The image encoding device can search for the directionality with the largest accumulation in the directionality table, and then derive the intra directional prediction mode corresponding to (mapped to) the searched directionality as the intra prediction mode of the current block.

[0224] As another example, the intra prediction mode of the current block can be derived using the reference block of the current block, the template of the reference block, and the template of the current block.

[0225] A video encoding device can generate intra prediction mode candidates of a current block based on reference lines and intra prediction mode candidates of a template of a current block, as in the TIMD method described above. Here, each intra prediction mode candidate can include a reference line candidate and an intra prediction mode candidate. The video encoding device can generate a prediction signal based on each intra prediction mode candidate of the current block using the reference lines in the template of the reference block. The video encoding device can calculate a cost between the reference block and the generated prediction signal, and determine the intra prediction mode and reference lines of the current block based on the intra prediction mode candidate corresponding to the minimum cost.

[0226] An image encoding device may configure a first table based on the directionality within the template of a current block, and configure a second table based on the directionality within the template of a reference block and the reference block. The image encoding device may configure a mapping table or a directional table by combining the first table and the second table. The image encoding device may derive the intra directional prediction mode with the largest accumulation in the mapping table as the intra prediction mode of the current block. Alternatively, the image encoding device may search for the directionality with the largest accumulation in the directional table, and then derive the intra directional prediction mode corresponding to (mapped to) the searched directionality as the intra prediction mode of the current block.

[0227] The video encoding device generates an intra prediction signal of the current block based on the intra prediction mode (S2208).

[0228] The video encoding device generates weights of an inter prediction signal and an intra prediction signal based on the directionality of the intra prediction mode (S2210).

[0229] The video encoding device generates sub-blocks by dividing the current block horizontally or vertically based on the directionality of the intra prediction mode. As shown in Table 2 or Table 3, the video encoding device can assign a greater weight to the intra prediction signal as each sub-block approaches the reference line of the current block.

[0230] The video encoding device generates final prediction signals of the current block by weighting inter prediction signals and intra prediction signals based on weights (S2212).

[0231] The video encoding device encodes information indicating a motion vector (S2214).

[0232] The video encoding device generates a residual block by subtracting the final prediction signals from the current block. The video encoding device transforms / quantizes the residual block to generate quantized transform coefficients and encodes the quantized transform coefficients.

[0233] FIG. 23 is a flowchart illustrating a method for an image decoding device to restore a current block according to one embodiment of the present disclosure.

[0234] The video decoding device decodes information indicating the motion vector of the current block (S2300). The information indicating the motion vector may be a flag or an index.

[0235] The video decoding device constructs a list of motion vector candidates at predefined locations around and not adjacent to the current block.

[0236] A video decoding device can decode a flag indicating the use of template matching from a bitstream.

[0237] If the flag indicating the use of template matching is true, the video decoding device generates a reference block candidate indicated by each motion vector candidate in the motion vector candidate list. The video decoding device can perform template-based motion compensation on each candidate motion vector based on the template of the current block and the template of the reference block candidate. The video decoding device can rearrange candidate motion vectors in the motion vector candidate list by applying template matching to the template of the prediction block according to the motion compensation and the template of the current block.

[0238] The video decoding device determines the motion vector of the current block from a motion vector candidate list based on information indicating the motion vector of the current block (S2302).

[0239] The video decoding device generates an inter prediction signal of the current block based on a reference block that exists in a reference picture and is indicated by a motion vector of the current block (S2304).

[0240] The video decoding device derives an intra prediction mode of the current block based on a reference block and a template of the reference block (S2306).

[0241] As an example, the intra prediction mode of the current block can be derived using the reference block of the current block and the template of the reference block.

[0242] The video decoding device can configure intra prediction mode candidates of the current block based on information about the surrounding restored area of ​​the current block, the size of the current block, or the aspect ratio of the current block, as in the aforementioned TIMD method. The video decoding device can generate a prediction signal based on each intra prediction mode candidate of the current block using a reference line in the template of the reference block. The video decoding device can calculate a cost between the reference block and the generated prediction signal, and determine the intra prediction mode candidate corresponding to the minimum cost as the intra prediction mode of the current block.

[0243] A video decoding device can use a reference block and a template of the reference block to calculate the directionality of a block having a preset size, as in the aforementioned DIMD method, and construct a mapping table or a directionality table based on the calculated directionality. The video decoding device can derive the intra directional prediction mode with the largest accumulation in the mapping table as the intra prediction mode of the current block. The video decoding device can search for the directionality with the largest accumulation in the directionality table, and then derive the intra directional prediction mode corresponding to (mapped to) the searched directionality as the intra prediction mode of the current block.

[0244] As another example, the intra prediction mode of the current block can be derived using the reference block of the current block, the template of the reference block, and the template of the current block.

[0245] A video decoding device can generate intra prediction mode candidates of a current block based on reference lines and intra prediction mode candidates of a template of a current block, as in the TIMD method described above. Here, each intra prediction mode candidate can include a reference line candidate and an intra prediction mode candidate. The video decoding device can generate a prediction signal based on each intra prediction mode candidate of the current block using the reference lines in the template of the reference block. The video decoding device can calculate a cost between the reference block and the generated prediction signal, and determine the intra prediction mode and reference lines of the current block based on the intra prediction mode candidate corresponding to the minimum cost.

[0246] An image decoding device may configure a first table based on the directionality within the template of a current block, and configure a second table based on the directionality within the template of a reference block and the reference block. The image decoding device may configure a mapping table or a directional table by combining the first table and the second table. The image decoding device may derive the intra directional prediction mode with the largest accumulation in the mapping table as the intra prediction mode of the current block. Alternatively, the image decoding device may search for the directionality with the largest accumulation in the directional table, and then derive the intra directional prediction mode corresponding to (mapped to) the searched directionality as the intra prediction mode of the current block.

[0247] The video decoding device generates an intra prediction signal of the current block based on the intra prediction mode (S2308).

[0248] The video decoding device generates weights of the inter prediction signal and the intra prediction signal based on the directionality of the intra prediction mode (S2310).

[0249] The video decoding device generates sub-blocks by dividing the current block horizontally or vertically based on the directionality of the intra prediction mode. As shown in Table 2 or Table 3, the video decoding device can assign a greater weight to the intra prediction signal as each sub-block approaches the reference line of the current block.

[0250] The image decoding device generates final prediction signals of the current block by weighting inter prediction signals and intra prediction signals based on weights (S2312).

[0251] The video decoding device decodes quantized transform coefficients from the bitstream and dequantizes / inversely transforms the quantized transform coefficients to restore the residual block. The video decoding device adds the final prediction signals of the current block and the residual block to generate a restored block of the current block.

[0252] Although the flowchart / timing diagram of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of ​​one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart / timing diagram without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart / timing diagram is not limited to a chronological order.

[0253] It should be understood that the exemplary embodiments described above can be implemented in many different ways. The functions or methods described in one or more examples can be implemented in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as "units" to further emphasize their implementation independence.

[0254] Meanwhile, the various functions or methods described in this embodiment may also be implemented as instructions stored on a non-transitory storage medium that can be read and executed by one or more processors. Non-transitory storage media include, for example, all types of storage devices that store data in a form readable by a computer system. For example, non-transitory storage media include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical drives, magnetic hard drives, and solid-state drives (SSDs).

[0255] The above description is merely an example of the technical idea of ​​the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of ​​the present embodiment, but rather to explain it, and the scope of the technical idea of ​​the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.

[0256]

[0257]

[0258] CROSS-REFERENCE TO RELATED APPLICATION

[0259] This patent application claims priority to Korean patent application No. 10-2023-0170532, filed on November 30, 2023, and Korean patent application No. 10-2024-0146612, filed on October 24, 2024, the entire contents of which are incorporated herein by reference.

Claims

1. A method for restoring a current block performed by a video decoding device, A step of determining a motion vector of the current block from a motion vector candidate list based on information indicating the motion vector of the current block; A step of generating an inter prediction signal of the current block based on a reference block existing in a reference picture and indicated by a motion vector of the current block; A step of deriving an intra prediction mode of the current block based on the reference block and the template of the reference block; A step of generating an intra prediction signal of the current block based on the intra prediction mode; and A step of generating final prediction signals of the current block based on the inter prediction signal and the intra prediction signal. A method comprising:

2. In paragraph 1, A step of decoding information indicating the motion vector from a bitstream; and A step of constructing a list of motion vector candidates at predetermined locations around and not adjacent to the current block. A method further comprising:

3. In paragraph 1, A step of generating a reference block candidate indicated by each motion vector candidate in the above motion vector candidate list; and A step of performing template-based motion compensation on each candidate motion vector based on the template of the current block and the template of the reference block candidate. A method further comprising:

4. In paragraph 3, A step of rearranging candidate motion vectors in the motion vector candidate list by applying template matching to the template of the prediction block according to the above motion compensation and the template of the current block. A method further comprising:

5. In paragraph 1, The step of deriving the above intra prediction mode is: A step of generating a prediction signal based on each intra prediction mode candidate of the current block using a reference line within the template of the reference block; A step of calculating a cost between the above reference block and the generated prediction signal; and A step of determining an intra prediction mode candidate corresponding to the minimum cost as the intra prediction mode of the current block. A method comprising:

6. In paragraph 5, A method further comprising the step of configuring intra prediction mode candidates of the current block based on information of a restored area surrounding the current block, the size of the current block, or the aspect ratio of the current block.

7. In paragraph 1, The step of deriving the above intra prediction mode is: A step of calculating the directionality of a block having a preset size using the above reference block and the template of the reference block; A step of accumulating the calculated directionality to form a table; and A step of deriving the most accumulated directional prediction mode in the above table as the intra prediction mode of the current block. A method comprising:

8. In paragraph 1, The step of deriving the above intra prediction mode is: A step of generating a prediction signal based on each intra prediction mode candidate of the current block by using a reference line within the template of the reference block; A step of calculating a cost between the above reference block and the generated prediction signal; and A step of determining the intra prediction mode and reference line of the current block based on a group of intra prediction mode candidates corresponding to the minimum cost. A method comprising:

9. In paragraph 8, Further comprising a step of generating intra prediction mode candidates of the current block based on reference lines of the template of the current block and intra prediction mode candidates, A method, wherein each intra prediction mode candidate group includes a reference line candidate and an intra prediction mode candidate.

10. In paragraph 1, The step of deriving the above intra prediction mode is: A step of constructing a first table based on the directionality within the template of the current block; A step of configuring a second table based on the directionality within the above reference block and the template of the above reference block; A step of combining the first table and the second table to create a combined table; and A step of deriving the most accumulated directional prediction mode from the above combination table as the intra prediction mode of the current block. A method comprising:

11. In paragraph 1, The step of generating the above final prediction signals is: A step of generating weights of the inter prediction signal and the intra prediction signal based on the directionality of the intra prediction mode; and A step of generating final prediction signals of the current block by weighting the inter prediction signal and the intra prediction signal based on the above weights. A method comprising:

12. In paragraph 11, The step of generating the above weights is: A step of generating sub-blocks by dividing the current block horizontally or vertically based on the directionality of the intra prediction mode; and A step of assigning a larger weight to the intra prediction signal as each sub-block approaches the reference line of the current block. A method comprising:

13. A method for encoding a current block performed by a video encoding device, A step of obtaining information indicating a motion vector of the current block; A step of determining a motion vector of the current block from a motion vector candidate list based on information indicating the motion vector; A step of generating an inter prediction signal of the current block based on a reference block existing in a reference picture and indicated by a motion vector of the current block; A step of deriving an intra prediction mode of the current block based on the reference block and the template of the reference block; A step of generating an intra prediction signal of the current block based on the intra prediction mode; and A step of generating final prediction signals of the current block based on the inter prediction signal and the intra prediction signal. A method comprising:

14. In paragraph 13, A step of constructing a list of motion vector candidates at predetermined locations around and not adjacent to the current block; and A step of encoding information indicating the above motion vector; A method further comprising:

15. In paragraph 13, A step of generating a reference block candidate indicated by each motion vector candidate in the above motion vector candidate list; and A step of performing template-based motion compensation on each candidate motion vector in the motion vector candidate list based on the template of the current block and the template of the reference block. A method comprising:

16. In paragraph 13, The step of generating the above final prediction signals is: A step of generating weights of the inter prediction signal and the intra prediction signal based on the directionality of the intra prediction mode; and A step of generating final prediction signals of the current block by weighting the inter prediction signal and the intra prediction signal based on the above weights. A method further comprising:

17. A method for providing video data to a video decoding device, A step of encoding the above video data into a bitstream; and A step of transmitting the above bitstream to the image decoding device Including, The step of encoding the above video data is: A step of obtaining information indicating the motion vector of the current block; A step of determining a motion vector of the current block from a motion vector candidate list based on information indicating the motion vector; A step of generating an inter prediction signal of the current block based on a reference block existing in a reference picture and indicated by a motion vector of the current block; A step of deriving an intra prediction mode of the current block based on the reference block and the template of the reference block; A step of generating an intra prediction signal of the current block based on the intra prediction mode; and A step of generating final prediction signals of the current block based on the inter prediction signal and the intra prediction signal. A method comprising:

Citation Information

Patent Citations

  • Device And Method For Setting Mosquito Net Or Wind Proof Sheet Using Velcro

    KR1020240175491A

  • Measuring apparatus for roller of firing furnace

    KR1020250043878A

  • Method for image processing and apparatus for implementing the same

    US20230011999A1

  • Video signal processing method based on template matching, and device therefor

    WO2023182781A1