Video coding method and apparatus, and recording medium having bitstream stored therein

The DIMD-based method addresses the challenge of efficiently compressing high-resolution videos by inducing accurate intra prediction modes through filter applications and HOG table generation, resulting in improved compression efficiency and predictive block accuracy.

WO2025095612A1PCT designated stage expired Publication Date: 2025-05-08DONG A UNIV RES FOUND FOR IND ACAD COOP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016892
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-15
Filing Date
2024-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current video compression technologies face challenges in efficiently compressing high-resolution videos, particularly in deriving accurate intra prediction modes for improved compression efficiency.

Method used

The proposed method and device use a DIMD (Decoder-Intra Mode Derivation) approach to induce intra predictive modes for video blocks by applying filters to peripheral regions of the current block, generating HOG tables, and combining predictive blocks based on intra prediction modes and planar modes.

Benefits of technology

This approach enhances video compression efficiency by deriving more accurate intra prediction modes, reducing memory complexity, and improving the accuracy of predictive blocks, thereby leading to better compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016892_08052025_PF_FP_ABST
    Figure KR2024016892_08052025_PF_FP_ABST
Patent Text Reader

Abstract

An image encoding / decoding method and apparatus according to the present disclosure may: derive an intra prediction mode for the current block by applying a filter to a template of the current block; derive a weight for the intra prediction mode; and generate a final prediction block on the basis of a prediction block generated on the basis of the intra prediction mode, and the weight.
Need to check novelty before this filing date? Find Prior Art

Description

Video coding method and device, and recording medium storing bitstream

[0001] The present invention relates to a video signal processing method and device.

[0002] The market demand for high-resolution video is growing, necessitating technologies capable of efficiently compressing high-resolution images. To address this market need, the ISO / IEC's Moving Picture Expert Group (MPEG) and the ITU-T's Video Coding Expert Group (VCEG) jointly formed the Joint Collaborative Team on Video Coding (JCT-VC). They completed development of the HEVC (High Efficiency Video Coding) video compression standard in January 2013 and have been actively conducting research and development on next-generation compression standards.

[0003] Video compression largely consists of intraprediction, interprediction, transform, quantization, entropy coding, and in-loop filtering. Among these, intraprediction refers to a technique that generates a prediction block for the current block using reconstructed pixels surrounding the current block. The encoder encodes the intraprediction mode used for intraprediction, and the decoder performs intraprediction by reconstructing the encoded intraprediction mode.

[0004] The present disclosure provides a method and device for deriving an intra prediction mode based on DIMD.

[0005] The present disclosure provides a method and device for generating an intra prediction block based on DIMD.

[0006] The video decoding method and device according to the present disclosure can induce an intra prediction mode for the current block by applying a filter to a template of the current block, induce a weight for the intra prediction mode, and generate a prediction block generated based on the intra prediction mode and a final prediction block of the current block based on the weight.

[0007] In the image decoding method and device according to the present disclosure, the template is a peripheral area adjacent to the current block, and the peripheral area may include at least one of a left area, an upper area, or an upper left area.

[0008] In the image decoding method and device according to the present disclosure, the range of the template to which the filter is applied can be variably determined based on the height and width of the current block.

[0009] In the image decoding method and device according to the present disclosure, the center sample among the reference samples input to the filter may belong to at least one of the first reference sample line that is 1 sample away from the boundary of the current block or the second reference sample line that is 2 samples away from the boundary of the current block.

[0010] In the video decoding method and device according to the present disclosure, if there is an unavailable sample among the reference samples input to the filter, the unavailable sample is replaced with an available sample, and the available sample can be generated based on a predetermined interpolation method or a predetermined intra prediction mode.

[0011] In the video decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include the step of generating a HoG table for a template of the current block, and the HoG table may further include a frequency for each intra prediction mode.

[0012] In the image decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include a step of generating a HoG of the current block based on a DIMD HoG of a surrounding block.

[0013] In the video decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include a step of generating an HoG for a reference block specified by a predetermined block vector.

[0014] In the image decoding method and device according to the present disclosure, the final prediction block can be generated based on a weighted sum of a prediction block generated based on the intra prediction mode and a prediction block generated based on the planar mode.

[0015] In the image decoding method and device according to the present disclosure, the Planar mode can be adaptively induced into any one of a general Planar mode, a vertical Planar mode, or a horizontal Planar mode.

[0016] The video encoding method and device according to the present disclosure can derive an intra prediction mode for the current block by applying a filter to a template of the current block, derive a weight for the intra prediction mode, and generate a prediction block generated based on the intra prediction mode and a final prediction block of the current block based on the weight.

[0017] A computer-readable recording medium according to the present invention can store a bitstream encoded by the image encoding method.

[0018] According to the present disclosure, the compression efficiency of a video encoder / decoder can be increased by giving directionality to the planar mode in a weighted sum with the planar mode.

[0019] According to the present disclosure, by expanding or changing the application range of the filter in DIMD, an intra prediction mode closer to the original can be derived and the compression efficiency of a video encoder / decoder can be increased.

[0020] According to the present disclosure, by proposing DIMD that takes into account the frequency of intra prediction modes, a more accurate intra prediction mode can be derived.

[0021] According to the present disclosure, the performance of a video encoder / decoder can be improved and memory complexity can be reduced by using a modified HoG table generation method.

[0022] According to the present disclosure, the performance of a video encoder / decoder can be improved by increasing the accuracy of a prediction block through DIMD merge mode or DIMD fusion mode.

[0023] According to the present disclosure, the accuracy of the final prediction block can be improved by considering not only the surrounding area of ​​the current block but also the variation of the reference block.

[0024] FIG. 1 is a block diagram showing an image encoding device according to the present disclosure.

[0025] FIG. 2 is a block diagram showing an image decoding device according to the present disclosure.

[0026] FIG. 3 illustrates a DIMD-based prediction block generation method performed in an image encoding / decoding device according to the present disclosure.

[0027] The video decoding method and device according to the present disclosure can induce an intra prediction mode for the current block by applying a filter to a template of the current block, induce a weight for the intra prediction mode, and generate a prediction block generated based on the intra prediction mode and a final prediction block of the current block based on the weight.

[0028] In the image decoding method and device according to the present disclosure, the template is a peripheral area adjacent to the current block, and the peripheral area may include at least one of a left area, an upper area, or an upper left area.

[0029] In the image decoding method and device according to the present disclosure, the range of the template to which the filter is applied can be variably determined based on the height and width of the current block.

[0030] In the image decoding method and device according to the present disclosure, the center sample among the reference samples input to the filter may belong to at least one of the first reference sample line that is 1 sample away from the boundary of the current block or the second reference sample line that is 2 samples away from the boundary of the current block.

[0031] In the video decoding method and device according to the present disclosure, if there is an unavailable sample among the reference samples input to the filter, the unavailable sample is replaced with an available sample, and the available sample can be generated based on a predetermined interpolation method or a predetermined intra prediction mode.

[0032] In the video decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include the step of generating a HoG table for a template of the current block, and the HoG table may further include a frequency for each intra prediction mode.

[0033] In the image decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include a step of generating a HoG of the current block based on a DIMD HoG of a surrounding block.

[0034] In the video decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include a step of generating an HoG for a reference block specified by a predetermined block vector.

[0035] In the image decoding method and device according to the present disclosure, the final prediction block can be generated based on a weighted sum of a prediction block generated based on the intra prediction mode and a prediction block generated based on the planar mode.

[0036] In the image decoding method and device according to the present disclosure, the Planar mode can be adaptively induced into any one of a general Planar mode, a vertical Planar mode, or a horizontal Planar mode.

[0037] The video encoding method and device according to the present disclosure can derive an intra prediction mode for the current block by applying a filter to a template of the current block, derive a weight for the intra prediction mode, and generate a prediction block generated based on the intra prediction mode and a final prediction block of the current block based on the weight.

[0038] A computer-readable recording medium according to the present invention can store a bitstream encoded by the image encoding method.

[0039] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings attached to this specification so that those skilled in the art can easily implement the present invention. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In addition, in the drawings, parts irrelevant to the description have been omitted to clearly explain the present invention, and similar parts have been designated with similar reference numerals throughout the specification.

[0040] Throughout this specification, when a part is said to be 'connected' to another part, this includes not only cases where they are directly connected, but also cases where they are electrically connected with another element in between.

[0041] Additionally, whenever a part throughout this specification is said to "include" a component, this does not mean that other components are excluded, but rather that other components may be included, unless otherwise specifically stated.

[0042] Additionally, while terms such as first, second, etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another.

[0043] Additionally, in the embodiments of the devices and methods described herein, some components of the devices or some steps of the methods may be omitted. Furthermore, the order of some components of the devices or some steps of the methods may be changed. Furthermore, other components or other steps may be inserted into some components of the devices or some steps of the methods.

[0044] Additionally, some components or some steps of the first embodiment of the present invention may be added to the second embodiment of the present invention, or some components or some steps of the second embodiment may be replaced.

[0045] In addition, the components shown in the embodiments of the present invention are independently depicted to represent different characteristic functions, and this does not mean that each component is composed of separate hardware or a single software component. That is, each component is described by listing each component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function. Such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0046] In this specification, a block can be variously expressed as a unit, an area, a unit, a partition, etc., and a sample can be variously expressed as a pixel, a pel, a pixel, etc.

[0047] Hereinafter, embodiments of the present invention will be described in more detail with reference to the attached drawings. In describing the present invention, duplicate descriptions of identical components will be omitted.

[0048] FIG. 1 is a block diagram showing an image encoding device according to the present disclosure.

[0049] Referring to FIG. 1, a video encoding device (100) may include a picture segmentation unit (110), a prediction unit (120, 125), a transformation unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse transformation unit (145), a filter unit (150), and a memory (155).

[0050] The picture segmentation unit (110) can segment the input picture into at least one processing unit. At this time, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). Hereinafter, in the embodiments of the present disclosure, the coding unit may be used to mean a unit that performs encoding or a unit that performs decoding.

[0051] A prediction unit may be divided into at least one square or rectangular shape of the same size within a single coding unit, or may be divided such that one prediction unit among the divided prediction units within a single coding unit has a different shape and / or size from another prediction unit. When a prediction unit that performs intra prediction based on a coding unit is generated and is not the minimum coding unit, intra prediction can be performed without being divided into a plurality of NxN prediction units.

[0052] The prediction unit (120, 125) may include an inter prediction unit (120) that performs inter prediction or inter prediction, and an intra prediction unit (125) that performs intra prediction or intra prediction. It may determine whether to use inter prediction or intra prediction for a prediction unit, and determine specific information (e.g., intra prediction mode, motion vector, reference picture, etc.) according to each prediction method. A residual value (residual block) between the generated prediction block and the original block may be input to the transformation unit (130). In addition, prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value by the entropy encoding unit (165) and transmitted to the decoder.

[0053] The inter prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous or subsequent pictures of the current picture, and in some cases, may predict a prediction unit based on information of a portion of an encoded region within the current picture. The inter prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.

[0054] The reference picture interpolation unit can receive reference picture information from the memory (155) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.

[0055] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm) can be used to derive a motion vector. The motion vector can have a motion vector value in units of 1 / 2 or 1 / 4 pixels based on the interpolated pixel. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various methods such as Skip Mode, Merge Mode, AMVP Mode, Intra Block Copy Mode, and Affine Mode can be used as motion prediction methods.

[0056] The intra prediction unit (125) can generate a prediction unit based on reference pixel information surrounding the current block, which is pixel information within the current picture. If the surrounding block of the current prediction unit is a block on which inter prediction has been performed and the reference pixel is a pixel on which inter prediction has been performed, the reference pixel included in the block on which inter prediction has been performed can be replaced and used with reference pixel information of the surrounding block on which intra prediction has been performed. That is, if the reference pixel is not available, the unavailable reference pixel information can be replaced and used with at least one reference pixel among the available reference pixels.

[0057] Additionally, a residual block containing residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130).

[0058] In the transformation unit (130), the residual block including the residual value information of the prediction unit generated through the original block and the prediction unit (120, 125) can be transformed using a transformation method such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on the intra prediction mode information of the prediction unit used to generate the residual block.

[0059] The quantization unit (135) can quantize values ​​converted to the frequency domain by the transformation unit (130). The quantization coefficients can vary depending on the block or the importance of the image. The values ​​produced by the quantization unit (135) can be provided to the dequantization unit (140) and the reordering unit (160).

[0060] The rearrangement unit (160) can perform rearrangement of coefficient values ​​for quantized residual values.

[0061] The rearrangement unit (160) can change a two-dimensional block-shaped coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the rearrangement unit (160) can change the two-dimensional block-shaped coefficient into a one-dimensional vector form by scanning from the DC coefficient to the coefficient of the high-frequency region using a zig-zag scan method. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional block-shaped coefficient in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficient in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, it is possible to determine which scan method among the zig-zag scan, the vertical scan, and the horizontal scan is to be used.

[0062] The entropy encoding unit (165) can perform entropy encoding based on the values ​​produced by the rearrangement unit (160). Entropy encoding can use various encoding methods such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC). In this regard, the entropy encoding unit (165) can encode residual value coefficient information of the encoding unit from the rearrangement unit (160) and the prediction units (120, 125). In addition, according to the present disclosure, it is possible to signal and transmit information indicating that motion information is derived and used on the decoder side and information on a technique used to derive motion information.

[0063] The inverse quantization unit (140) and the inverse transformation unit (145) inversely quantize the values ​​quantized in the quantization unit (135) and inversely transform the values ​​transformed in the transformation unit (130). The residual values ​​generated in the inverse quantization unit (140) and the inverse transformation unit (145) can be combined with the predicted prediction units predicted through the motion estimation unit, motion compensation unit, and intra prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.

[0064] The filter unit (150) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter). The deblocking filter may remove block distortion caused by boundaries between blocks in a restored picture. The offset correction unit may correct the offset from the original image on a pixel-by-pixel basis for the image on which deblocking has been performed. In order to perform offset correction for a specific picture, a method may be used in which the pixels included in the image are divided into a certain number of regions, the regions to be offset are determined, and the offset is applied to the regions, or the offset is applied by considering edge information of each pixel. The ALF (Adaptive Loop Filtering) may be performed based on a value obtained by comparing the filtered restored image with the original image. After dividing the pixels included in the image into a predetermined group, one filter to be applied to the group is determined, and filtering may be performed differentially for each group.

[0065] The memory (155) can store a restored block or picture produced through the filter unit (150), and the stored restored block or picture can be provided to the prediction unit (120, 125) when performing inter prediction.

[0066] FIG. 2 is a block diagram showing an image decoding device according to the present disclosure.

[0067] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (210), a rearrangement unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).

[0068] When a video bitstream is input to a video encoding device, the input bitstream can be decoded in the opposite procedure to that of the video encoding device.

[0069] The entropy decoding unit (210) can perform entropy decoding in a procedure opposite to that of the entropy encoding unit of the video encoder. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied in response to the method performed in the video encoder.

[0070] The entropy decoding unit (210) can decode information related to intra prediction and inter prediction performed in the encoder.

[0071] The reordering unit (215) can perform reordering based on the method by which the bitstream entropy-decoded by the entropy decoding unit (210) is reordered by the encoding unit. The coefficients expressed in the form of a one-dimensional vector can be reordered by restoring them back to coefficients in the form of a two-dimensional block.

[0072] The inverse quantization unit (220) can perform inverse quantization based on the quantization parameters provided by the encoder and the coefficient values ​​of the rearranged block.

[0073] The inverse transform unit (225) can perform inverse transform, i.e., inverse DCT, inverse DST, and inverse KLT, on the transforms performed by the transform unit, i.e., DCT, DST, and KLT, on the quantization result performed by the image encoder. The inverse transform can be performed based on the transmission unit determined by the image encoder. In the inverse transform unit (225) of the image decoder, a transform technique (e.g., DCT, DST, KLT) can be selectively performed according to a plurality of pieces of information, such as a prediction method, the size of the current block, and the prediction direction.

[0074] The prediction unit (230, 235) can generate a prediction block based on prediction block generation related information provided from the entropy decoding unit (210) and previously decoded block or picture information provided from the memory (245).

[0075] As described above, when performing intra prediction or intra prediction in the same manner as the operation in the image encoder, if the size of the prediction unit and the size of the transformation unit are the same, intra prediction for the prediction unit is performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side. However, when performing intra prediction, if the size of the prediction unit and the size of the transformation unit are different, intra prediction can be performed using reference pixels based on the transformation unit. In addition, intra prediction using NxN division only for the minimum coding unit can be used.

[0076] The prediction unit (230, 235) may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (210), prediction mode information of an intra prediction method, and motion prediction-related information of an inter prediction method, and may distinguish a prediction unit from a current encoding unit and determine whether the prediction unit performs inter prediction or intra prediction. On the other hand, if the encoder (100) does not transmit motion prediction-related information for the inter prediction, but instead transmits information indicating that motion information is to be derived and used on the decoder side and information on a technique used to derive motion information, the prediction unit determination unit determines whether the inter prediction unit (230) performs prediction based on the information transmitted from the encoder (100).

[0077] The inter prediction unit (230) can perform inter prediction on the current prediction unit based on information included in at least one picture among the previous picture or the subsequent picture of the current picture including the current prediction unit, using information required for inter prediction of the current prediction unit provided by the image encoder. In order to perform inter prediction, it can be determined based on the encoding unit whether the motion prediction method of the prediction unit included in the corresponding encoding unit is one of Skip Mode, Merge Mode, AMVP Mode, Intra Block Copy Mode, and Affine Mode.

[0078] The intra prediction unit (235) can generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit that has performed intra prediction, intra prediction can be performed based on intra prediction mode information of the prediction unit provided by the image encoder.

[0079] The intra prediction unit (235) may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a unit that performs filtering on the reference pixels of the current block and can determine whether to apply the filter based on the prediction mode of the current prediction unit and apply it. AIS filtering can be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the image encoder. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.

[0080] The reference pixel interpolation unit can interpolate the reference pixel to generate a reference pixel of a pixel unit less than an integer value when the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on the pixel value interpolated from the reference pixel. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.

[0081] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.

[0082] Information about whether a deblocking filter has been applied to a corresponding block or picture can be received from a video encoding device, and if a deblocking filter has been applied, information about whether a strong or weak filter has been applied. The deblocking filter of the video decoder can receive information related to the deblocking filter provided by the video encoder, and the video decoder can perform deblocking filtering on the corresponding block.

[0083] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and information on the offset value. ALF can be applied to the encoding unit based on information on whether ALF is applied and ALF coefficient information provided from the encoder. This ALF information can be provided by being included in a specific parameter set.

[0084] The memory (245) can store a restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to an output unit.

[0085] The decoder-side intra mode derivation (DIMD) according to the present disclosure may be a method for deriving an intra prediction mode based on the gradient of sample values ​​between reference samples belonging to a surrounding area of ​​a current block. A predetermined filter may be used to calculate the gradient. The filter may have a predetermined size of NxM, and N and M may be positive integers. N and M may have the same value or different values.

[0086] For example, DIMD can extract a gradient by applying a filter of a certain size (e.g., a Sobel filter of size 3x3) to the surrounding area of ​​the current block (hereinafter referred to as a template). The gradient can be extracted by applying the filter to reference samples within the template while shifting by 1 space from the lower left to the upper right of the template (stride=1).

[0087] A Histogram of Gradient (HoG) can be generated based on the extracted gradient. At this time, the y-axis of the HoG is defined as the sum of the absolute values ​​of the horizontal gradient (gradient x, Gx) and the vertical gradient (gradient y, Gy), and the x-axis of the HoG can be defined as an intra prediction mode that maps to the values ​​and signs of Gx and Gy based on a pre-defined table. The HoG can be a method of expressing the features of an image or video as a histogram of the amount of change (Gradient) according to the direction. The generated HoG can be rearranged in ascending order based on the y-axis.

[0088] Based on the HoG, one or more intra prediction modes can be derived for the current block. One or more prediction blocks can be generated for the current block based on the one or more intra prediction modes. If multiple intra prediction modes are derived for the current block, the final prediction block of the current block can be generated through a weighted sum of the multiple prediction blocks for the current block. Alternatively, the final prediction block of the current block can be generated through a weighted sum between one or more prediction blocks for the current block and a prediction block generated based on the Planar mode. This process can be performed identically in the decoder and the encoder. Hereinafter, a method for generating a final prediction block of the current block based on DIMD will be described in detail with reference to FIG. 3.

[0089] Referring to FIG. 3, an intra prediction mode for the current block can be derived by applying a filter to the template of the current block (S300).

[0090] A template according to the present disclosure may be a surrounding region adjacent to a current block. The surrounding region may include at least one of a left region, an upper region, or an upper left region. For example, if the left region is unavailable (or the current block borders the left edge of the current picture), a filter may be applied only to the upper region. Alternatively, if the upper region is unavailable (or the current block borders the upper edge of the current picture), a filter may be applied only to the left region.

[0091] For convenience of explanation, it is assumed that the template includes a left region, a top region, and a top-left region. A HoG buffer for the template can be created / initialized. The HoG buffer can be created / initialized for each region belonging to the template. For example, the HoG buffer can include a left HoG buffer for the left region, an upper HoG buffer for the top region, an upper-left HoG buffer for the top-left region, and a full HoG buffer for the entire template region. The left HoG buffer can store the HoG for the left region, the upper HoG buffer can store the HoG for the top region, the upper-left HoG buffer can store the HoG for the top-left region, and the full HoG buffer can store the HoG for the template.

[0092] A HoG can be created by applying filters to each area within a template. Filters can be applied in the following order: left area, top area, top left area. However, this is merely an example, and the order of filter application is not limited thereto. For example, filters can be applied in the following order: left area, top left area, top area. Alternatively, filters can be applied in the following order: top area, top left area, left area. Alternatively, filters can be applied in the following order: top area, left area, top left area.

[0093] By applying a filter to each region of the template, the values ​​of Gx and Gy can be derived. Based on the derived values ​​of Gx and Gy, the intra prediction mode and amplitude of the variation stored in the corresponding HoG buffer can be derived.

[0094] The range of the template to which the filter is applied can be determined based on the height and width of the current block. For example, the height of the left region may be the same as the height of the current block. The width of the top region may be the same as the width of the current block. However, this is not limited thereto. If there are available reference sample(s) in the lower left peripheral area and / or the upper right peripheral area of ​​the current block, the range of application of the filter can be expanded. At this time, the expanded range can have a size of N-samples. Here, N can be an integer of 1, 2, 3, 4, or more. For example, if there are available reference sample(s) in the lower left peripheral area of ​​the current block, the left region of the template can be expanded downward by a size of 4-samples. If there are available reference sample(s) in the upper right peripheral area of ​​the current block, the upper region of the template can be expanded rightward by a size of 4-samples.

[0095] Among the reference samples input to the above filter, the center sample may belong to the 1st reference sample line that is 1 sample away from the boundary of the current block. That is, if the reference sample line adjacent to the current block is called the 0th reference sample line, the 1st reference sample line may be adjacent to the 0th reference sample line. Hereinafter, the kth reference sample line may be defined as being adjacent to the (k-1)th reference sample line. The above filter may utilize the weights of existing filters such as a Sobel filter and a Gaussian filter, or may utilize arbitrary weights according to the user.

[0096] According to the present disclosure, an intra prediction mode can be derived that can generate a prediction block closer to the original image by extending or changing a template to which a filter is applied.

[0097] For example, the template may include at least one of a left region, a top region, or an upper left region. Here, the left region may be composed of three reference sample lines (i.e., a 0th reference sample line, a 1st reference sample line, and a 2nd reference sample line). The lengths of the three reference sample lines may be equal to the height of the current block. The top region may be composed of three reference sample lines (i.e., a 0th reference sample line, a 1st reference sample line, and a 2nd reference sample line). The lengths of the three reference sample lines may be equal to the width of the current block. The upper left region may be defined as a 3x3 region.

[0098] The above template can be extended by N samples. N can be an integer greater than or equal to 1. For example, if the template is extended by 1 sample, the left region of the extended template can include a third reference sample line in addition to the three reference sample lines. The upper region of the extended template can include a third reference sample line in addition to the three reference sample lines. The upper left region of the extended template can be defined as a 4x4 region.

[0099] Among the reference samples input to the filter, the center sample may include a reference sample belonging to a first reference sample line that is 1 sample away from the boundary of the current block and a reference sample belonging to a second reference sample line that is 2 samples away from the boundary of the current block. Alternatively, among the reference samples input to the filter, the center sample may be at least one reference sample belonging only to the second reference sample line that is 2 samples away from the boundary of the current block. Alternatively, among the reference samples input to the filter, the center sample may be at least one reference sample belonging to a 0th reference sample line adjacent to the boundary of the current block.

[0100] If there are unavailable samples among the reference samples input to the filter (for example, if an internal sample of the current block that has not been decoded / encoded must be input to the filter), the unavailable samples can be replaced with samples generated through a predetermined method.

[0101] For example, unavailable samples can be replaced with samples generated through a predetermined interpolation. Here, the interpolation can be performed based on a method such as nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation.

[0102] Alternatively, internal samples of the current block that are not decoded / encoded can be generated based on predetermined intra prediction mode(s), and the generated internal samples can be input to the filter. Here, the predetermined intra prediction mode(s) can include at least one of a planar mode, a DC mode, a vertical mode, and a horizontal mode. Any one of the aforementioned predetermined intra prediction mode(s) can be selectively used based on the size of the current block. The size of the current block can be defined by width, height, a ratio of width and height, a product of width and height, a maximum value / minimum value of width and height, etc.

[0103] Alternatively, the internal samples of the current block that have not been decoded / encoded may be generated based on a combination of at least two of the aforementioned intra prediction modes. For example, the internal samples may be generated as a weighted sum between samples generated based on a vertical mode and samples generated based on a horizontal mode.

[0104] Whether a template is expanded may be determined based on whether the current block is a square block. Alternatively, the scope / size of a template for DIMD may be determined differently depending on whether the current block is a square block. Alternatively, the scope / size of a template for DIMD may be determined differently depending on whether the width of the current block is greater than its height.

[0105] For example, if the current block is a square block, the extended template may not be applied to the current block. If the current block is a non-square block, the extended template described above may be applied to the current block.

[0106] Alternatively, if the current block is a square block, the extended template may not be applied to the current block. If the current block is a non-square block, the extended template may be applied to some areas. For example, if the width of the current block is greater than the height, the left area may be extended by N samples, but the top area may not be extended. If only the left area is extended by 1 sample, the left area of ​​the extended template may include a third reference sample line in addition to the three reference sample lines, and the top area of ​​the extended template may consist of three reference sample lines. The top left area of ​​the extended template may be defined as a 4x3 area. On the other hand, if the width of the current block is less than the height, the left area may not be extended, but the top area may be extended by N samples. If only the upper region is expanded by 1 sample, the left region of the expanded template may consist of three reference sample lines, and the upper region of the expanded template may include a third reference sample line in addition to the three reference sample lines. The upper left region of the expanded template may be defined as a 3x4 region.

[0107] Alternatively, if the current block is a square block, the extended template described above may be applied to the current block. If the current block is a non-square block, the extended template may not be applied to the current block.

[0108] Alternatively, if the current block is a square block, the extended template described above may be applied to the current block. If the current block is a non-square block, a template with an extended portion may be applied. The template with an extended portion is as described above.

[0109] Alternatively, if the width of the current block is greater than the height, the template of the current block may include at least one of a left region, a top region, or an upper left region. Here, the left region may be composed of two reference sample lines (e.g., the 0th and 1st reference sample lines). Meanwhile, the upper region may be composed of three reference sample lines (e.g., the 0th, 1st, and 2nd reference sample lines). The upper left region may be defined as a 2x3 region.

[0110] If the left region consists of two reference sample lines, a process of generating internal samples to replace unavailable samples may be involved, as discussed above. The generated internal samples may be one or more sample rows located at the leftmost side of the current block.

[0111] Alternatively, if the width of the current block is smaller than the height, the template of the current block may include at least one of a left region, a top region, or an upper left region. Here, the left region may be composed of three reference sample lines (e.g., the 0th, 1st, and 2nd reference sample lines). Meanwhile, the upper region may be composed of two reference sample lines (e.g., the 0th and 1st reference sample lines). The upper left region may be defined as a 3x2 region.

[0112] If the upper region consists of these two reference sample lines, a process of generating internal samples to replace unavailable samples may be involved, as discussed above. The generated internal samples may be one or more of the uppermost sample rows within the current block.

[0113] According to the present disclosure, by expanding or changing the application range of the filter in DIMD, an intra prediction mode closer to the original can be derived and the compression efficiency of a video encoder / decoder can be increased.

[0114] By adding up the magnitude of the variation for each intra prediction mode in the HoG for the left region, the HoG for the upper region, and the HoG for the upper left region, we can generate a HoG for the template.

[0115] The top M intra prediction modes with the largest values ​​of the magnitude of the variation in the HoG for the template can be derived, where M can be an integer of 1, 2, 3, 4, 5, or higher.

[0116] For example, the HoG for the left area (Left_HoG), the HoG for the upper area (Above_HoG), and the HoG for the upper left area (LeftAbove_HoG) can be generated as shown in Table 1 below.

[0117]

[0118] When the magnitude of the variation (Ampl) for each intra prediction mode (IPM) is added up in the HoG for each region in Table 1, the HoG for the template can be generated as in Table 2.

[0119]

[0120] The top five intra prediction modes with the largest values ​​of variation magnitude can be derived from the HoG for the template. According to Table 2, intra prediction modes with values ​​of 53, 44, 13, 64, and 49 can be derived.

[0121] The top M intra prediction modes can be derived by considering the frequency of the intra prediction modes derived by applying a filter to the template of the current block. Additionally, the frequency of outlier occurrence of the intra prediction mode can be reduced by performing filtering on the magnitude (Ampl) of the change in HoG. At this time, 3-tap filtering, interpolation filtering, etc. can be applied as the filtering. A specific directionality can be assigned to the top M intra prediction modes based on a predetermined reference mode, and in this case, a mode having a different directionality from the remaining modes among the top M intra prediction modes is called an outlier. For example, if the directionality is assigned based on mode 34, mode 13 among the top 5 intra prediction modes, i.e., {53, 55, 13, 65, 47}, has a different directionality from the other modes.

[0122] For example, frequencies for each intra prediction mode can be added to the HoG table, and the top M intra prediction modes can be derived based on these frequencies. Table 3 shows an example of a HoG table with added frequencies. Here, Freq can represent the frequency of the intra prediction mode.

[0123]

[0124] Alternatively, the top M intra prediction modes can be derived based on the magnitude (Ampl) and frequency (freq) of the variation for each intra prediction mode. In this case, the intra prediction mode can be determined based on the size of the current block and / or additional weights. Furthermore, filtering can be applied to the Ampl values ​​in the HoG table, and the intra prediction mode can be derived based on the Ampl values ​​after filtering.

[0125] Any one of the derivation methods defined identically for the encoder / decoder may be adaptively utilized. The derivation methods may include at least one of a derivation method based on the magnitude of the variation, a derivation method based on the frequency, or a derivation method based on both the magnitude and the frequency of the variation.

[0126] Which of the above derivation methods is used can be determined by comparing the bit rate-distortion cost based on the intra prediction mode derived through each derivation method. Alternatively, information indicating one of the above derivation methods can be encoded and signaled.

[0127] For example, a flag may be signaled indicating whether a derivation method based on the magnitude of the change is applied. If the flag is set to the first value, the derivation method based on the magnitude of the change may be applied. If the flag is set to the second value, the derivation method based on the frequency may be applied.

[0128] Alternatively, any one of the aforementioned derivation methods may be selectively utilized based on the size of the current block.

[0129] Prediction blocks can also be generated using only intra prediction modes less than or equal to the top M in the HoG, but using MRL (Multi-Reference Line). Here, MRL can be a prediction technique that selectively uses any one of the reference sample lines available to the current block. The reference sample lines available to the current block can be defined as the 0th to Lth reference sample lines, and L can be an integer greater than or equal to 1.

[0130] For each of the above-described intra prediction modes, a prediction block of the current block can be generated based on a single pre-selected reference sample line. Furthermore, for the same intra prediction mode, a prediction block can be generated based on two or more reference sample lines.

[0131] Alternatively, a DIMD merge HoG can be generated based on the DIMD HoG of the surrounding blocks of the current block. This HoG generation method is called DIMD merge mode. The DIMD HoG of the surrounding block may refer to the HoG used by the surrounding block to derive an intra prediction mode based on DIMD. The surrounding block may include at least one of the left block, the upper block, the upper left block, the lower left block, or the upper right block of the current block. The intra prediction mode can be derived based on the DIMD merge mode.

[0132] For example, the HoG of the current block (hereinafter referred to as a DIMD merge HoG) may be generated by averaging the DIMD HoGs of at least two neighboring blocks adjacent to the current block. The buffer used to generate the DIMD merge HoG may store the DIMD HoGs of at most N neighboring block(s), where N may be an integer greater than or equal to 1.

[0133] The RD (Rate-Distortion) cost between the prediction block generated through the general DIMD mode and the prediction block generated through the DIMD merge mode can be compared to select one of the two modes. Information indicating the selected mode can be signaled through the bitstream. The general DIMD mode may refer to a mode that generates an HoG by applying a filter to the template of the current block described above.

[0134] To reduce the memory complexity of DIMD merge mode, a modified HoG table generation method can be used to improve video encoder / decoder performance and reduce memory complexity. The HoG buffer in DIMD merge mode can store the Ampls of all HoGs of neighboring blocks. This can require high memory complexity. To further reduce memory complexity, the buffer can store only the Ampls for the top M intra prediction modes(es) in the HoGs of neighboring blocks.

[0135] The buffer storing the DIMD merge HoG can be generated using only the top M intra prediction modes within the buffer storing the DIMD HoG of the surrounding blocks. For example, only the Ampl values ​​for the top 5 intra prediction modes within the buffer storing the DIMD HoG of the surrounding blocks can be used.

[0136] An intra prediction mode can also be derived based on the DIMD fusion mode, which is a combination of the regular DIMD mode and the DIMD merge mode. By fusion between the regular DIMD mode and the DIMD merge mode, the performance of the video encoder / decoder can be improved by increasing the accuracy of the predicted block. Here, the fusion can be performed based on the weighted sum of the magnitudes of the changes (Ampl) between the HoG according to the regular DIMD mode and the HoG according to the DIMD merge mode.

[0137] Alternatively, the DIMD fusion mode may be defined as a mode that generates a final prediction block through a weighted sum between a prediction block generated through the general DIMD mode and a prediction block generated through the DIMD merge mode.

[0138] Information indicating whether the above DIMD fusion mode is applied may be encoded and signaled through a bitstream. If the information is a first value, the DIMD fusion mode is applied, and if the information is a second value, the DIMD fusion mode may not be applied. If the information is a second value, either the general DIMD mode or the DIMD merge mode may be applied instead of the DIMD fusion mode.

[0139] For example, the Ampl value of the DIMD fusion HoG can be generated by averaging the Ampl values ​​between the HoG of the general DIMD mode and the HoG of the DIMD merge mode. The DIMD fusion HoG can be calculated as shown in the following mathematical expression 1. Here, i can represent the number of the intra prediction mode.

[0140]

[0141] Alternatively, the Ampl value of the DIMD fusion HoG can be generated by performing a weighted average of the Ampl values ​​between the HoG of the general DIMD mode and the HoG of the DIMD merge mode. Here, the weight for the weighted average can be determined based on the number of surrounding blocks used to generate the HoG of the DIMD merge mode.

[0142] For example, when the number of surrounding blocks used to generate the HoG of the DIMD merge mode is 5, the Ampl value of the DIMD fusion HoG can be calculated as in the following mathematical expression 2. Here, W1 and W2 can have weight values ​​of 1 / 6 and 5 / 6, respectively.

[0143]

[0144] A HoG can be derived based on another region within the current picture that is a pre-restored region and does not belong to the aforementioned template, and an intra prediction mode can be derived based on the derived HoG. Here, the other region can be a reference block specified based on a block vector.

[0145] A block vector list for the current block can be generated. The block vector list can be composed of multiple block vector candidates. The block vector specifying the reference block can be derived based on any one of the multiple block vector candidates. Alternatively, two or more block vector candidates can be selected, and reference blocks corresponding to the selected block vector candidates can be specified, respectively.

[0146] The above block vector list can be generated based on the movement information (or block vector information) of surrounding blocks adjacent to the current block.

[0147] The above block vector list may be generated based on the intra-block copy (IBC) information of the surrounding blocks. If the surrounding blocks are encoded in the IBC mode, the surrounding blocks may have block vectors for the IBC mode. The IBC information of the surrounding blocks may include block vectors for the IBC mode of the surrounding blocks.

[0148] The above block vector list may be generated based on intra-template matching prediction (intra-TMP) information of surrounding blocks. Intra-TMP searches for an area matching or most similar to the template of the current block within a pre-restored area within the current picture, and determines a block using the area as a template as a reference block. At this time, the intra-TMP information may include a block vector for specifying the determined reference block.

[0149] The above block vector list may be generated based on at least two of motion information of surrounding blocks, block vector information, IBC information, or intra TMP information.

[0150] The aforementioned surrounding blocks may include at least one of a left block, an upper block, an upper left block, a lower left block, or an upper right block.

[0151] One block vector candidate can be selected from the block vector list, and an HoG for a reference block specified by the selected block vector candidate can be derived. Alternatively, two or more block vector candidates can be selected from the block vector list, and an HoG for each of the reference blocks specified by the selected block vector candidates can be derived.

[0152] For example, a filter can be applied to a reference block to extract a gradient, and an HoG can be generated based on the extracted gradient. This has been discussed previously, and a redundant description will be omitted here. In other words, the method for generating a HoG based on the aforementioned template can be applied in the same or similar manner when generating a HoG based on the reference block.

[0153] According to the present disclosure, the accuracy of the final prediction block can be improved by considering not only the surrounding area of ​​the current block but also the variation of the reference block.

[0154] Referring to FIG. 3, weights for the derived intra prediction mode can be derived (S310).

[0155] Weights for intra prediction modes can be derived based on the magnitude of the HoG variation (Ampl) relative to the template. For example, the weights can be derived as shown in Equation 3. Equation 3 assumes that five intra prediction modes are derived for the current block.

[0156]

[0157] In Equation 3, wDIMD1 to wDIMD5 may represent weights for the top five intra prediction modes derived based on DIMD. wDIMD6 may represent a weight for the planar mode. Table 4 below shows examples of weights for each intra prediction mode derived using Equation 3.

[0158]

[0159] The above-derived weights can be adjusted based on the dependency of regions within the template. The dependency can indicate which region among the left region, top region, and top-left region is more affected by the previously derived intra-prediction mode.

[0160] For example, if the top region is significantly affected, an index of 1 may be assigned, if the left region is significantly affected, an index of 2 may be assigned, and if the top and left regions are equally affected, an index of 0 may be assigned. The values ​​of the indices representing the above dependence may be defined as in the following mathematical expression 4.

[0161]

[0162] In mathematical expression 4, ampl LEFT represents the magnitude of the change corresponding to the intra prediction mode in the HoG for the left region, and ampl ABOVE represents the magnitude of the change corresponding to the intra prediction mode in the HoG for the upper region, and ampl TOTAL can represent the magnitude of the variation corresponding to the intra prediction mode in the HoG for the template. LocDep can be an index indicating the region on which the intra prediction mode depends.

[0163] According to mathematical formula 4, ampl LEFT Go (ampl TOTAL / 3), LocDep is set to 1, which may indicate that the corresponding intra prediction mode has a high dependence on the upper region. ampl ABOVE Go (ampl TOTAL / 3), LocDep is set to 2, which may indicate that the intra prediction mode has a high dependence on the left region. Otherwise (i.e., ampl LEFT and ampl ABOVE Both (ampl TOTAL / 3) is greater than or equal to ), LocDep is set to 0, which may indicate that the corresponding intra prediction mode has equal dependence on the left region and the top region.

[0164] For the top five intra prediction modes derived above, it is assumed that the HoG for the left region (Left_HoG), the HoG for the upper region (Above_HoG), and the HoG for the template (Total_HoG) are as shown in Table 5 below.

[0165]

[0166] According to Table 5, ampl for intra prediction mode 53 LEFT is (ampl TOTAL / 3) smaller than ampl ABOVE is (ampl TOTAL / 3), so the 53rd intra prediction mode has a high dependence on the upper region, and an index of 1 can be assigned to the 53rd intra prediction mode.

[0167] Through the aforementioned process, an index representing the degree of dependence can be derived for each intra prediction mode. For example, the indices for the top five intra prediction modes according to Table 5 can be derived as shown in Table 6.

[0168]

[0169] For intra prediction mode, the pre-derived weights (Weight) can be adjusted based on the index (LocDep). This can be used to generate the final predicted block for the current block. The weights applied to the predicted block can be designed to decrease as the distance from a highly dependent region increases. For example, weight adjustment can be performed as shown in Equation 5 below.

[0170]

[0171] Referring to FIG. 3, a final prediction block of the current block can be generated based on a prediction block generated based on an intra prediction mode and a weight corresponding to the intra prediction mode (S320).

[0172] As discussed above, when multiple intra prediction modes are derived for the current block in S300, prediction blocks corresponding to the multiple intra prediction modes can be generated respectively. At this time, the final prediction block of the current block can be generated through a weighted sum of the multiple prediction blocks for the current block. The weights for the weighted sum may be derived in S310. Alternatively, the final prediction block of the current block can be generated through a weighted sum between one or more prediction blocks for the current block and a prediction block generated based on the Planar mode. The weights for the weighted sum may be derived in S310.

[0173] For example, a final prediction block can be generated through a weighted sum of prediction blocks based on five pre-derived intra prediction modes and the planar mode for the current block. Equation 6 is an example showing a formula for generating the final prediction block.

[0174]

[0175] Planar mode can generate a prediction block based on at least two of four reference samples. Here, the four reference samples can include an upper reference sample, a left reference sample, a lower left reference sample, and an upper right reference sample of a sample that is a current prediction target (hereinafter referred to as the current sample). Planar mode can be divided into three modes. For example, it can be divided into a general planar mode using the four reference samples, a horizontal planar mode using the left reference sample and the upper right reference sample, and a vertical planar mode using the upper reference sample and the lower left reference sample.

[0176] In the method for generating a final prediction block based on DIMD according to the present disclosure, only the general Planar mode may be used as the Planar mode. Alternatively, in the method for generating a final prediction block based on DIMD according to the present disclosure, any one of the three modes described above may be selectively used. Alternatively, at least one of the horizontal Planar mode or the vertical Planar mode may be additionally used in addition to the general Planar mode.

[0177] DIMD is a method that implicitly derives an intra prediction mode for the current block in the encoder / decoder and generates a prediction block based on this. Multiple prediction blocks are generated based on M or more intra prediction modes derived from DIMD and planar modes, and the final prediction block is generated through a weighted sum. Here, M can be an integer greater than or equal to 1.

[0178] In the present disclosure, any one of a plurality of planar mode candidates can be selectively utilized based on a predetermined condition. Here, the plurality of planar mode candidates can include at least two of the aforementioned general planar mode, horizontal planar mode, or vertical planar mode. As the condition, an index indicating the aforementioned dependency (hereinafter, "LocDep"), the directionality of an intra prediction mode derived based on DIMD, etc. can be utilized.

[0179] Horizontal Planar mode and vertical Planar mode can use different or modified weights than the standard Planar mode. Alternatively, new weights can be created by combining weights.

[0180] For example, let's assume that the LocDep values ​​for each selected intra prediction mode are derived as shown in Table 7. In this case, if the proportion / frequency of LocDep with a value of 1 is greater, the vertical planar mode can be used. Conversely, if the proportion / frequency of LocDep with a value of 2 is greater, the horizontal planar mode can be used. Alternatively, if the proportion / frequency of LocDep with a value of 1 is greater, the horizontal planar mode can be used. Conversely, if the proportion / frequency of LocDep with a value of 2 is greater, the vertical planar mode can be used.

[0181]

[0182] When calculating the LocDep ratio for each intra prediction mode, additional weighting can be applied based on the difference between the height and width of the current block. For example, if the height of the current block is greater than its width, a LocDep value (i.e., 1) indicating a higher dependence on the left region can be given a greater weight. Conversely, if the width of the current block is greater than its height, a LocDep value (i.e., 2) indicating a higher dependence on the top region can be given an additional weight.

[0183] The top M intra prediction mode(s) derived based on DIMD can be classified into vertical directional modes and horizontal directional modes based on a given reference mode. Modes with a value smaller than the reference mode can be classified as horizontal directional modes, and modes with a value larger than the reference mode can be classified as vertical directional modes. Either the vertical planar mode or the horizontal planar mode can be used for the directional ratio of the top M intra prediction mode(s). Either the vertical planar mode or the horizontal planar mode can be selectively used based on whether the number of vertical directional modes among the top M intra prediction modes is greater than the number of horizontal directional modes.

[0184] For example, according to Table 7, when the reference mode is mode 34, three of the five intra prediction modes have vertical directionality, so the vertical Planar mode can be used.

[0185] When checking the directional ratio of intra prediction modes, if the ratio between the width and height of the current block is greater than K times, additional weighting may be applied to each intra prediction mode. Here, K may be an integer greater than or equal to 1.

[0186] For example, if the width of the current block is more than twice the height, the horizontally oriented modes that are smaller than the reference mode can be weighted twice because the correlation decreases as the distance between the left reference sample and the current sample in the current block increases along the x-axis.

[0187] Alternatively, if the width of the current block is more than twice the height, the correlation between the left reference sample and the current sample within the current block decreases as the distance along the x-axis increases. Therefore, to compensate for the left reference samples, the weights for horizontal directional modes corresponding to cases where the width is smaller than the reference mode can be doubled.

[0188] According to the present disclosure, the compression efficiency of a video encoder / decoder can be increased by imparting directionality to the Planar mode.

[0189] Alternatively, if the number of intra prediction modes derived from S300 is 1, the final prediction block can be generated using the intra prediction mode as a single mode without weighting with the planar mode. In this case, the weight derivation process of S310 can be omitted.

[0190] The final prediction block of the current block can also be generated by weighting the prediction blocks generated based on the general DIMD mode and the prediction blocks generated based on the DIMD merge mode. For example, the final prediction block can be generated by taking the average of each prediction block. The weights for the weighted sum can be determined based on the information from the surrounding blocks used to generate the HoG in the DIMD merge mode.

[0191] The method of generating the final prediction block described above can be performed identically in the encoder and decoder.

[0192] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combinations of two or more.

[0193] Additionally, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, the embodiments may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.

[0194] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.

Claims

1. A step of deriving an intra prediction mode for the current block by applying a filter to the template of the current block; A step of deriving weights for the intra prediction mode; and A method for decoding an image, comprising the step of generating a final prediction block of the current block based on the weight and the prediction block generated based on the intra prediction mode.

2. In paragraph 1, The above template is a surrounding area adjacent to the current block, A method for decoding an image, wherein the peripheral area includes at least one of a left area, an upper area, or an upper left area.

3. In paragraph 1, A method for decoding an image, wherein the range of the template to which the filter is applied is variably determined based on the height and width of the current block.

4. In paragraph 1, A method for decoding an image, wherein a center sample among reference samples input to the filter belongs to at least one of a first reference sample line located 1 sample away from the boundary of the current block or a second reference sample line located 2 samples away from the boundary of the current block.

5. In paragraph 1, If there is an unavailable sample among the reference samples input to the above filter, the unavailable sample is replaced with an available sample, A method for decoding an image, wherein the above available samples are generated based on a predetermined interpolation method or a predetermined intra prediction mode.

6. In paragraph 1, The step of deriving the intra prediction mode includes the step of generating a HoG table for the template of the current block, A method for decoding an image, wherein the above HoG table further includes a frequency for each intra prediction mode.

7. In paragraph 1, A method for decoding an image, wherein the step of deriving the intra prediction mode includes the step of generating a HoG of the current block based on a DIMD HoG of a surrounding block.

8. In paragraph 1, A method for decoding an image, wherein the step of deriving the intra prediction mode includes the step of generating an HoG for a reference block specified by a predetermined block vector.

9. In paragraph 1, A method for decoding an image, wherein the final prediction block is generated based on a weighted sum of a prediction block generated based on the intra prediction mode and a prediction block generated based on the planar mode.

10. In paragraph 9, An image decoding method, wherein the above Planar mode is adaptively induced into any one of a general Planar mode, a vertical Planar mode, or a horizontal Planar mode.

11. A step of applying a filter to the template of the current block to derive an intra prediction mode for the current block; A step of deriving weights for the intra prediction mode; and A video encoding method, comprising the step of generating a final prediction block of the current block based on the prediction block generated based on the intra prediction mode and the weight.

12. A computer-readable storage medium storing a bitstream generated based on the image encoding method according to Article 11.

Citation Information

Patent Citations

  • Method and apparatus for processing a video signal

    KR1020180005120A

  • Method and apparatus of improvement for decoder-derived intra prediction in video coding system

    WO2023198112A1

  • KR20200132829A