Video coding method and device, and recording medium storing bitstream
The multi-prediction and linear filter-based intra prediction methods improve video coding accuracy and efficiency by employing block vector-based and adaptive filtering techniques for high-resolution images.
Patent Information
- Application Number
- PCT/KR2025/099838
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-03-18
- Publication Date
- 2025-09-25
AI Technical Summary
Existing video compression technologies face challenges in achieving high prediction accuracy for high-resolution images, particularly in intra prediction methods, which are crucial for efficient video coding.
A multi-prediction based intra prediction method and device utilizing block vector-based and linear filter-based approaches to generate prediction blocks, incorporating adaptive weighting and multiple candidate lists for improved prediction accuracy.
Enhances prediction accuracy and compression efficiency by leveraging multiple prediction methods and adaptive filtering techniques, optimizing video coding for high-resolution images.
Smart Images

Figure KR2025099838_25092025_PF_FP_ABST
Abstract
Description
Video coding method and device, and recording medium storing bitstream
[0001] The present invention relates to a video signal processing method and device.
[0002] The market demand for high-resolution video is growing, necessitating technologies capable of efficiently compressing high-resolution images. To address this market need, the ISO / IEC's Moving Picture Expert Group (MPEG) and the ITU-T's Video Coding Expert Group (VCEG) jointly formed the Joint Collaborative Team on Video Coding (JCT-VC). They completed development of the HEVC (High Efficiency Video Coding) video compression standard in January 2013 and have been actively conducting research and development on next-generation compression standards.
[0003] Video compression largely consists of intraprediction, interprediction, transform, quantization, entropy coding, and in-loop filtering. Among these, intraprediction refers to a technique that generates a prediction block for the current block using reconstructed pixels surrounding the current block. The encoder encodes the intraprediction mode used for intraprediction, and the decoder performs intraprediction by reconstructing the encoded intraprediction mode.
[0004] The present disclosure provides a multi-prediction based intra prediction method and device.
[0005] The present disclosure provides a block vector-based intra prediction method and device.
[0006] The present disclosure provides a linear filter-based intra prediction method and device.
[0007] The video decoding method and device according to the present disclosure can generate a first prediction block of a current block, generate a second prediction block of the current block, and generate a prediction block of the current block based on the first prediction block and the second prediction block.
[0008] In the image decoding method and device according to the present disclosure, the first prediction block can be derived based on at least one of an intra prediction mode of a neighboring block of the current block or a size of the current block.
[0009] In the video decoding method and device according to the present disclosure, the first prediction block may be generated based on a predetermined intra prediction mode. Here, the predetermined intra prediction mode may be derived based on the amount of variation in sample values between reference samples belonging to a surrounding area of the current block.
[0010] In the video decoding method and device according to the present disclosure, the second prediction block may be generated based on a pre-defined mode, and the pre-defined mode may be a planar mode or a DC mode.
[0011] In the video decoding method and device according to the present disclosure, the second prediction block can be generated based on at least one of a plurality of candidates belonging to a block vector-based candidate list.
[0012] In the video decoding method and device according to the present disclosure, a candidate belonging to the block vector-based candidate list can be derived based on a block vector of a candidate block decoded before the current block. Here, the candidate block can include at least one of a neighboring block adjacent to the current block or a neighboring block not adjacent to the current block.
[0013] In the image decoding method and device according to the present disclosure, the second prediction block can be generated by applying a predetermined linear filter to a reference block specified by at least one of the plurality of candidates.
[0014] In the video decoding method and device according to the present disclosure, the prediction block of the current block can be generated through a weighted sum of the first prediction block and the second prediction block. Here, the weight for the weighted sum can be adaptively determined as one of the weights available to the current block.
[0015] The video encoding method and device according to the present disclosure can generate a first prediction block of a current block, generate a second prediction block of the current block, and generate a prediction block of the current block based on the first prediction block and the second prediction block.
[0016] A computer-readable recording medium according to the present disclosure can store a bitstream encoded by the image encoding method.
[0017] According to the present disclosure, prediction accuracy can be improved through intra prediction based on multiple predictions.
[0018] According to the present disclosure, prediction accuracy can be improved through block vector-based intra prediction.
[0019] According to the present disclosure, prediction accuracy can be improved by applying an additional linear filter to a prediction block.
[0020] FIG. 1 is a block diagram showing an image encoding device according to the present disclosure.
[0021] FIG. 2 is a block diagram showing an image decoding device according to the present disclosure.
[0022] FIG. 3 illustrates an intra prediction method according to the present disclosure.
[0023] FIG. 4 is an example according to the present disclosure, showing the locations of surrounding blocks available to the current block.
[0024] FIG. 5 is an example according to the present disclosure, showing candidate blocks available for the current block.
[0025] FIG. 6 illustrates a linear filter as an example according to the present disclosure.
[0026] The video decoding method and device according to the present disclosure can generate a first prediction block of a current block, generate a second prediction block of the current block, and generate a prediction block of the current block based on the first prediction block and the second prediction block.
[0027] In the image decoding method and device according to the present disclosure, the first prediction block can be derived based on at least one of an intra prediction mode of a neighboring block of the current block or a size of the current block.
[0028] In the video decoding method and device according to the present disclosure, the first prediction block may be generated based on a predetermined intra prediction mode. Here, the predetermined intra prediction mode may be derived based on the amount of variation in sample values between reference samples belonging to a surrounding area of the current block.
[0029] In the video decoding method and device according to the present disclosure, the second prediction block may be generated based on a pre-defined mode, and the pre-defined mode may be a planar mode or a DC mode.
[0030] In the video decoding method and device according to the present disclosure, the second prediction block can be generated based on at least one of a plurality of candidates belonging to a block vector-based candidate list.
[0031] In the video decoding method and device according to the present disclosure, a candidate belonging to the block vector-based candidate list can be derived based on a block vector of a candidate block decoded before the current block. Here, the candidate block can include at least one of a neighboring block adjacent to the current block or a neighboring block not adjacent to the current block.
[0032] In the image decoding method and device according to the present disclosure, the second prediction block can be generated by applying a predetermined linear filter to a reference block specified by at least one of the plurality of candidates.
[0033] In the video decoding method and device according to the present disclosure, the prediction block of the current block can be generated through a weighted sum of the first prediction block and the second prediction block. Here, the weight for the weighted sum can be adaptively determined as one of the weights available to the current block.
[0034] The video encoding method and device according to the present disclosure can generate a first prediction block of a current block, generate a second prediction block of the current block, and generate a prediction block of the current block based on the first prediction block and the second prediction block.
[0035] A computer-readable recording medium according to the present disclosure can store a bitstream encoded by the image encoding method.
[0036] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings attached to this specification so that those skilled in the art can easily implement the present invention. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In addition, in the drawings, parts irrelevant to the description have been omitted to clearly explain the present invention, and similar parts have been designated with similar reference numerals throughout the specification.
[0037] Throughout this specification, when a part is said to be 'connected' to another part, this includes not only cases where they are directly connected, but also cases where they are electrically connected with another element in between.
[0038] Additionally, whenever a part throughout this specification is said to "include" a component, this does not mean that other components are excluded, but rather that other components may be included, unless otherwise specifically stated.
[0039] Additionally, while terms such as first, second, etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another.
[0040] Additionally, in the embodiments of the devices and methods described herein, some components of the devices or some steps of the methods may be omitted. Furthermore, the order of some components of the devices or some steps of the methods may be changed. Furthermore, other components or other steps may be inserted into some components of the devices or some steps of the methods.
[0041] Additionally, some components or some steps of the first embodiment of the present invention may be added to the second embodiment of the present invention, or some components or some steps of the second embodiment may be replaced.
[0042] In addition, the components shown in the embodiments of the present invention are independently depicted to represent different characteristic functions, and this does not mean that each component is composed of separate hardware or a single software component. That is, each component is described by listing each component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function. Such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0043] In this specification, a block can be variously expressed as a unit, an area, a unit, a partition, etc., and a sample can be variously expressed as a pixel, a pel, a pixel, etc.
[0044] Hereinafter, embodiments of the present invention will be described in more detail with reference to the attached drawings. In describing the present invention, duplicate descriptions of identical components will be omitted.
[0045] FIG. 1 is a block diagram showing an image encoding device according to the present disclosure.
[0046] Referring to FIG. 1, a video encoding device (100) may include a picture segmentation unit (110), a prediction unit (120, 125), a transformation unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse transformation unit (145), a filter unit (150), and a memory (155).
[0047] The picture segmentation unit (110) can segment the input picture into at least one processing unit. At this time, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). Hereinafter, in the embodiments of the present disclosure, the coding unit may be used to mean a unit that performs encoding or a unit that performs decoding.
[0048] A prediction unit may be divided into at least one square or rectangular shape of the same size within a single coding unit, or may be divided such that one prediction unit among the divided prediction units within a single coding unit has a different shape and / or size from another prediction unit. When a prediction unit that performs intra prediction based on a coding unit is generated and is not the minimum coding unit, intra prediction can be performed without being divided into a plurality of NxN prediction units.
[0049] The prediction unit (120, 125) may include an inter prediction unit (120) that performs inter prediction or inter prediction, and an intra prediction unit (125) that performs intra prediction or intra prediction. It may determine whether to use inter prediction or intra prediction for a prediction unit, and determine specific information (e.g., intra prediction mode, motion vector, reference picture, etc.) according to each prediction method. A residual value (residual block) between the generated prediction block and the original block may be input to the transformation unit (130). In addition, prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value by the entropy encoding unit (165) and transmitted to the decoder.
[0050] The inter prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous or subsequent pictures of the current picture, and in some cases, may predict a prediction unit based on information of a portion of an encoded region within the current picture. The inter prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.
[0051] The reference picture interpolation unit can receive reference picture information from the memory (155) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.
[0052] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm) can be used to derive a motion vector. The motion vector can have a motion vector value in units of 1 / 2 or 1 / 4 pixels based on the interpolated pixel. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various methods such as Skip Mode, Merge Mode, AMVP Mode, Intra Block Copy Mode, and Affine Mode can be used as motion prediction methods.
[0053] The intra prediction unit (125) can generate a prediction unit based on reference pixel information surrounding the current block, which is pixel information within the current picture. If the surrounding block of the current prediction unit is a block on which inter prediction has been performed and the reference pixel is a pixel on which inter prediction has been performed, the reference pixel included in the block on which inter prediction has been performed can be replaced and used with reference pixel information of the surrounding block on which intra prediction has been performed. That is, if the reference pixel is not available, the unavailable reference pixel information can be replaced and used with at least one reference pixel among the available reference pixels.
[0054] Additionally, a residual block containing residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130).
[0055] In the transformation unit (130), the residual block including the residual value information of the prediction unit generated through the original block and the prediction unit (120, 125) can be transformed using a transformation method such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on the intra prediction mode information of the prediction unit used to generate the residual block.
[0056] The quantization unit (135) can quantize values converted to the frequency domain by the transformation unit (130). The quantization coefficients can vary depending on the block or the importance of the image. The values produced by the quantization unit (135) can be provided to the dequantization unit (140) and the reordering unit (160).
[0057] The rearrangement unit (160) can perform rearrangement of coefficient values for quantized residual values.
[0058] The rearrangement unit (160) can change a two-dimensional block-shaped coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the rearrangement unit (160) can change the two-dimensional block-shaped coefficient into a one-dimensional vector form by scanning from the DC coefficient to the coefficient of the high-frequency region using a zig-zag scan method. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional block-shaped coefficient in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficient in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, it is possible to determine which scan method among the zig-zag scan, the vertical scan, and the horizontal scan is to be used.
[0059] The entropy encoding unit (165) can perform entropy encoding based on the values produced by the rearrangement unit (160). Entropy encoding can use various encoding methods such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC). In this regard, the entropy encoding unit (165) can encode residual value coefficient information of the encoding unit from the rearrangement unit (160) and the prediction units (120, 125). In addition, according to the present disclosure, it is possible to signal and transmit information indicating that motion information is derived and used on the decoder side and information on a technique used to derive motion information.
[0060] The inverse quantization unit (140) and the inverse transformation unit (145) inversely quantize the values quantized in the quantization unit (135) and inversely transform the values transformed in the transformation unit (130). The residual values generated in the inverse quantization unit (140) and the inverse transformation unit (145) can be combined with the predicted prediction units predicted through the motion estimation unit, motion compensation unit, and intra prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.
[0061] The filter unit (150) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter). The deblocking filter may remove block distortion caused by boundaries between blocks in a restored picture. The offset correction unit may correct the offset from the original image on a pixel-by-pixel basis for the image on which deblocking has been performed. In order to perform offset correction for a specific picture, a method may be used in which the pixels included in the image are divided into a certain number of regions, the regions to be offset are determined, and the offset is applied to the regions, or the offset is applied by considering edge information of each pixel. The ALF (Adaptive Loop Filtering) may be performed based on a value obtained by comparing the filtered restored image with the original image. After dividing the pixels included in the image into a predetermined group, one filter to be applied to the group is determined, and filtering may be performed differentially for each group.
[0062] The memory (155) can store a restored block or picture produced through the filter unit (150), and the stored restored block or picture can be provided to the prediction unit (120, 125) when performing inter prediction.
[0063] FIG. 2 is a block diagram showing an image decoding device according to the present disclosure.
[0064] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (210), a rearrangement unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).
[0065] When a video bitstream is input to a video encoding device, the input bitstream can be decoded in the opposite procedure to that of the video encoding device.
[0066] The entropy decoding unit (210) can perform entropy decoding in a procedure opposite to that of the entropy encoding unit of the video encoder. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied in response to the method performed in the video encoder.
[0067] The entropy decoding unit (210) can decode information related to intra prediction and inter prediction performed in the encoder.
[0068] The reordering unit (215) can perform reordering based on the method by which the bitstream entropy-decoded by the entropy decoding unit (210) is reordered by the encoding unit. The coefficients expressed in the form of a one-dimensional vector can be reordered by restoring them back to coefficients in the form of a two-dimensional block.
[0069] The inverse quantization unit (220) can perform inverse quantization based on the quantization parameters provided by the encoder and the coefficient values of the rearranged block.
[0070] The inverse transform unit (225) can perform inverse transform, i.e., inverse DCT, inverse DST, and inverse KLT, on the transforms performed by the transform unit, i.e., DCT, DST, and KLT, on the quantization result performed by the image encoder. The inverse transform can be performed based on the transmission unit determined by the image encoder. In the inverse transform unit (225) of the image decoder, a transform technique (e.g., DCT, DST, KLT) can be selectively performed according to a plurality of pieces of information, such as a prediction method, the size of the current block, and the prediction direction.
[0071] The prediction unit (230, 235) can generate a prediction block based on prediction block generation related information provided from the entropy decoding unit (210) and previously decoded block or picture information provided from the memory (245).
[0072] As described above, when performing intra prediction or intra prediction in the same manner as the operation in the image encoder, if the size of the prediction unit and the size of the transformation unit are the same, intra prediction for the prediction unit is performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side. However, when performing intra prediction, if the size of the prediction unit and the size of the transformation unit are different, intra prediction can be performed using reference pixels based on the transformation unit. In addition, intra prediction using NxN division only for the minimum coding unit can be used.
[0073] The prediction unit (230, 235) may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (210), prediction mode information of an intra prediction method, and motion prediction-related information of an inter prediction method, and may distinguish a prediction unit from a current encoding unit and determine whether the prediction unit performs inter prediction or intra prediction. On the other hand, if the encoder (100) does not transmit motion prediction-related information for the inter prediction, but instead transmits information indicating that motion information is to be derived and used on the decoder side and information on a technique used to derive motion information, the prediction unit determination unit determines whether the inter prediction unit (230) performs prediction based on the information transmitted from the encoder (100).
[0074] The inter prediction unit (230) can perform inter prediction on the current prediction unit based on information included in at least one picture among the previous picture or the subsequent picture of the current picture including the current prediction unit, using information required for inter prediction of the current prediction unit provided by the image encoder. In order to perform inter prediction, it can be determined based on the encoding unit whether the motion prediction method of the prediction unit included in the corresponding encoding unit is one of Skip Mode, Merge Mode, AMVP Mode, Intra Block Copy Mode, and Affine Mode.
[0075] The intra prediction unit (235) can generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit that has performed intra prediction, intra prediction can be performed based on intra prediction mode information of the prediction unit provided by the image encoder.
[0076] The intra prediction unit (235) may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a unit that performs filtering on the reference pixels of the current block and can determine whether to apply the filter based on the prediction mode of the current prediction unit and apply it. AIS filtering can be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the image encoder. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.
[0077] The reference pixel interpolation unit can interpolate the reference pixel to generate a reference pixel of a pixel unit less than an integer value when the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on the pixel value interpolated from the reference pixel. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.
[0078] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.
[0079] Information about whether a deblocking filter has been applied to a corresponding block or picture can be received from a video encoding device, and if a deblocking filter has been applied, information about whether a strong or weak filter has been applied. The deblocking filter of the video decoder can receive information related to the deblocking filter provided by the video encoder, and the video decoder can perform deblocking filtering on the corresponding block.
[0080] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and information on the offset value. ALF can be applied to the encoding unit based on information on whether ALF is applied and ALF coefficient information provided from the encoder. This ALF information can be provided by being included in a specific parameter set.
[0081] The memory (245) can store a restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to an output unit.
[0082] FIG. 3 illustrates an intra prediction method according to the present disclosure.
[0083] Referring to FIG. 3, a first prediction block of the current block can be generated (S300).
[0084] One or more intra prediction modes can be derived for the current block, and a first prediction block can be generated based on the derived intra prediction modes. Hereinafter, a method for deriving intra prediction modes for the current block will be described.
[0085] Example 1
[0086] A histogram of occurrences (HoC) can be generated based on at least one of the intra prediction mode of the surrounding block or the amplitude for the surrounding block.
[0087] According to the present disclosure, a neighboring block may include at least one of a neighboring block that is spatially / temporally adjacent to the current block or a neighboring block that is not adjacent to the current block. Here, the adjacent block may include at least one of a left neighboring block, an upper neighboring block, an upper-left neighboring block, an upper-right neighboring block, or a lower-left neighboring block of the current block. For example, the positions of neighboring blocks available to the current block may be illustrated as in FIG. 4. The numbers in FIG. 4 may indicate priorities of the available neighboring blocks.
[0088] A neighboring block can have one or more intra prediction modes. For example, if a neighboring block is encoded with Decoder-side Intra Mode Derivation (DIMD), the neighboring block can have five intra prediction modes. If a neighboring block is encoded with Template-based Intra Mode Derivation (TIMD) mode, the neighboring block can have two intra prediction modes. If a neighboring block is encoded with Spatial Geometric Partitioning Mode (SGPM) mode, the neighboring block can have two intra prediction modes. The HoC may be generated using all of the intra prediction modes of the neighboring block, or may be generated using only some of the intra prediction modes (e.g., one intra prediction mode) of the intra prediction modes of the neighboring block. If a neighboring block is encoded in Matrix-based Intra Prediction (MIP) mode, Intra Template Matching Prediction (Intra TMP) mode, and / or Intra Block Copy (IBC) mode, the neighboring block may not be used to generate the HoC.
[0089] The x-axis and y-axis of the HoC according to the present disclosure can be defined by the intra prediction mode and amplitude, respectively. That is, the HoC can also be defined as a candidate list including multiple candidates. Each candidate can be defined by the intra prediction mode and amplitude derived from the surrounding blocks. Here, the amplitude can represent the block size, and the block size can be defined by the width, the height, the product of the width and the height, the sum of the width and the height, etc.
[0090] At most N intra prediction modes can be derived from HoC, where N can be an integer of 1, 2, 3, 4, 5, or a larger number. For example, the top five intra prediction modes with the largest amplitudes can be derived from HoC. N prediction blocks can be generated based on the N intra prediction modes, and a first prediction block can be generated through a weighted sum of the N prediction blocks.
[0091] Example 2
[0092] An intra prediction mode can be derived based on the gradient of sample values between reference samples in the surrounding area of the current block, and this is called DIMD (Decoder-side Intra Mode Derivation). DIMD can extract a gradient by applying a filter of a predetermined size to the surrounding area of the current block (hereinafter referred to as a template). A Histogram of Gradient (HoG) can be generated based on the extracted gradient. At this time, the y-axis of the HoG is defined as the sum of the absolute values of the horizontal gradient (gradient x, Gx) and the vertical gradient (gradient y, Gy), and the x-axis of the HoG can be defined as an intra prediction mode that is mapped to the values and signs of Gx and Gy based on a pre-defined table. The HoG can be a method of expressing the features of an image or video as a histogram of the gradient according to the direction. One or more intra prediction modes can be derived for the current block based on the HoG. Based on one or more intra prediction modes, one or more prediction blocks can be generated for the current block. If multiple intra prediction modes are derived for the current block, a first prediction block of the current block can be generated through a weighted sum of the multiple prediction blocks for the current block. This process can be performed identically in the decoder and the encoder, and a method for deriving intra prediction modes based on DIMD will be described in detail below.
[0093] An intra prediction mode for the current block can be derived by applying a filter to a template of the current block. The template according to the present disclosure may be a surrounding region adjacent to the current block. The surrounding region may include at least one of a left region, an upper region, or an upper left region. For example, if the left region is not available (or the current block borders the left edge of the current picture), the filter may be applied only to the upper region. Alternatively, if the upper region is not available (or the current block borders the upper edge of the current picture), the filter may be applied only to the left region.
[0094] Hereinafter, for convenience of explanation, it is assumed that the template includes a left region, a top region, and an upper left region. A HoG buffer for the template can be created / initialized. The HoG buffer can be created / initialized for each region belonging to the template. For example, the HoG buffer can include a left HoG buffer for the left region, an upper HoG buffer for the top region, an upper left HoG buffer for the upper left region, and a full HoG buffer for the entire template region. The left HoG buffer can store the HoG for the left region, the upper HoG buffer can store the HoG for the top region, the upper left HoG buffer can store the HoG for the upper left region, and the full HoG buffer can store the HoG for the template.
[0095] A HoG can be created by applying filters to each area within a template. Filters can be applied in the following order: left area, top area, top left area. However, this is merely an example, and the order of filter application is not limited thereto. For example, filters can be applied in the following order: left area, top left area, top area. Alternatively, filters can be applied in the following order: top area, top left area, left area. Alternatively, filters can be applied in the following order: top area, left area, top left area.
[0096] By applying a filter to each region of the template, the values of Gx and Gy can be derived. Based on the derived values of Gx and Gy, the intra prediction mode and the magnitude of the variation (Ampl) stored in the corresponding HoG buffer can be derived.
[0097] The range of the template to which the filter is applied can be determined based on the height and width of the current block. For example, the height of the left region may be the same as the height of the current block. The width of the top region may be the same as the width of the current block. However, this is not limited thereto. If there are available reference sample(s) in the lower left peripheral area and / or the upper right peripheral area of the current block, the range of application of the filter can be expanded. At this time, the expanded range can have a size of N-samples. Here, N can be an integer of 1, 2, 3, 4, or more. For example, if there are available reference sample(s) in the lower left peripheral area of the current block, the left region of the template can be expanded downward by a size of 4-samples. If there are available reference sample(s) in the upper right peripheral area of the current block, the upper region of the template can be expanded rightward by a size of 4-samples.
[0098] Among the reference samples input to the above filter, the center sample may belong to the 1st reference sample line that is 1 sample away from the boundary of the current block. That is, if the reference sample line adjacent to the current block is called the 0th reference sample line, the 1st reference sample line may be adjacent to the 0th reference sample line. Hereinafter, the kth reference sample line may be defined as being adjacent to the (k-1)th reference sample line. The above filter may utilize the weights of existing filters such as a Sobel filter and a Gaussian filter, or may utilize arbitrary weights according to the user.
[0099] According to the present disclosure, an intra prediction mode can be derived that can generate a prediction block closer to the original image by extending or changing a template to which a filter is applied.
[0100] For example, the template may include at least one of a left region, a top region, or an upper left region. Here, the left region may be composed of three reference sample lines (i.e., a 0th reference sample line, a 1st reference sample line, and a 2nd reference sample line). The lengths of the three reference sample lines may be equal to the height of the current block. The top region may be composed of three reference sample lines (i.e., a 0th reference sample line, a 1st reference sample line, and a 2nd reference sample line). The lengths of the three reference sample lines may be equal to the width of the current block. The upper left region may be defined as a 3x3 region.
[0101] The above template can be extended by N samples. N can be an integer greater than or equal to 1. For example, if the template is extended by 1 sample, the left region of the extended template can include a third reference sample line in addition to the three reference sample lines. The upper region of the extended template can include a third reference sample line in addition to the three reference sample lines. The upper left region of the extended template can be defined as a 4x4 region.
[0102] Among the reference samples input to the filter, the center sample may include a reference sample belonging to a first reference sample line that is 1 sample away from the boundary of the current block and a reference sample belonging to a second reference sample line that is 2 samples away from the boundary of the current block. Alternatively, among the reference samples input to the filter, the center sample may be at least one reference sample belonging only to the second reference sample line that is 2 samples away from the boundary of the current block. Alternatively, among the reference samples input to the filter, the center sample may be at least one reference sample belonging to a 0th reference sample line adjacent to the boundary of the current block.
[0103] If there are unavailable samples among the reference samples input to the filter (for example, if an internal sample of the current block that has not been decoded / encoded must be input to the filter), the unavailable samples can be replaced with samples generated through a predetermined method.
[0104] For example, unavailable samples can be replaced with samples generated through a predetermined interpolation. Here, the interpolation can be performed based on a method such as nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation.
[0105] Alternatively, the internal samples of the current block that have not been decoded / encoded can be generated based on predetermined intra prediction mode(s), and the generated internal samples can be input to the filter. Here, the predetermined intra prediction mode(s) can include at least one of a planar mode, a DC mode, a vertical mode, and a horizontal mode. Any one of the aforementioned predetermined intra prediction mode(s) can be selectively used based on the size of the current block. The size of the current block can be defined by the width, the height, the ratio of the width and the height, the product of the width and the height, the maximum / minimum values of the width and the height, etc.
[0106] Alternatively, the internal samples of the current block that have not been decoded / encoded may be generated based on a combination of at least two of the aforementioned intra prediction modes. For example, the internal samples may be generated as a weighted sum between samples generated based on a vertical mode and samples generated based on a horizontal mode.
[0107] Whether a template is expanded may be determined based on whether the current block is a square block. Alternatively, the scope / size of a template for DIMD may be determined differently depending on whether the current block is a square block. Alternatively, the scope / size of a template for DIMD may be determined differently depending on whether the width of the current block is greater than its height.
[0108] For example, if the current block is a square block, the extended template may not be applied to the current block. If the current block is a non-square block, the extended template described above may be applied to the current block.
[0109] Alternatively, if the current block is a square block, the extended template may not be applied to the current block. If the current block is a non-square block, the extended template may be applied to some areas. For example, if the width of the current block is greater than the height, the left area may be extended by N samples, but the top area may not be extended. If only the left area is extended by 1 sample, the left area of the extended template may include a third reference sample line in addition to the three reference sample lines, and the top area of the extended template may consist of three reference sample lines. The top left area of the extended template may be defined as a 4x3 area. On the other hand, if the width of the current block is less than the height, the left area may not be extended, but the top area may be extended by N samples. If only the upper region is expanded by 1 sample, the left region of the expanded template may consist of three reference sample lines, and the upper region of the expanded template may include a third reference sample line in addition to the three reference sample lines. The upper left region of the expanded template may be defined as a 3x4 region.
[0110] Alternatively, if the current block is a square block, the extended template described above may be applied to the current block. If the current block is a non-square block, the extended template may not be applied to the current block.
[0111] Alternatively, if the current block is a square block, the extended template described above may be applied to the current block. If the current block is a non-square block, a template with an extended portion may be applied. The template with an extended portion is as described above.
[0112] Alternatively, if the width of the current block is greater than the height, the template of the current block may include at least one of a left region, a top region, or an upper left region. Here, the left region may be composed of two reference sample lines (e.g., the 0th and 1st reference sample lines). Meanwhile, the upper region may be composed of three reference sample lines (e.g., the 0th, 1st, and 2nd reference sample lines). The upper left region may be defined as a 2x3 region.
[0113] If the left region consists of two reference sample lines, a process of generating internal samples to replace unavailable samples may be involved, as discussed above. The generated internal samples may be one or more sample rows located at the leftmost side of the current block.
[0114] Alternatively, if the width of the current block is smaller than the height, the template of the current block may include at least one of a left region, a top region, or an upper left region. Here, the left region may be composed of three reference sample lines (e.g., the 0th, 1st, and 2nd reference sample lines). Meanwhile, the upper region may be composed of two reference sample lines (e.g., the 0th and 1st reference sample lines). The upper left region may be defined as a 3x2 region.
[0115] If the upper region consists of these two reference sample lines, a process of generating internal samples to replace unavailable samples may be involved, as discussed above. The generated internal samples may be one or more of the uppermost sample rows within the current block.
[0116] According to the present disclosure, by expanding or changing the application range of the filter in DIMD, an intra prediction mode closer to the original can be derived and the compression efficiency of a video encoder / decoder can be increased.
[0117] By adding up the magnitude of the variation for each intra prediction mode in the HoG for the left region, the HoG for the upper region, and the HoG for the upper left region, we can generate a HoG for the template.
[0118] The top M intra prediction modes with the largest values of the magnitude of the variation in the HoG for the template can be derived, where M can be an integer of 1, 2, 3, 4, 5, or higher.
[0119] For example, the HoG for the left area (Left_HoG), the HoG for the upper area (Above_HoG), and the HoG for the upper left area (LeftAbove_HoG) can be generated as shown in Table 1 below.
[0120]
[0121] When the magnitude of the variation (Ampl) for each intra prediction mode (IPM) is added up in the HoG for each region in Table 1, the HoG for the template can be generated as in Table 2.
[0122]
[0123] The top five intra prediction modes with the largest values of variation magnitude can be derived from the HoG for the template. According to Table 2, intra prediction modes with values of 53, 44, 13, 64, and 49 can be derived.
[0124] Weights for the above-described intra prediction modes can be derived. Weights for the intra prediction modes can be derived based on the magnitude of the HoG variation (Ampl) relative to the template. For example, the weights can be derived as shown in Equation 3. Equation 1 assumes that five intra prediction modes are derived for the current block.
[0125]
[0126] In Equation 1, wDIMD1 to wDIMD5 may represent weights for the top five intra prediction modes derived based on DIMD. wDIMD6 may represent a weight for the planar mode. Table 3 below shows examples of weights for each intra prediction mode derived using Equation 1.
[0127]
[0128] The above-derived weights can be adjusted based on the dependency of regions within the template. The dependency can indicate which region among the left region, top region, and top-left region is more affected by the previously derived intra-prediction mode.
[0129] For example, if the top region is significantly affected, an index of 1 may be assigned, if the left region is significantly affected, an index of 2 may be assigned, and if the top and left regions are equally affected, an index of 0 may be assigned. The values of the indices representing the above dependencies may be defined as in the following mathematical expression 2.
[0130]
[0131] In Equation 2, amplLEFT may represent the magnitude of the variation corresponding to the corresponding intra prediction mode in the HoG for the left region, amplABOVE may represent the magnitude of the variation corresponding to the corresponding intra prediction mode in the HoG for the upper region, and amplTOTAL may represent the magnitude of the variation corresponding to the corresponding intra prediction mode in the HoG for the template. LocDep may be an index indicating the region on which the corresponding intra prediction mode depends.
[0132] According to Equation 2, when amplLEFT is less than (amplTOTAL / 3), LocDep is set to 1, which may indicate that the corresponding intra prediction mode has a high dependency on the upper region. When amplABOVE is less than (amplTOTAL / 3), LocDep is set to 2, which may indicate that the corresponding intra prediction mode has a high dependency on the left region. Otherwise (i.e., when both amplLEFT and amplABOVE are greater than or equal to (amplTOTAL / 3)), LocDep is set to 0, which may indicate that the corresponding intra prediction mode has an equal dependency on the left region and the upper region.
[0133] For the top five intra prediction modes derived above, it is assumed that the HoG for the left region (Left_HoG), the HoG for the upper region (Above_HoG), and the HoG for the template (Total_HoG) are as shown in Table 4 below.
[0134]
[0135] According to Table 4, amplLEFT for intra prediction mode 53 is less than (amplTOTAL / 3) and amplABOVE is equal to (amplTOTAL / 3), so intra prediction mode 53 has a high dependence on the upper region, and an index of 1 can be assigned to intra prediction mode 53.
[0136] Through the aforementioned process, an index representing the degree of dependence can be derived for each intra prediction mode. For example, the indices for the top five intra prediction modes according to Table 4 can be derived as shown in Table 5.
[0137]
[0138] For intra prediction mode, the derived weights (Weight) can be adjusted based on the index (LocDep). This can be used to generate a prediction block for the current block. The weights applied to the prediction block can be designed to decrease as the block moves away from a highly dependent region. For example, weight adjustment can be performed as shown in Equation 3 below.
[0139]
[0140] A first prediction block can be generated based on a prediction block generated based on an intra prediction mode and a weight corresponding to the intra prediction mode.
[0141] As previously discussed, when multiple intra prediction modes are derived for the current block, prediction blocks corresponding to the multiple intra prediction modes can be generated. At this time, the multiple prediction blocks for the current block can be weighted and combined based on the weights to generate the first prediction block of the current block.
[0142] Referring to FIG. 3, a second prediction block of the current block can be generated (S310).
[0143] The second prediction block according to the present disclosure can be generated based on a predefined mode. Here, the predefined mode can be a non-directional mode, such as a planar mode or a DC mode. Alternatively, the predefined mode can be a directional mode, such as a horizontal mode, a vertical mode, or a diagonal mode.
[0144] For example, a second prediction block may be generated based on a planar mode. The planar mode may generate a prediction block based on at least two of four reference samples. Here, the four reference samples may include an upper reference sample, a left reference sample, a lower left reference sample, and an upper right reference sample of a sample that is a current prediction target (hereinafter, referred to as a current sample). The planar mode may be divided into three modes. For example, the planar mode may be divided into a general planar mode using the four reference samples, a horizontal planar mode using the left reference sample and the upper right reference sample, and a vertical planar mode using the upper reference sample and the lower left reference sample.
[0145] When a first prediction block is generated based on DIMD according to the present disclosure, only the general planar mode may be used as the planar mode. Alternatively, when a first prediction block is generated based on DIMD according to the present disclosure, any one of the three modes described above may be selectively used.
[0146] In the present disclosure, any one of a plurality of planar mode candidates can be selectively utilized based on a predetermined condition. The plurality of planar mode candidates may include at least two of the aforementioned general planar mode, horizontal planar mode, or vertical planar mode. As the condition, an index indicating the aforementioned dependency (hereinafter, "LocDep"), the directionality of an intra prediction mode derived based on DIMD, and the like can be utilized.
[0147] Horizontal and vertical planner modes can use different weights than the standard planner mode. Alternatively, new weights can be created by combining weights.
[0148] For example, let's assume that the LocDep values for each selected intra prediction mode are derived as shown in Table 6. In this case, if the proportion / frequency of LocDep with a value of 1 is greater, the vertical planar mode can be used. Conversely, if the proportion / frequency of LocDep with a value of 2 is greater, the horizontal planar mode can be used. Alternatively, if the proportion / frequency of LocDep with a value of 1 is greater, the horizontal planar mode can be used. Conversely, if the proportion / frequency of LocDep with a value of 2 is greater, the vertical planar mode can be used.
[0149]
[0150] When calculating the LocDep ratio for each intra prediction mode, additional weighting can be applied based on the difference between the height and width of the current block. For example, if the height of the current block is greater than its width, a LocDep value (i.e., 1) indicating a higher dependence on the left region can be given a greater weight. Conversely, if the width of the current block is greater than its height, a LocDep value (i.e., 2) indicating a higher dependence on the top region can be given an additional weight.
[0151] The top M intra prediction mode(s) derived based on DIMD can be classified into vertical directional modes and horizontal directional modes based on a given reference mode. Modes having a value smaller than that of the reference mode can be classified as horizontal directional modes, and modes having a value larger than that of the reference mode can be classified as vertical directional modes. Either the vertical planar mode or the horizontal planar mode can be used for the directional ratio of the top M intra prediction mode(s). Either the vertical planar mode or the horizontal planar mode can be selectively used based on whether the number of vertical directional modes among the top M intra prediction modes is greater than the number of horizontal directional modes.
[0152] For example, according to Table 6, when the reference mode is mode 34, three of the five intra prediction modes have vertical directionality, so the vertical planar mode can be used.
[0153] When checking the directional ratio of intra prediction modes, if the ratio between the width and height of the current block is greater than K times, additional weighting may be applied to each intra prediction mode. Here, K may be an integer greater than or equal to 1.
[0154] For example, if the width of the current block is more than twice the height, the horizontally oriented modes that are smaller than the reference mode can be weighted twice because the correlation decreases as the distance between the left reference sample and the current sample in the current block increases along the x-axis.
[0155] Alternatively, if the width of the current block is more than twice the height, the correlation decreases as the distance between the left reference sample and the current sample within the current block increases along the x-axis. Therefore, to compensate for the left reference samples, the weights for horizontal directional modes corresponding to cases where the width is smaller than the reference mode can be doubled.
[0156] According to the present disclosure, the compression efficiency of a video encoder / decoder can be increased by giving directionality to the planar mode.
[0157] The second prediction block according to the present disclosure may also be generated through block vector-based prediction. Block vector-based prediction can generate a block vector-based candidate list, and generate the second prediction block based on at least one of the multiple candidates included in the block vector-based candidate list. The encoder and decoder can generate the block vector-based candidate list using the same method.
[0158] At least one of the plurality of candidates in the candidate list based on block vectors may be derived based on the block vector of a block (hereinafter, referred to as a candidate block) that has been encoded / decoded before the current block. The candidate block may include at least one of a neighboring block adjacent to the current block or a neighboring block that is not adjacent to the current block. For example, candidate blocks available to the current block may be illustrated as in FIG. 5. In FIG. 5, the left drawing illustrates a candidate block adjacent to the current block, and the right drawing illustrates a candidate block that is not adjacent to the current block. The numbers in FIG. 5 may indicate priorities among the candidate blocks.
[0159] Alternatively, at least one of a plurality of candidates belonging to a candidate list based on a block vector may be derived by performing template matching based on a template of a current block within a predetermined search range. The predetermined search range may belong to a region that has been previously restored before the current block within a picture to which the current block belongs. The template of the current block is a surrounding region adjacent to the current block, wherein the surrounding region may include at least one of an upper surrounding region, a left surrounding region, or an upper-left surrounding region. A predetermined cost may be calculated for each search position within the search range, and a candidate may be derived based on a block vector corresponding to the top M costs in ascending order of the calculated costs. Here, M may be an integer greater than or equal to 1. The cost may be a difference between the template of the current block and the template at the search position.
[0160] The block vector-based candidate list according to the present disclosure may further include one or more predefined modes. The predefined modes may include planar modes. The predefined modes may be inserted as the first candidate (e.g., a candidate with an index of 0) in the block vector-based candidate list. Alternatively, the predefined modes may be inserted as the last candidate in the block vector-based candidate list.
[0161] The encoder can encode an index indicating at least one candidate among a plurality of candidates in a block vector-based candidate list and signal the encoded index to the decoder. The decoder can select at least one candidate among the candidates in the block vector-based candidate list based on the signaled index, and generate a second prediction block based on the selected candidate.
[0162] Alternatively, a cost may be calculated for each candidate in a candidate list based on a block vector. Here, the cost may be the difference between the predicted value and the restored value of the surrounding area of the current block (e.g., SAD, SATD). The predicted value of the surrounding area may be derived based on the predicted information of the candidate (i.e., block vector information, mode information). The size of the surrounding area may be NxM, and N and M may be integers greater than or equal to 1. In addition, the number of surrounding areas used to calculate the cost may be K or more, and K may be an integer greater than or equal to 1. For example, if the current block is an 8x8 block, the surrounding area may include at least one of the left area of 4x8 or the top area of 8x4. A candidate with the smallest cost among the calculated costs may be selected, and a second prediction block may be generated based on the selected candidate.
[0163] A reference block can be specified based on the block vector of the selected candidate, and a second prediction block can be generated based on the specified reference block. If the selected candidate is in planar mode, the second prediction block can be generated based on the planar mode.
[0164] The second prediction block according to the present disclosure may be generated as a weighted sum between a prediction block generated based on the aforementioned planar mode and a prediction block generated through prediction based on the aforementioned block vector.
[0165] For example, if the planar mode is derived from the surrounding blocks of the current block, block vector-based prediction can be used together.
[0166] As discussed above, a second prediction block may be generated based on at least one of planar mode or block vector-based prediction, and a linear filter may additionally be applied to the second prediction block.
[0167] Specifically, the filter coefficients of a linear filter can be derived based on the template of the current block and the template of the reference block. Here, the reference block can be identified based on any one of multiple candidates belonging to a block vector-based candidate list. The filter coefficients of the linear filter can be applied to the reference block to generate a second prediction block.
[0168] The filter coefficients of the linear filter can be derived based on at least one of a current sample, a reference sample, or a neighboring sample adjacent to the reference sample. Here, the current sample may refer to a sample belonging to a template area of the current block. The reference sample may be a sample belonging to the template area of the reference block, and may be a sample corresponding to the position of the current sample. The neighboring samples adjacent to the reference sample may include at least one of an upper neighboring sample, a left neighboring sample, a lower neighboring sample, or a right neighboring sample.
[0169] For example, all or some of the samples within the template region of the reference block can be used as the reference samples. If there are a total of 64 samples within the template region of the reference block, the optimal filter coefficients can be derived through the Gaussian Elimination technique as shown in the following mathematical expression 4.
[0170]
[0171] In mathematical expression 4, C i represents the value of the reference sample, and N i , S i , E i and W i may represent the values of the upper, lower, right, and left surrounding samples of the corresponding reference sample, respectively (where i is an integer greater than or equal to 0 and less than or equal to 63). B may be a predetermined offset. For example, B may be a value set based on the bit depth. c0 to c5 may be filter coefficients applied to the reference sample, surrounding samples, and offset, respectively. ref i can represent the value of the current sample corresponding to the reference sample.
[0172] A second prediction block can be generated by applying a linear filter to samples of a reference block in a predetermined scan order (e.g., raster scan, TZ-search order).
[0173] A linear filter according to the present disclosure may be an N-tap linear filter. Here, N may be a fixed value that is identically predefined for an encoder and a decoder. For example, N may be 5. However, the present invention is not limited thereto, and N may be an integer of 3, 6, 7, 9, or a higher integer. A plurality of linear filters may be defined in the encoder and the decoder, and any one of the plurality of linear filters may be selectively used. The plurality of pre-filters may have different numbers of taps. Alternatively, the plurality of linear filters may have the same number of taps but different filter coefficients. Alternatively, the plurality of linear filters may have different numbers of taps and filter coefficients.
[0174] Figure 6 illustrates an example of a linear filter. All or some of the filter types illustrated in Figure 6 may be included in the plurality of linear filters described above. Various linear filters may be defined depending on the filter application area (or the positions of samples input to the linear filter). For example, a linear filter that inputs surrounding samples located above, below, left, and right of the current sample may be used. Alternatively, a linear filter that inputs surrounding samples located above, below, left, and right of the current sample may be used.
[0175] The filter application area may be expanded or reduced based on the size of the current block and / or the reference block. Here, the reference block may refer to a block that belongs to the current picture and has been encoded / decoded before the current block. The filter application area may be expanded or reduced based on the height and width ratio of the current block and / or the surrounding blocks. Alternatively, the filter application area may be adjusted to an asymmetrical area based on the height and width ratio of the current block and / or the surrounding blocks.
[0176] A flag may be explicitly signaled to indicate whether a linear filter is applied when generating the second prediction block. Alternatively, a linear filter may be implicitly applied without explicitly signaling a flag.
[0177] The method for generating the second prediction block described above can be performed in the same manner in the encoder and decoder.
[0178] Referring to FIG. 3, a prediction block of the current block can be generated based on the first prediction block and the second prediction block of the current block (S320).
[0179] A prediction block of the current block can be generated through a weighted sum between the first prediction block and the second prediction block.
[0180] The weights available for the above weighted sum may include at least one of 4 / 64, 8 / 64, 16 / 64, or 32 / 64. The weighted sum may be performed based on any one of the available weights.
[0181] For example, prediction blocks can be generated based on available weights, and one of the available weights can be implicitly derived by comparing the Rate-Distortion (RD) cost. Alternatively, the encoder can encode an index indicating one of the available weights and signal it to the decoder. The decoder can then specify one of the available weights based on the signaled index. In this case, the signaled bit can have a value of N, where N can be an integer greater than or equal to 1.
[0182] When the above index is explicitly signaled, a first bit of index information may be used to indicate whether the default weight is applied, and a second bit of index information may be used to indicate any of the remaining weights excluding the default weight. That is, up to three bits may be signaled to signal the index. The default weight may be defined as 16 / 64, but this is only an example, and other weights may be set as the default weight.
[0183] For example, the encoder can generate prediction blocks based on available weights and calculate a predetermined cost for each prediction block. Here, the cost can be defined as the difference (SAD) between the prediction block and the original block. An index indicating the weight corresponding to the minimum of the calculated costs can be encoded and signaled to the decoder. For example, the SADs for weights of 16 / 64 and 8 / 64 can be calculated, and the weight corresponding to the smallest SAD can be signaled.
[0184] Alternatively, weights can be implicitly derived based on the difference in values corresponding to the y-axis of the histogram. Here, the y-axis can be defined as ampl, and ampl can be understood as the magnitude of the aforementioned change or the amplitude relative to the surrounding blocks.
[0185] For example, if the ampl value of the planar mode is greater than the sum of the ampl values for the intra prediction mode(s) derived from S300, the weight for the planar mode may be assigned a value greater than the default weight (e.g., 16 / 64). Alternatively, if the ampl value of the planar mode is greater than the sum of the ampl values for the intra prediction mode(s) derived from S300, the weight for the planar mode may be assigned a value less than the default weight (e.g., 16 / 64).
[0186] If the ampl value of the planar mode is less than the sum of the ampl values for the intra prediction mode(s) derived from S300, the weight value for the planar mode may be assigned less than the default weight (e.g., 16 / 64). Alternatively, if the ampl value of the planar mode is less than the sum of the ampl values for the intra prediction mode(s) derived from S300, the weight value for the planar mode may be assigned greater than the default weight (e.g., 16 / 64).
[0187] Weights can be derived by comparing the ampl values of the intra prediction mode(s) derived from S300 with a predetermined threshold. Here, the threshold can be determined by multiplying or dividing the ampl value of the planar mode by an arbitrary value.
[0188] Weights can also be implicitly derived by comparing RD costs for surrounding template regions. A first prediction value for the surrounding template region can be generated based on an intra prediction mode derived from a histogram, such as the aforementioned HoC or HoG. A second prediction value for the surrounding template region can be generated based on the aforementioned planar mode (or block vector). Predictions for the surrounding template region can be generated by weighting the first and second prediction values based on the available weights. The difference (SAD) between the predicted value and the reconstructed value can be calculated for the surrounding template region. A SAD can be calculated for each predicted value, and weights can be derived based on the smallest SAD among the calculated SADs.
[0189] If the number of intra prediction modes derived from S300 is 1, the first prediction block can be set as the prediction block of the current block. That is, the weighted sum process with the second prediction block generated based on the planar mode or block vector-based prediction can be omitted.
[0190] The above-described prediction method can be applied equally to an image encoding device and an image decoding device.
[0191] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combinations of two or more.
[0192] Additionally, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, the embodiments may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0193] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.
Claims
1. A step of generating a first prediction block of the current block; generating a second prediction block of the current block; and An image decoding method, comprising the step of generating a prediction block of the current block based on the first prediction block and the second prediction block.
2. In paragraph 1, A method for decoding an image, wherein the first prediction block is derived based on at least one of an intra prediction mode of a surrounding block of the current block or a size of the current block.
3. In paragraph 1, The above first prediction block is generated based on a predetermined intra prediction mode, A method for decoding an image, wherein the above-described intra prediction mode is derived based on the amount of change in sample values between reference samples belonging to a surrounding area of the current block.
4. In paragraph 1, The above second prediction block is generated based on a pre-defined mode, A method for decoding an image, wherein the above-described mode is a planar mode or a DC mode.
5. In paragraph 1, A method for decoding an image, wherein the second prediction block is generated based on at least one of a plurality of candidates belonging to a candidate list based on a block vector.
6. In paragraph 5, A candidate belonging to the candidate list based on the above block vector is derived based on the block vector of the candidate block decrypted before the current block, A method for decoding an image, wherein the candidate block includes at least one of a neighboring block adjacent to the current block or a neighboring block not adjacent to the current block.
7. In paragraph 5, A method for decoding an image, wherein the second prediction block is generated by applying a predetermined linear filter to a reference block specified by at least one of the plurality of candidates.
8. In paragraph 1, The prediction block of the current block is generated through a weighted sum of the first prediction block and the second prediction block, A method for decoding an image, wherein the weight for the above weighted sum is adaptively determined as one of the weights available to the current block.
9. A step of generating a first prediction block of the current block; generating a second prediction block of the current block; and A video encoding method, comprising the step of generating a prediction block of the current block based on the first prediction block and the second prediction block.
10. A computer-readable storage medium for storing a bitstream generated by the image encoding method according to Article 9.
Citation Information
Patent Citations
Method and apparatus for processing a video signal
KR1020180005121A
Semiconductor device
KR1020250003146A
A puzzle type electric mat
KR102701191B1
Installation structure of air conditioner attached to the exterior of a building in a drawer style
KR102711720B1
KR20220077095A