Video encoding / decoding method and apparatus, and recording medium having bitstream stored therein
The method enhances video encoding efficiency by using extrapolation filters with adaptive modes to derive and signal coefficients for improved prediction blocks, addressing inefficiencies in high-resolution image compression.
Patent Information
- Application Number
- PCT/KR2025/000217
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2025-01-06
- Publication Date
- 2025-07-17
AI Technical Summary
Existing video compression technologies face challenges in efficiently compressing high-resolution images, particularly in deriving and signaling information for adaptive use of extrapolation filters in inter prediction, which affects encoding efficiency.
A method and device for generating prediction blocks using extrapolation filters (EIP) with adaptive filter type and coefficients, derived based on pre-defined modes such as general, merge, and decoder-side EIP modes, utilizing motion information and block vectors to enhance encoding efficiency.
Improves video encoding efficiency by effectively deriving and signaling EIP coefficients, allowing for better prediction block generation and adaptive filter application.
Smart Images

Figure KR2025000217_17072025_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device and recording medium storing bitstream
[0001] The present invention relates to a video signal processing method and device.
[0002] The market demand for high-resolution video is growing, necessitating technologies capable of efficiently compressing high-resolution images. To address this market need, the ISO / IEC's Moving Picture Expert Group (MPEG) and the ITU-T's Video Coding Expert Group (VCEG) jointly formed the Joint Collaborative Team on Video Coding (JCT-VC). They completed development of the HEVC (High Efficiency Video Coding) video compression standard in January 2013 and have been actively conducting research and development on next-generation compression standards.
[0003] Video compression largely consists of intra-prediction, inter-prediction, transform, quantization, entropy coding, and in-loop filtering. Among these, inter-prediction refers to a technology that generates a prediction block for the current block based on adjacent frames that have been encoded / decoded. The encoder encodes the inter-prediction mode used for inter-prediction, and the decoder decodes the encoded inter-prediction mode to perform inter-prediction.
[0004] The present disclosure provides a method and device for generating a prediction block based on EIP.
[0005] The present disclosure provides a method and device for deriving filter types and coefficients for EIP.
[0006] The present disclosure seeks to provide a method and device for signaling information for adaptive use of EIP.
[0007] The video decoding method and device according to the present disclosure can derive coefficients of a filter for a current block and generate a prediction block of the current block based on the derived coefficients. The coefficients of the filter can be derived based on a predefined mode, and the predefined mode can be a general EIP mode, a merge EIP mode, a BV EIP mode, or a decoder-side EIP mode.
[0008] In the image decoding method and device according to the present disclosure, when the pre-defined mode is the general EIP mode, the coefficients of the filter can be derived based on the reference area of the current block.
[0009] In the image decoding method and device according to the present disclosure, the reference area can be specified based on motion information of one or more surrounding blocks of the current block.
[0010] In the image decoding method and device according to the present disclosure, the reference area can be specified based on at least one of the size of the current block or the size of the filter.
[0011] In the image decoding method and device according to the present disclosure, the filter type of the filter can be determined based on type information signaled through a bitstream.
[0012] In the video decoding method and device according to the present disclosure, when the pre-defined mode is the merge EIP mode, the coefficients of the filter for the current block can be derived based on any one of a plurality of candidates belonging to a candidate list.
[0013] In the video decoding method and device according to the present disclosure, when the pre-defined mode is the BV EIP mode, the step of deriving coefficients of the filter may include the steps of generating a candidate list for the current block, determining a reference block based on a block vector (BV) of any one of a plurality of candidates belonging to the candidate list, and deriving coefficients by applying the filter to the reference block.
[0014] In the video decoding method and device according to the present disclosure, when the pre-defined mode is the decoder side EIP mode, the coefficients of the filter can be derived based on the template of the current block.
[0015] In the video decoding method and device according to the present disclosure, any one of the general EIP mode, the merge EIP mode, the BV EIP mode, or the decoder side EIP mode can be selectively used based on information signaled through a bitstream.
[0016] The video encoding method and device according to the present disclosure can derive coefficients of a filter for a current block and generate a prediction block of the current block based on the derived coefficients. The coefficients of the filter can be derived based on a predefined mode, and the predefined mode can be a general EIP mode, a merge EIP mode, a BV EIP mode, or a decoder-side EIP mode.
[0017] A computer-readable recording medium according to the present disclosure can store a bitstream encoded by the image encoding method.
[0018] According to the present disclosure, encoding efficiency can be improved by efficiently deriving coefficients of a filter for EIP.
[0019] According to the present disclosure, information for adaptively applying EIP can be effectively signaled.
[0020] FIG. 1 is a block diagram showing an image encoding device according to the present disclosure.
[0021] FIG. 2 is a block diagram showing an image decoding device according to the present disclosure.
[0022] FIG. 3 illustrates an intra prediction method based on an extrapolation filter as an embodiment according to the present disclosure.
[0023] Figure 4 illustrates an example of a filter type for EIP.
[0024] FIG. 5 is an example according to the present disclosure, showing the range of a portion of the area.
[0025] The video decoding method and device according to the present disclosure can derive coefficients of a filter for a current block and generate a prediction block of the current block based on the derived coefficients. The coefficients of the filter can be derived based on a predefined mode, and the predefined mode can be a general EIP mode, a merge EIP mode, a BV EIP mode, or a decoder-side EIP mode.
[0026] In the image decoding method and device according to the present disclosure, when the pre-defined mode is the general EIP mode, the coefficients of the filter can be derived based on the reference area of the current block.
[0027] In the image decoding method and device according to the present disclosure, the reference area can be specified based on motion information of one or more surrounding blocks of the current block.
[0028] In the image decoding method and device according to the present disclosure, the reference area can be specified based on at least one of the size of the current block or the size of the filter.
[0029] In the image decoding method and device according to the present disclosure, the filter type of the filter can be determined based on type information signaled through a bitstream.
[0030] In the video decoding method and device according to the present disclosure, when the pre-defined mode is the merge EIP mode, the coefficients of the filter for the current block can be derived based on any one of a plurality of candidates belonging to a candidate list.
[0031] In the video decoding method and device according to the present disclosure, when the pre-defined mode is the BV EIP mode, the step of deriving coefficients of the filter may include the steps of generating a candidate list for the current block, determining a reference block based on a block vector (BV) of any one of a plurality of candidates belonging to the candidate list, and deriving coefficients by applying the filter to the reference block.
[0032] In the video decoding method and device according to the present disclosure, when the pre-defined mode is the decoder side EIP mode, the coefficients of the filter can be derived based on the template of the current block.
[0033] In the video decoding method and device according to the present disclosure, any one of the general EIP mode, the merge EIP mode, the BV EIP mode, or the decoder side EIP mode can be selectively used based on information signaled through a bitstream.
[0034] The video encoding method and device according to the present disclosure can derive coefficients of a filter for a current block and generate a prediction block of the current block based on the derived coefficients. The coefficients of the filter can be derived based on a predefined mode, and the predefined mode can be a general EIP mode, a merge EIP mode, a BV EIP mode, or a decoder-side EIP mode.
[0035] A computer-readable recording medium according to the present disclosure can store a bitstream encoded by the image encoding method.
[0036] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings attached to this specification so that those skilled in the art can easily implement the present invention. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In addition, in the drawings, parts irrelevant to the description have been omitted to clearly explain the present invention, and similar parts have been designated with similar reference numerals throughout the specification.
[0037] Throughout this specification, when a part is said to be 'connected' to another part, this includes not only cases where they are directly connected, but also cases where they are electrically connected with another element in between.
[0038] Additionally, whenever a part throughout this specification is said to "include" a component, this does not mean that other components are excluded, but rather that other components may be included, unless specifically stated otherwise.
[0039] Additionally, while terms such as "first," "second," etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another.
[0040] Additionally, in the embodiments of the devices and methods described herein, some components of the devices or some steps of the methods may be omitted. Furthermore, the order of some components of the devices or some steps of the methods may be changed. Furthermore, other components or other steps may be inserted into some components of the devices or some steps of the methods.
[0041] Additionally, some components or some steps of the first embodiment of the present invention may be added to the second embodiment of the present invention, or some components or some steps of the second embodiment may be replaced.
[0042] In addition, the components shown in the embodiments of the present invention are independently depicted to represent different characteristic functions, and this does not mean that each component is composed of separate hardware or a single software component. That is, each component is described by listing each component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function. Such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0043] In this specification, a block can be variously expressed as a unit, an area, a unit, a partition, etc., and a sample can be variously expressed as a pixel, a pel, a pixel, etc.
[0044] Hereinafter, embodiments of the present invention will be described in more detail with reference to the attached drawings. In describing the present invention, duplicate descriptions of identical components will be omitted.
[0045] FIG. 1 is a block diagram showing an image encoding device according to the present disclosure.
[0046] Referring to FIG. 1, a video encoding device (100) may include a picture segmentation unit (110), a prediction unit (120, 125), a transformation unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse transformation unit (145), a filter unit (150), and a memory (155).
[0047] The picture segmentation unit (110) can segment the input picture into at least one processing unit. At this time, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). Hereinafter, in the embodiments of the present disclosure, the coding unit may be used to mean a unit that performs encoding or a unit that performs decoding.
[0048] A prediction unit may be divided into at least one square or rectangular shape of the same size within a single coding unit, or may be divided such that one prediction unit among the divided prediction units within a single coding unit has a different shape and / or size from another prediction unit. When a prediction unit that performs intra prediction based on a coding unit is generated and is not the minimum coding unit, intra prediction can be performed without being divided into a plurality of NxN prediction units.
[0049] The prediction unit (120, 125) may include an inter prediction unit (120) that performs inter prediction or inter prediction, and an intra prediction unit (125) that performs intra prediction or intra prediction. It may determine whether to use inter prediction or intra prediction for a prediction unit, and determine specific information (e.g., intra prediction mode, motion vector, reference picture, etc.) according to each prediction method. A residual value (residual block) between the generated prediction block and the original block may be input to the transformation unit (130). In addition, prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value by the entropy encoding unit (165) and transmitted to the decoder.
[0050] The inter prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous or subsequent pictures of the current picture, and in some cases, may predict a prediction unit based on information of a portion of an encoded region within the current picture. The inter prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.
[0051] The reference picture interpolation unit can receive reference picture information from the memory (155) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.
[0052] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm) can be used to derive a motion vector. The motion vector can have a motion vector value in units of 1 / 2 or 1 / 4 pixels based on the interpolated pixel. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various methods such as Skip Mode, Merge Mode, AMVP Mode, Affine Mode, and Affine Merge Mode can be used as motion prediction methods.
[0053] The intra prediction unit (125) can generate a prediction unit based on reference pixel information surrounding the current block, which is pixel information within the current picture. If the surrounding block of the current prediction unit is a block on which inter prediction has been performed and the reference pixel is a pixel on which inter prediction has been performed, the reference pixel included in the block on which inter prediction has been performed can be replaced and used with reference pixel information of the surrounding block on which intra prediction has been performed. That is, if the reference pixel is not available, the unavailable reference pixel information can be replaced and used with at least one reference pixel among the available reference pixels.
[0054] Additionally, a residual block containing residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130).
[0055] In the transformation unit (130), the residual block including the residual value information of the prediction unit generated through the original block and the prediction unit (120, 125) can be transformed using a transformation method such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on the intra prediction mode information of the prediction unit used to generate the residual block.
[0056] The quantization unit (135) can quantize values converted to the frequency domain by the transformation unit (130). The quantization coefficients may vary depending on the block or the importance of the image. The values produced by the quantization unit (135) can be provided to the dequantization unit (140) and the reordering unit (160).
[0057] The rearrangement unit (160) can perform rearrangement of coefficient values for quantized residual values.
[0058] The rearrangement unit (160) can change a two-dimensional block-shaped coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the rearrangement unit (160) can change the two-dimensional block-shaped coefficient into a one-dimensional vector form by scanning from the DC coefficient to the coefficient of the high-frequency region using a zig-zag scan method. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional block-shaped coefficient in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficient in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, it is possible to determine which scan method among the zig-zag scan, the vertical scan, and the horizontal scan is to be used.
[0059] The entropy encoding unit (165) can perform entropy encoding based on the values produced by the rearrangement unit (160). Entropy encoding can use various encoding methods such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC). In this regard, the entropy encoding unit (165) can encode residual value coefficient information of the encoding unit from the rearrangement unit (160) and the prediction units (120, 125). In addition, according to the present disclosure, it is possible to signal and transmit information indicating that motion information is derived and used on the decoder side and information on a technique used to derive motion information.
[0060] The inverse quantization unit (140) and the inverse transformation unit (145) inversely quantize the values quantized in the quantization unit (135) and inversely transform the values transformed in the transformation unit (130). The residual values generated in the inverse quantization unit (140) and the inverse transformation unit (145) can be combined with the predicted prediction units predicted through the motion estimation unit, motion compensation unit, and intra prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.
[0061] The filter unit (150) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter). The deblocking filter may remove block distortion caused by boundaries between blocks in a restored picture. The offset correction unit may correct the offset from the original image on a pixel-by-pixel basis for the image on which deblocking has been performed. In order to perform offset correction for a specific picture, a method may be used in which the pixels included in the image are divided into a certain number of regions, the regions to be offset are determined, and the offset is applied to the regions, or the offset is applied by considering edge information of each pixel. The ALF (Adaptive Loop Filtering) may be performed based on a value obtained by comparing the filtered restored image with the original image. After dividing the pixels included in the image into a predetermined group, one filter to be applied to the group is determined, and filtering may be performed differentially for each group.
[0062] The memory (155) can store a restored block or picture produced through the filter unit (150), and the stored restored block or picture can be provided to the prediction unit (120, 125) when performing inter prediction.
[0063] FIG. 2 is a block diagram showing an image decoding device according to the present disclosure.
[0064] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (210), a rearrangement unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).
[0065] When a video bitstream is input to a video encoding device, the input bitstream can be decoded in the opposite procedure to that of the video encoding device.
[0066] The entropy decoding unit (210) can perform entropy decoding in a procedure opposite to that of the entropy encoding unit of the video encoder. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied in response to the method performed in the video encoder.
[0067] The entropy decoding unit (210) can decode information related to intra prediction and inter prediction performed in the encoder.
[0068] The reordering unit (215) can perform reordering based on the method by which the bitstream entropy-decoded by the entropy decoding unit (210) is reordered by the encoding unit. The coefficients expressed in the form of a one-dimensional vector can be reordered by restoring them back to coefficients in the form of a two-dimensional block.
[0069] The inverse quantization unit (220) can perform inverse quantization based on the quantization parameters provided by the encoder and the coefficient values of the rearranged block.
[0070] The inverse transform unit (225) can perform inverse transform, i.e., inverse DCT, inverse DST, and inverse KLT, on the transforms performed by the transform unit, i.e., DCT, DST, and KLT, on the quantization result performed by the image encoder. The inverse transform can be performed based on the transmission unit determined by the image encoder. In the inverse transform unit (225) of the image decoder, a transform technique (e.g., DCT, DST, KLT) can be selectively performed according to a plurality of pieces of information, such as a prediction method, the size of the current block, and the prediction direction.
[0071] The prediction unit (230, 235) can generate a prediction block based on the prediction block generation related information provided by the entropy decoding unit (210) and the previously decoded block or picture information provided by the memory (245).
[0072] As described above, when performing intra prediction or intra prediction in the same manner as the operation in the image encoder, if the size of the prediction unit and the size of the transformation unit are the same, intra prediction for the prediction unit is performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side. However, when performing intra prediction, if the size of the prediction unit and the size of the transformation unit are different, intra prediction can be performed using reference pixels based on the transformation unit. In addition, intra prediction using NxN division only for the minimum coding unit can be used.
[0073] The prediction unit (230, 235) may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (210), prediction mode information of an intra prediction method, and motion prediction-related information of an inter prediction method, and may distinguish a prediction unit from a current encoding unit and determine whether the prediction unit performs inter prediction or intra prediction. On the other hand, if the encoder (100) does not transmit motion prediction-related information for the inter prediction, but instead transmits information indicating that motion information is to be derived and used on the decoder side and information on a technique used to derive motion information, the prediction unit determination unit determines whether the inter prediction unit (230) performs prediction based on the information transmitted from the encoder (100).
[0074] The inter prediction unit (230) can perform inter prediction on the current prediction unit based on information included in at least one picture among the previous picture or the subsequent picture of the current picture including the current prediction unit, using information required for inter prediction of the current prediction unit provided by the image encoder. In order to perform inter prediction, it can be determined based on the encoding unit whether the motion prediction method of the prediction unit included in the corresponding encoding unit is one of Skip Mode, Merge Mode, AMVP Mode, Affine Mode, and Affine Merge Mode.
[0075] The intra prediction unit (235) can generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit that has performed intra prediction, intra prediction can be performed based on intra prediction mode information of the prediction unit provided by the image encoder.
[0076] The intra prediction unit (235) may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a unit that performs filtering on the reference pixels of the current block and can determine whether to apply the filter based on the prediction mode of the current prediction unit and apply it. AIS filtering can be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the image encoder. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.
[0077] The reference pixel interpolation unit can interpolate the reference pixel to generate a reference pixel of a pixel unit less than an integer value when the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on the pixel value interpolated from the reference pixel. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.
[0078] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.
[0079] Information about whether a deblocking filter has been applied to a corresponding block or picture can be received from a video encoding device, and if a deblocking filter has been applied, information about whether a strong or weak filter has been applied. The deblocking filter of the video decoder can receive information related to the deblocking filter provided by the video encoder, and the video decoder can perform deblocking filtering on the corresponding block.
[0080] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and information on the offset value. ALF can be applied to the encoding unit based on information on whether ALF is applied and ALF coefficient information provided from the encoder. This ALF information can be provided by being included in a specific parameter set.
[0081] The memory (245) can store a restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to an output unit.
[0082] FIG. 3 illustrates an intra prediction method based on an extrapolation filter as an embodiment according to the present disclosure.
[0083] Extrapolation Filter-based Intra Prediction (EIP) may be a method of applying a predetermined filter to a reference region for the current block to derive coefficients, and generating a prediction block of the current block based on the derived coefficients. Here, the reference region is a region encoded / decoded before the current block, and may include a surrounding region of the current block.
[0084] The above filter may have a size of NxM. N and M may each be an integer greater than or equal to 1. N and M may be the same value. Alternatively, N and M may be different values. The filter may derive coefficients by shifting by a size of K pixels within the reference area. Here, K may be an integer greater than or equal to 1. The coefficients of the filter may be derived based on the least square method or the Gaussian elimination method. For example, the size of the filter for deriving coefficients may be 4x4, and the coefficients of the filter may be derived by shifting by a size of 1 pixel. The coefficients of the derived filter may be applied to the current block to generate a prediction block. Hereinafter, a method for generating a prediction block based on EIP will be described in detail.
[0085] The coefficients of the filter for the current block can be derived (S300).
[0086] The coefficients of the filter can be derived based on a predefined mode identical to that of the decoder / decoder. The predefined mode can be any one of the general EIP mode, the merge EIP mode, the BV EIP mode, or the decoder-side EIP mode.
[0087] For the general EIP mode, N filter types can be defined. Here, N can be an integer greater than or equal to 1. As an example, FIG. 4 illustrates an example of a filter type for EIP. Referring to FIG. 4, the filter can have coefficient values for 16 pixels. Here, 15 of the 16 pixels can be pixels input to the filter, and the remaining 1 pixel can be a pixel output from the filter. In FIG. 4, the white area can represent the output value of the EIP. To calculate the output value of the EIP, the pixel values of the gray area (i.e., the pixel values of the sub / decoded reference area) can be used as input. However, this is only an example, and the number of pixels input to the filter can be less than or greater than 16. The number of pixels input to the filter can also be expressed by the input length of the filter or the number of filter coefficients.
[0088] The above reference area is an area belonging to the same picture as the current block, and may be the entire area that was sub- / decoded before the current block. Alternatively, the reference area may be defined as a portion of the entire area that was sub- / decoded before the current block.
[0089] The above-mentioned partial region may be specified based on motion information of one or more neighboring blocks of the current block. Here, the motion information may include at least one of a motion vector, a reference picture index, or prediction direction information. Alternatively, the above-mentioned partial region may be specified based on a block vector of one or more neighboring blocks of the current block. Here, the block vector may indicate a reference block belonging to the same picture as the current block. This may be distinguished from a motion vector indicating a reference block belonging to a different picture from the current block.
[0090] Alternatively, the partial region may be a peripheral region adjacent to the current block. Here, the peripheral region may include at least one of a left peripheral region, an upper peripheral region, an upper-left peripheral region, a lower-left peripheral region, or an upper-right peripheral region. The size of the partial region may be specified based on at least one of the size of the current block or the size of the filter. For example, the size of the partial region may be specified based on the following mathematical expression 1.
[0091] [Mathematical Formula 1]
[0092] D = min(width, height) + Size filter - 1
[0093] In mathematical expression 1, D can represent the size of some area. Width and height represent the width and height of the current block, respectively, and Size filter can represent the size of the filter used to derive the coefficients. If the size of the filter is NxM, Size filter can be determined based on at least one of N or M. For example, Size filter can be determined as N, M, (N*M), (N+M), Min(N, M), or Max(N, M).
[0094] FIG. 5 is an example according to the present disclosure, showing the range of a portion of the area.
[0095] In Fig. 5, it is assumed that the current block has a size of 8x8 and the filter size is 4x4.
[0096] As shown in Fig. 5(a), some areas may include a left peripheral area, an upper peripheral area, an upper left peripheral area, a lower left peripheral area, and an upper right peripheral area. Since the current block has a size of 8x8 and the filter size is 4x4, the size (D) of some areas according to mathematical expression 1 can be derived as 11. That is, the width and height of the left peripheral area can be derived as 11 and 8, respectively. The width and height of the upper peripheral area can be derived as 8 and 11, respectively. The width and height of the upper left peripheral area can be derived as 11 and 11, respectively. The width and height of the lower left peripheral area can be derived as 11 and 11, respectively. Alternatively, the width and height of the lower left peripheral area can be derived as 11 and 8, respectively. The width and height of the upper right peripheral area can be derived as 11 and 11, respectively. Alternatively, the width and height of the upper right peripheral area can be derived as 8 and 11, respectively.
[0097] As shown in Fig. 5(b), some areas may include a top peripheral area, an upper left peripheral area, and an upper right peripheral area. Since the current block has a size of 8x8 and the filter size is 4x4, the size (D) of some areas according to mathematical expression 1 can be derived as 11. That is, the width and height of the top peripheral area can be derived as 8 and 11, respectively. The width and height of the upper left peripheral area can be derived as 11 and 11, respectively. The width and height of the upper right peripheral area can be derived as 11 and 11, respectively. Alternatively, the width and height of the upper right peripheral area can be derived as 8 and 11, respectively.
[0098] As shown in Fig. 5(c), some areas may include a left peripheral area, an upper left peripheral area, and a lower left peripheral area. Since the current block has a size of 8x8 and the filter has a size of 4x4, the size (D) of some areas according to mathematical expression 1 can be derived as 11. That is, the width and height of the left peripheral area can be derived as 11 and 8, respectively. The width and height of the upper left peripheral area can be derived as 11 and 11, respectively. The width and height of the lower left peripheral area can be derived as 11 and 11, respectively. Alternatively, the width and height of the lower left peripheral area can be derived as 11 and 8, respectively.
[0099] The location and / or range of some regions for deriving the coefficients of the filter are not limited to some regions illustrated in FIG. 5. For example, some regions may include only the left peripheral region, the top peripheral region, and the upper left peripheral region. Alternatively, some regions may include only the left peripheral region and the top peripheral region. Alternatively, the location and / or range of some regions may be variably determined based on the properties of the current block. Here, the properties of the current block may include at least one of the number of pixels belonging to the current block, the ratio of the width and the height of the current block, the maximum or minimum value of the width and the height of the current block, whether at least one of the width or the height of the current block is greater than a first threshold, whether at least one of the width or the height of the current block is less than a second threshold, the component type of the current block, or whether the current block borders a coding unit tree (CTU).
[0100] In the Merge EIP mode, if a neighboring block of the current block is encoded / decoded in the EIP mode, the filter of the current block can be derived based on the filter used by the neighboring block. For example, the filter type and coefficients of the filter for the current block can be set to the filter type and coefficients of the filter for the neighboring block. The neighboring block can include at least one of an upper neighboring block, a left neighboring block, a lower left neighboring block, an upper right neighboring block, or an upper left neighboring block. Alternatively, only neighboring blocks at a specific position (e.g., a left neighboring block or an upper neighboring block) can be used as neighboring blocks for the Merge EIP mode.
[0101] A candidate list can be generated for the current block. The candidate list can include multiple candidates. The multiple candidates can include at least one of a spatial candidate, a temporal candidate, or a history-based candidate. The spatial candidate can include at least one of an adjacent candidate or a non-adjacent candidate. Here, an adjacent candidate can refer to a candidate derived based on a neighboring block adjacent to the current block, and a non-adjacent candidate can refer to a candidate derived based on a block that belongs to the same picture as the current block but is not adjacent to the current block. The temporal candidate can refer to a candidate derived based on a block that belongs to a different picture from the current block. A history-based candidate can refer to a candidate derived based on a block that was previously encoded / decoded in the current block and has a filter for the EIP mode. The candidate list can be generated in the same manner by the encoder and the decoder. The encoder can encode an index that specifies one of the multiple candidates in the candidate list into the bitstream. The decoder can identify one of multiple candidates in the candidate list based on an index signaled through the bitstream. The current block can inherit the filter (i.e., filter type, coefficients) of the identified candidate.
[0102] For example, a cost may be calculated for each of multiple candidates in a candidate list. Here, the cost may be defined as the difference between a value generated by applying a filter to a template of the current block and a pre-reconstructed value of the template. The difference may be defined as the Sum of Absolute Difference (SAD). The top K candidates may be selected in ascending order of SAD among the multiple candidates. Assuming that the candidate list consists of 12 candidates, K may be less than or equal to 12. For example, K may be 6, but this is merely an example. An index specifying one of the selected K candidates may be encoded in the bitstream.
[0103] BV EIP mode may be a mode that derives coefficients by applying a filter to a reference block that is specified based on the block vector of the surrounding blocks.
[0104] A candidate list for the BV EIP mode can be generated for the current block. The candidate list can include multiple candidates. The multiple candidates can include at least one of a spatial candidate, a temporal candidate, or a history-based candidate. The spatial candidate can include at least one of an adjacent candidate or a non-adjacent candidate. Here, an adjacent candidate can refer to a candidate derived based on a neighboring block adjacent to the current block, and a non-adjacent candidate can refer to a candidate derived based on a block that belongs to the same picture as the current block but is not adjacent to the current block. The temporal candidate can refer to a candidate derived based on a block that belongs to a different picture from the current block. A history-based candidate can refer to a candidate derived based on a block that has been previously encoded / decoded and has a block vector (BV) in the current block. The candidate list can be generated in the same manner by the encoder and the decoder. The encoder can encode an index that specifies one of the multiple candidates in the candidate list into a bitstream. The decoder can identify one of multiple candidates in the candidate list based on an index signaled through the bitstream. A reference block can be determined based on the block vector of the identified candidate, and a filter can be applied to the determined reference block to derive coefficients.
[0105] For example, a cost may be calculated for each of multiple candidates in a candidate list. Here, the cost may be defined as the difference between a value generated by applying a filter to a template of the current block and a pre-reconstructed value of the template. The difference may be defined as the Sum of Absolute Difference (SAD). The top K candidates may be selected in ascending order of SAD among the multiple candidates. Assuming that the candidate list consists of 12 candidates, K may be less than or equal to 12. For example, K may be 6, but this is merely an example. An index specifying one of the selected K candidates may be encoded in the bitstream.
[0106] The decoder-side EIP mode may be a mode that calculates filter types and / or coefficients based on templates previously encoded / decoded before the current block. Filter coefficients may be derived based on a predetermined search area. Here, the search area may correspond to the aforementioned reference area. The template may have a size of N pixels, where N may be an integer greater than or equal to 1. Filter coefficients may be derived for each predefined filter type. Here, the predefined filter type may include at least one of the three filter types discussed with reference to FIG. 4.
[0107] A filter of the above-described filter type can be applied to the template of the current block to generate a predicted value of the template. The difference (e.g., SAD) between the predicted value of the template and the restored value can be calculated for each filter type. The filter type corresponding to the smallest SAD can be selected.
[0108] Below, we will examine a method of signaling information for adaptively applying EIP according to the present disclosure.
[0109] The encoder can determine whether EIP is applied to the current block. If it is determined that EIP is applied to the current block, the encoder can encode a flag (EIP flag) indicating that EIP is applied to the current block into the bitstream. The decoder can determine whether EIP is applied to the current block based on the flag signaled through the bitstream. If the flag indicates that EIP is applied to the current block, the coefficient derivation process of the aforementioned filter can be performed.
[0110] A normal EIP mode and a merge EIP mode can be defined as coefficient derivation methods in the encoder / decoder. If the encoder encodes a flag indicating that EIP is applied to the current block, mode information indicating a method for deriving filter coefficients may also be additionally encoded. If the flag indicates that EIP is applied to the current block, the decoder can obtain mode information indicating a method for deriving filter coefficients from the bitstream. For example, a first value of mode information may indicate that the normal EIP mode is used, and a second value of mode information may indicate that the merge EIP mode is used.
[0111] When the mode information is encoded as a first value, the encoder can encode type information indicating a filter type and signal it to the decoder. The decoder can determine a filter type of a filter for a current block based on the signaled type information. The filter of the determined filter type can be applied to a reference region to derive coefficients. When the mode information is encoded as a second value, the encoder can encode an index for specifying one candidate from a candidate list and signal it to the decoder. The decoder can specify one candidate from the candidate list based on the signaled index and inherit a filter from the specified candidate.
[0112] A BV EIP mode can be defined as a coefficient derivation method in the encoder / decoder. If the encoder encodes a flag indicating that EIP is applied to the current block, mode information indicating a method for deriving filter coefficients may be additionally encoded. If the flag indicates that EIP is applied to the current block, the decoder may obtain mode information indicating a method for deriving filter coefficients from a bitstream. For example, the mode information as a first value may indicate that the BV EIP mode is used, and the mode information as a second value may indicate that a mode other than the BV EIP mode (e.g., a normal EIP mode or a merge EIP mode) is used. Conversely, the mode information as a second value may indicate that the BV EIP mode is used, and the mode information as a first value may indicate that a mode other than the BV EIP mode is used.
[0113] When mode information indicating that the BV EIP mode is used is encoded, the encoder can encode and signal to the decoder at least one of an index for specifying a candidate from the candidate list or type information regarding the type of the filter. The decoder can specify a candidate from the candidate list based on the signaled index, and can specify a reference block based on a block vector of the specified candidate. A filter can be applied to the specified reference block to derive coefficients. Here, the filter may be determined based on the type information.
[0114] If mode information indicating that the general EIP mode is used is encoded, the encoder can encode type information indicating a filter type and signal it to the decoder. The decoder can determine the filter type of the filter for the current block based on the signaled type information. The filter of the determined filter type can be applied to the reference region to derive coefficients. If mode information indicating that the merge EIP mode is used is encoded, the encoder can encode an index for specifying one candidate from the candidate list and signal it to the decoder. The decoder can specify one candidate from the candidate list based on the signaled index and inherit a filter from the specified candidate.
[0115] The encoder / decoder may define the normal EIP mode, merge EIP mode, and BV EIP mode as coefficient derivation methods. If the encoder encodes a flag indicating that EIP is applied to the current block, mode information indicating a method for deriving filter coefficients may also be additionally encoded. Here, the mode information may include at least one of a first flag indicating whether to use the merge EIP mode or a second flag indicating whether to use the BV EIP mode.
[0116] For example, if the value of the first flag is 0, this may indicate that the normal EIP mode is used, and if the value of the first flag is 1, this may indicate that the normal EIP mode is not used. If the value of the first flag is 1, BV EIP mode or merge EIP mode may be used. If the value of the second flag is 1, this may indicate that BV EIP mode is used, and if the value of the second flag is 0, this may indicate that merge EIP mode is used.
[0117] The encoder can determine whether the generic EIP mode is used. If it is determined that the generic EIP mode is used, the encoder can signal by encoding a first flag as 0. If the first flag is encoded as 0, the encoder can signal by encoding type information regarding the type of the filter. Conversely, if it is determined that the generic EIP mode is not used, the encoder can signal by encoding the first flag as 1. If it is determined that the generic EIP mode is not used, the encoder can determine whether the BV EIP mode is used. If it is determined that the BV EIP mode is used, the encoder can signal by encoding a second flag as 1. If the second flag is encoded as 1, the encoder can signal by encoding at least one of an index for specifying a candidate from the candidate list or type information regarding the type of the filter. On the other hand, if it is determined that the BV EIP mode is not used (or if it is determined that the merge EIP mode is used), the encoder can signal by encoding the second flag as 0 and can signal by encoding an index for specifying one candidate from the candidate list.
[0118] The decoder can obtain a flag indicating whether EIP is applied to the current block from the bitstream. If the flag indicates that EIP is applied to the current block, the first flag can be obtained from the bitstream. If the value of the first flag is 0 (or if the normal EIP mode is used), type information can be obtained from the bitstream.
[0119] If the value of the first flag is 1 (or, if the general EIP mode is not used), the second flag can be obtained from the bitstream. If the value of the second flag is 1 (or, if the BV EIP mode is used), at least one of an index for specifying one candidate from the candidate list or type information regarding the type of the filter can be obtained from the bitstream. If the value of the second flag is 0 (or, if the merge EIP mode is used), an index for specifying one candidate from the candidate list can be obtained from the bitstream.
[0120] A decoder-side EIP mode may be defined as a coefficient derivation method in a decoder / decoder. A flag (hereinafter referred to as a third flag) for adaptively using the decoder-side EIP mode may be defined. A third flag having a value of 1 may indicate that the decoder-side EIP mode is used, and a third flag having a value of 0 may indicate that the decoder-side EIP mode is not used. The encoder may determine whether the decoder-side EIP mode is used in the current block, and may encode and signal the value of the third flag based on the determination. The decoder may determine whether the decoder-side EIP mode is used based on the value of the signaled third flag.
[0121] The above third flag may be set / decoded before the flag indicating whether EIP is applied. The above flag may be set / decoded when the third flag indicates that the decoder-side EIP mode is not used. The process after the above flag is set / decoded is the same as described above, and a duplicate description will be omitted here.
[0122] Based on the coefficients derived above, a prediction block of the current block can be generated (S310).
[0123] The above filter can be applied to the current block in a predetermined scan order, wherein the scan order can be any one of raster scan, horizontal scan, or vertical scan.
[0124] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combinations of two or more.
[0125] Additionally, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, the embodiments may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0126] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.
Claims
1. A step for deriving the coefficients of the filter for the current block; and A step of generating a prediction block of the current block based on the above derived coefficients, The coefficients of the above filter are derived based on the pre-defined mode, A method for decoding a video, wherein the above-described mode is a normal EIP mode, a merge EIP mode, a BV EIP mode, or a decoder side EIP mode.
2. In paragraph 1, An image decoding method, wherein the coefficients of the filter are derived based on the reference area of the current block when the above-described mode is the general EIP mode.
3. In paragraph 2, An image decoding method, wherein the above reference area is specified based on motion information of one or more surrounding blocks of the current block.
4. In paragraph 2, A method for decoding an image, wherein the reference area is specified based on at least one of the size of the current block or the size of the filter.
5. In paragraph 2, A method for decoding an image, wherein the filter type of the above filter is determined based on type information signaled through a bitstream.
6. In paragraph 1, A method for decoding an image, wherein when the above-described mode is the merge EIP mode, the coefficients of the filter for the current block are derived based on any one of a plurality of candidates belonging to a candidate list.
7. In paragraph 1, If the above-definition mode is the BV EIP mode, the step of deriving the coefficients of the filter is, A step of generating a candidate list for the current block; A step of determining a reference block based on a block vector (BV) of any one of a plurality of candidates belonging to the above candidate list; and An image decoding method, comprising a step of deriving coefficients by applying the filter to the reference block.
8. In paragraph 1, A method for decoding an image, wherein the coefficients of the filter are derived based on the template of the current block, when the above-described mode is the decoder side EIP mode.
9. In paragraph 1, A video decoding method, wherein any one of the general EIP mode, the merge EIP mode, the BV EIP mode, or the decoder side EIP mode is selectively used based on information signaled through a bitstream.
10. A step for deriving the coefficients of the filter for the current block; and A step of generating a prediction block of the current block based on the above derived coefficients, The coefficients of the above filter are derived based on the pre-defined mode, A method of video encoding, wherein the above-described mode is a normal EIP mode, a merge EIP mode, a BV EIP mode, or a decoder side EIP mode.
11. A computer-readable recording medium storing a bitstream generated based on the image encoding method of Article 10.
Citation Information
Patent Citations
Image encoding / image decoding method and image encoding / image decoding apparatus
KR100977101B1
Competition-Based Intra Prediction Coding / Decoding Apparatus and Method Using Multiple Prediction Filters
KR101663762B1
Features of intra block copy prediction mode for video and image coding and decoding
KR1020160072181A
Microspeaker
KR1020240161880A