Video encoding / decoding method and device using adaptive interpolation filter based on sum of absolute value differences of reference samples

Adaptive interpolation filters in video codecs select between 4-tap and 8-tap filters based on block size and MeanSAD to enhance encoding/decoding efficiency for high-resolution video signals, addressing inefficiencies in existing codecs.

WO2025178347A1PCT designated stage Publication Date: 2025-08-28IND ACAD COOP GRP OF SEJONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/002349
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2025-02-18
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing video codecs face inefficiencies in encoding and decoding high-resolution video signals, particularly in block-based codecs like VVC, which require improved interpolation filters for better compression and decoding efficiency.

Method used

Adaptive interpolation filters are employed based on the sum of absolute value differences of reference samples, using 4-tap and 8-tap filters, with selection criteria including block size and mean value of absolute differences (MeanSAD), to enhance encoding/decoding efficiency.

Benefits of technology

The adaptive interpolation filters improve encoding/decoding efficiency by optimizing filter selection for different block sizes and resolutions, leading to better compression and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025002349_28082025_PF_FP_ABST
    Figure KR2025002349_28082025_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding / encoding method according to the present invention comprises the steps of: determining an intra-prediction mode used for intra-prediction of the current block; interpolating reference pixels used for the intra-prediction; and performing the intra-prediction of the current block on the basis of the interpolated reference pixels and the intra-prediction mode, wherein the types of an interpolation filter applied to the interpolation of the reference pixels include a 4-tap filter and an 8-tap filter, and the interpolation filter can be determined on the basis of the size of the current block and the mean value of absolute values of differences (MeanSAD) of the reference pixels of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and device using an adaptive interpolation filter based on the sum of absolute value differences of reference samples

[0001] The present invention relates to a method and device for encoding / decoding an image, and more particularly, to a method and device for encoding / decoding an image that adaptively selects an interpolation filter based on the sum of the absolute value differences of reference samples.

[0002] As video resolution increases, the need for more efficient and compressed video codecs also increases. The ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) formed the Joint Video Exploration Team (JVET) in October 2015 to develop a next-generation video coding standard, and the VVC / H.266 standardization was completed in July 2020. VVC is a video codec developed after High Efficiency Video Coding (HEVC / H.265) and offers a 39% bitrate reduction compared to HEVC. Like Advanced Video Coding (AVC / H.264) and HEVC, VVC is a block-based video codec. It was developed with the goal of being a codec suitable for various video formats, such as high-resolution, screen content, and 360° video.

[0003] The present invention aims to improve the encoding / decoding efficiency of a video signal.

[0004] The video decoding / encoding method of the present invention comprises the steps of: determining an intra-screen prediction mode used for intra-screen prediction of a current block; interpolating a reference pixel used for the intra-screen prediction; and performing intra-screen prediction of the current block based on the interpolated reference pixel and the intra-screen prediction mode, wherein the type of interpolation filter applied to the interpolation of the reference pixel includes a 4-tap filter and an 8-tap filter, and the interpolation filter can be determined based on the size of the current block and the mean value (MeanSAD) of absolute values ​​of difference values ​​of reference pixels of the current block.

[0005] In the image decoding / encoding method of the present invention, the 8-tap filter may include an 8-tap DCT-IF (Discrete Cosine Transform based Interpolation Filter) and an 8-tap SIF (Smoothing Interpolation Filter).

[0006] In the image decoding / encoding method of the present invention, the interpolation filter may be determined by further considering the result of comparing the resolution of the current picture including the current block with 4K resolution (3840x2160).

[0007] In the video decoding / encoding method of the present invention, when the resolution of the current picture is less than the 4K resolution, the interpolation filter is determined based on the size of the current block and the MeanSAD, and when the resolution of the current picture is greater than the 4K resolution, the interpolation filter can be determined based on the size of the current block regardless of the MeanSAD.

[0008] In the video decoding / encoding method of the present invention, when the resolution of the current picture is less than the 4K resolution, when the size of the current block is 2 and the MeanSAD is less than a threshold value, the interpolation filter may be determined as a 4-tap SIF, and when the size of the current block is 2 and the MeanSAD is greater than or equal to a threshold value, the interpolation filter may be determined as an 8-tap DCT-IF.

[0009] In the video decoding / encoding method of the present invention, when the resolution of the current picture is less than the 4K resolution, when the size of the current block is 5 or more and the MeanSAD is less than a threshold value, the interpolation filter may be determined as an 8-tap SIF, and when the size of the current block is 5 or more and the MeanSAD is greater than or equal to a threshold value, the interpolation filter may be determined as a 4-tap DCT-IF.

[0010] A method and device for encoding / decoding using an adaptive interpolation filter based on the sum of absolute value differences of reference samples according to the present invention can improve the encoding / decoding efficiency of a video signal.

[0011] FIG. 1 is a block diagram showing an image encoding device according to an embodiment of the present invention.

[0012] FIG. 2 is a block diagram showing an image decoding device (200) according to one embodiment of the present invention.

[0013] Figure 3 is a diagram showing h[n], y[n], and z[n] for obtaining 8-tap coefficients.

[0014] Figure 4 illustrates the integer reference samples used to derive the 8-tap SIF coefficients.

[0015] Figure 5 is a diagram showing the average correlation values ​​of reference samples for various video resolutions and each nTbS.

[0016] Figure 6 shows the magnitude response at 16 / 32 pixel locations for 4-tap DCT-IF, 4-tap SIF, 8-tap DCT-IF, and 8-tap SIF.

[0017] Figure 7 illustrates embodiments for 8-tap DCT interpolation filter coefficients.

[0018] Figure 8 illustrates an embodiment of 8-tap smoothing interpolation filter coefficients.

[0019] Figure 9 illustrates an example of a method for selecting an interpolation filter using SAD-based information when the resolution is less than 4K.

[0020] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.

[0021] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.

[0022] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0023] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0024] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. Hereinafter, identical components in the drawings will be designated by the same reference numerals, and redundant descriptions of identical components will be omitted.

[0025] The present invention proposes a method for obtaining post-transform data and a method for rearranging and scanning transform coefficients according to the signal-dependent KL (Karhunen-Loeve) or SVD (Singular Value Decomposition) transform using the covariance and correlation of each two-dimensional block (wherein the block means a residual signal block or a transformed block) when using a separable horizontal and vertical one-dimensional transform such as a signal-independent DCT-2 (Discrete Cosine Transform-2), DCT-8 (Discrete Cosine Transform-8), or DST-7 (Discrete Sine Transform-7) used in a video compression / reconstruction standard, unlike the case where the signal-independent DCT-2, DCT-8 (Discrete Cosine Transform-8), or DST-7 (Discrete Sine Transform-7) transform is used.

[0026]

[0027] FIG. 1 is a block diagram showing an image encoding device according to one embodiment of the present invention.

[0028] Referring to FIG. 1, an image encoding device (100) may include an image segmentation unit (101), an intra-screen prediction unit (102), an inter-screen prediction unit (103), a subtraction unit (104), a transformation unit (105), a quantization unit (106), an entropy encoding unit (107), an inverse quantization unit (108), an inverse transformation unit (109), an increase unit (110), a filter unit (111), and a memory (112).

[0029] Each component shown in Fig. 1 is independently depicted to indicate different characteristic functions in the video encoding device, and does not mean that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or one component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0030] Additionally, some components may not be essential components that perform essential functions of the present invention, but may be optional components merely used to enhance performance. The present invention may be implemented by including only components essential to implementing the essence of the present invention, excluding components used solely for performance enhancement. A structure that includes only essential components, excluding optional components used solely for performance enhancement, is also within the scope of the present invention.

[0031] The image segmentation unit (100) can segment an input image into at least one block. At this time, the input image can have various shapes and sizes, such as pictures, slices, tiles, and segments. The block can mean a coding unit (CU), a prediction unit (PU), or a transformation unit (TU). The segmentation can be performed based on at least one of a quadtree or a binary tree. A quadtree is a method of dividing an upper block into four sub-blocks whose width and height are half of those of the upper block. A binary tree is a method of dividing an upper block into two sub-blocks whose width or height is half of that of the upper block. Through the segmentation based on the binary tree described above, a block can have a shape that is not only square but also non-square.

[0032] The prediction unit (102, 103) may include an inter-prediction unit (103) that performs inter-prediction and an intra-prediction unit (102) that performs intra-prediction. It may be determined whether to use inter-prediction or intra-prediction for a prediction unit, and specific information (e.g., intra-prediction mode, motion vector, reference picture, etc.) according to each prediction method may be determined. At this time, the processing unit where the prediction is performed and the processing unit where the prediction method and specific contents are determined may be different. For example, the prediction method and prediction mode may be determined for a prediction unit, and the prediction may be performed for a transformation unit.

[0033] The residual value (residual block) between the generated prediction block and the original block can be input to the transformation unit (105). In addition, the prediction mode information, motion vector information, etc. used for prediction can be encoded together with the residual value by the entropy encoding unit (107) and transmitted to the decoder. When using a specific encoding mode, it is also possible to encode the original block as is and transmit it to the decoder without generating the prediction block through the prediction unit (102, 103).

[0034] The prediction unit (102) within the screen can generate a prediction block based on reference pixel information surrounding the current block, which is pixel information within the current picture. If the prediction mode of a block surrounding the current block on which intra prediction is to be performed is inter prediction, the reference pixel included in the surrounding block to which inter prediction is applied can be replaced with a reference pixel within another block surrounding which intra prediction is applied. That is, if the reference pixel is not available, the unavailable reference pixel information can be used by replacing it with at least one reference pixel among the available reference pixels.

[0035] In intra prediction, the prediction mode can have a directional prediction mode that uses reference pixel information according to the prediction direction, and a non-directional mode that does not use directional information when performing prediction. The mode for predicting luminance information and the mode for predicting chrominance information can be different, and the intra prediction mode information used to predict luminance information or the predicted luminance signal information can be utilized to predict chrominance information.

[0036] The prediction unit (102) within the screen may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a filter that performs filtering on the reference pixels of the current block and can adaptively determine whether to apply the filter depending on the prediction mode of the current prediction unit. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.

[0037] The reference pixel interpolation unit of the prediction unit (102) within the screen can interpolate the reference pixel to generate a reference pixel at a fractional unit position when the intra prediction mode of the prediction unit is a prediction unit that performs intra prediction based on the pixel value interpolated from the reference pixel. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.

[0038] The inter-screen prediction unit (103) generates a prediction block using the previously restored reference image and motion information stored in the memory (112). The motion information may include, for example, a motion vector, a reference picture index, a list 1 prediction flag, a list 0 prediction flag, etc.

[0039] A residual block containing residual value information, which is the difference value between the prediction unit generated in the prediction unit (102, 103) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130) and transformed.

[0040] The inter-screen prediction unit (103) can derive a prediction block based on information about at least one picture from among the previous or subsequent pictures of the current picture. Furthermore, the prediction block of the current block can also be derived based on information about a portion of the current picture in which encoding has been completed. The inter-screen prediction unit (103) according to one embodiment of the present invention can include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.

[0041] The reference picture interpolation unit can receive reference picture information from the memory (112) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.

[0042] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods can be used to derive a motion vector, such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm). The motion vector can have a motion vector value in units of 1 / 2 or 1 / 4 pixels based on the interpolated pixel. The motion prediction unit can predict the prediction block of the current block by using different motion prediction methods. Various methods can be used as motion prediction methods, such as the Skip method, the Merge method, and the AMVP (Advanced Motion Vector Prediction) method.

[0043] The subtraction unit (104) subtracts the block to be currently encoded from the prediction block generated from the intra-screen prediction unit (102) or inter-screen prediction unit (103) to generate a residual block of the current block.

[0044] In the transformation unit (105), the residual block including the residual data can be transformed using a transformation method such as DCT, DST, KLT (Karhunen Loeve Transform, KL), SVD, etc. At this time, the transformation method (or transformation kernel) can be determined based on the intra prediction mode of the prediction unit used to generate the residual block. For example, depending on the intra prediction mode, DCT may be used in the horizontal direction and DST may be used in the vertical direction.

[0045] The quantization unit (106) can quantize the values ​​converted to the frequency domain by the transformation unit (105). The quantization coefficients can vary depending on the block or the importance of the image. The values ​​produced by the quantization unit (106) can be provided to the inverse quantization unit (108) and the entropy encoding unit (107).

[0046] The above-described transformation unit (105) and / or quantization unit (106) may be optionally included in the image encoding device (100). That is, the image encoding device (100) may encode the residual block by performing at least one of transformation or quantization on the residual data of the residual block, or by skipping both transformation and quantization. Even if neither transformation nor quantization is performed in the image encoding device (100), or neither transformation nor quantization is performed, a block that enters the input of the entropy encoding unit (107) is typically referred to as a transformation block. The entropy encoding unit (107) entropy-encodes the input data. Entropy encoding may use various encoding methods, such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).

[0047] The entropy encoding unit (107) can encode various information such as coefficient information of a transform block, block type information, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. The coefficients of a transform block can be encoded in units of sub-blocks within the transform block.

[0048] For encoding the coefficients of a transform block, various syntax elements can be encoded, such as Last_sig, a syntax element indicating the position of the first non-zero coefficient according to the scan order, Coded_sub_blk_flag, a flag indicating whether there is at least one non-zero coefficient in the subblock, Sig_coeff_flag, a flag indicating whether the coefficient is non-zero, Abs_greaterN_flag, a flag indicating whether the absolute value of the coefficient is greater than N (where N can be a natural number such as 1, 2, 3, 4, or 5), and Sign_flag, a flag indicating the sign of the coefficient. The residual value of the coefficient that is not encoded by the above syntax elements alone can be encoded through the syntax element remaining_coeff.

[0049] The inverse quantization unit (108) and the inverse transformation unit (109) inversely quantize the values ​​quantized in the quantization unit (106) and inversely transform the values ​​transformed in the transformation unit (105). The residual values ​​generated in the inverse quantization unit (108) and the inverse transformation unit (109) can be combined with the prediction units predicted through the motion estimation unit, motion compensation unit, and the intra-screen prediction unit (102) included in the prediction unit (102, 103) to generate a reconstructed block. The multiplier (110) multiplies the prediction blocks generated in the prediction units (102, 103) and the residual blocks generated through the inverse transformation unit (109) to generate a reconstructed block.

[0050] The filter unit (111) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter).

[0051] A deblocking filter can remove block distortion caused by boundaries between blocks in a reconstructed picture. To determine whether to perform deblocking, a deblocking filter can be applied to the current block based on the pixels contained in several columns or rows within the block. When applying a deblocking filter to a block, a strong filter or a weak filter can be applied depending on the required deblocking filtering strength. Furthermore, when applying a deblocking filter, horizontal and vertical filtering can be processed in parallel when performing vertical and horizontal filtering.

[0052] The offset correction unit can correct the offset from the original image on a pixel-by-pixel basis for an image that has undergone deblocking. To perform offset correction for a specific picture, the pixels contained in the image can be divided into a certain number of regions, the regions to be offset can be determined, and the offset can be applied to those regions. Alternatively, the offset can be applied by considering the edge information of each pixel.

[0053] Adaptive Loop Filtering (ALF) can be performed based on the comparison of the filtered restored image with the original image. After dividing the pixels included in the image into predetermined groups, a filter to be applied to each group can be determined, and filtering can be performed differentially for each group. Information regarding whether to apply ALF can be transmitted by luminance signal for each coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied can vary depending on each block. Furthermore, an ALF filter of the same form (fixed form) can be applied regardless of the characteristics of the target block.

[0054] The memory (112) can store a restoration block or picture produced through the filter unit (111), and the stored restoration block or picture can be provided to the prediction unit (102, 103) when performing inter-screen prediction.

[0055]

[0056] Next, an image decoding device according to an embodiment of the present invention will be described with reference to the drawings. Fig. 2 is a block diagram illustrating an image decoding device (200) according to an embodiment of the present invention.

[0057] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (201), an inverse quantization unit (202), an inverse transformation unit (203), an amplification unit (204), a filter unit (205), a memory (206), and a prediction unit (207, 208).

[0058] When an image bitstream generated by an image encoding device (100) is input to an image decoding device (200), the input bitstream can be decoded according to a process opposite to the process performed in the image encoding device (100).

[0059] The entropy decoding unit (201) can perform entropy decoding in a procedure opposite to that of the entropy encoding unit (107) of the video encoding device (100). For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied in response to the method performed in the video encoder. The entropy decoding unit (201) can decode the syntax elements described above, namely, Last_sig, Coded_sub_blk_flag, Sig_coeff_flag, Abs_greaterN_flag, Sign_flag, and remaining_coeff. In addition, the entropy decoding unit (201) can decode information related to intra prediction and inter prediction performed in the video encoding device (100).

[0060] The inverse quantization unit (202) performs inverse quantization on a quantized transform block to generate a transform block. It operates substantially the same as the inverse quantization unit (108) of Fig. 1.

[0061] The inverse transform unit (203) performs an inverse transform on the transform block to generate a residual block. At this time, the transform method can be determined based on information regarding the prediction method (inter or intra prediction), the size and / or shape of the block, the intra prediction mode, etc. It operates substantially the same as the inverse transform unit (109) of FIG. 1.

[0062] The multiplication unit (204) multiplies the prediction block generated from the intra-screen prediction unit (207) or inter-screen prediction unit (208) and the residual block generated through the inverse transformation unit (203) to generate a restored block. It operates substantially the same as the multiplication unit (110) of Fig. 1.

[0063] The filter unit (205) reduces various types of noise occurring in restored blocks.

[0064] The filter unit (205) may include a deblocking filter, an offset correction unit, and an ALF.

[0065] Information on whether a deblocking filter has been applied to a corresponding block or picture from a video encoding device (100) and, if a deblocking filter has been applied, information on whether a strong filter or a weak filter has been applied can be provided. The deblocking filter of the video decoding device (200) can receive information related to the deblocking filter provided from the video encoding device (100) and perform deblocking filtering on the corresponding block in the video decoding device (200).

[0066] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and offset value information.

[0067] ALF can be applied to an encoding unit based on information on whether ALF is applied, ALF coefficient information, etc. provided from a video encoding device (100). This ALF information can be provided by being included in a specific parameter set. The filter unit (205) operates substantially the same as the filter unit (111) of FIG. 1.

[0068] The memory (206) stores the recovery block generated by the increase unit (204). It operates substantially the same as the memory (112) of FIG. 1.

[0069] The prediction unit (207, 208) can generate a prediction block based on the prediction block generation related information provided by the entropy decoding unit (201) and the previously decoded block or picture information provided by the memory (206).

[0070] The prediction unit (207, 208) may include an intra-screen prediction unit (207) and an inter-screen prediction unit (208). Although not separately illustrated, the prediction unit (207, 208) may further include a prediction unit determination unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (201), prediction mode information of an intra-prediction method, and motion prediction-related information of an inter-prediction method, and may distinguish a prediction unit from a current encoding unit and determine whether the prediction unit performs inter-prediction or intra-prediction. The inter-screen prediction unit (208) may perform inter-screen prediction on the current prediction unit based on information included in at least one of a previous picture or a subsequent picture of the current picture including the current prediction unit, using information necessary for inter-prediction of the current prediction unit provided from the video encoding device (100). Alternatively, inter-screen prediction can be performed based on information from some previously reconstructed region within the current picture containing the current prediction unit.

[0071] In order to perform inter-screen prediction, it is possible to determine whether the motion prediction method of the prediction unit included in the encoding unit is Skip Mode, Merge Mode, or AMVP Mode based on the encoding unit.

[0072] The prediction unit (207) within the screen generates a prediction block using pixels located around the block to be encoded and previously restored.

[0073] The prediction unit (207) within the screen may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a filter that performs filtering on the reference pixels of the current block and can adaptively determine whether to apply the filter depending on the prediction mode of the current prediction unit. AIS filtering may be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the image encoding device (100). If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.

[0074] The reference pixel interpolation unit of the prediction unit (207) within the screen can interpolate the reference pixel to generate a reference pixel at a fractional unit position when the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on the pixel value interpolated from the reference pixel. The generated reference pixel at the fractional unit position can be used as a prediction pixel of a pixel within the current block. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.

[0075] The on-screen prediction unit (207) operates substantially the same as the on-screen prediction unit (102) of FIG. 1.

[0076] The inter-screen prediction unit (208) generates an inter-screen prediction block using reference pictures and motion information stored in the memory (206). The inter-screen prediction unit (208) operates substantially the same as the inter-screen prediction unit (103) of FIG. 1.

[0077]

[0078] In the present invention, 8-tap DCT-IF and 8-tap SIF, which utilize more reference samples, can be applied instead of the 4-tap Discrete Cosine Transform-based interpolation filter (DCT-IF) and 4-tap Smoothing interpolation filter (SIF) that were previously used in VVC intra-screen prediction. In some cases, a 10-tap, 12-tap, 14-tap, 16-tap filter, etc., can be used instead of the 8-tap filter. For example, the type of interpolation filter applied to the block can be adaptively selected based on at least one of the block size, resolution, intra-screen prediction mode, SAD (Sum of Absolute Difference) of reference samples, correlation of reference samples, or frequency characteristics of reference samples.

[0079] The 8-tap DCT-IF can be obtained using the mathematical formula 1 below.

[0080] [Mathematical Formula 1]

[0081]

[0082]

[0083] A p / 32 pixel interpolation filter (when p=0,1,2,3,…,31, 1 / 32 fractional samples are used) can be obtained by replacing n=3+p / 32 in Equation 1 above with a linear combination of discrete cosine coefficients and x(m) (m=0,1,2,…,7).

[0084] For example, the 8-tap DCT-IF coefficients derived for fractional sample positions (0 / 32, 1 / 32, 2 / 32, …, 16 / 32) can be obtained using n=3+(0 / 32, 1 / 32, 2 / 32, …, 16 / 32). The 8-tap DCT-IF coefficients for (17 / 32, 18 / 32, 19 / 32, …, 31 / 32) can also be obtained in a similar way.

[0085] Figure 3 is a diagram showing h[n], y[n], and z[n] for obtaining 8-tap coefficients.

[0086] The 8-tap SIF coefficients can be obtained from the convolution of z[n] and a 1 / 32 fractional linear filter. Here, z[n] in FIG. 3 can be obtained from the convolution of h[n] and y[n] in Equations 2 and 3. Here, h[n] can be a 3-point [1, 2, 1] LPF (Low Pass filter).

[0087] Equations 2 and 3 illustrate the process of deriving y[n] and z[n]. Figure 3 shows h[n], y[n], and z[n], and the 8-tap SIF coefficients can be obtained through linear interpolation of z[n] and a 1 / 32 fractional linear filter.

[0088] [Equation 2]

[0089]

[0090] [Equation 3]

[0091]

[0092] The following equations 4 and 5 describe the procedure for calculating the 8-tap SIF coefficients, where g[n] = z[n-3], n = 0, 1, 2, …, 6.

[0093] [Equation 4]

[0094]

[0095] Mathematical expression 5 is when p=16 An example of deriving SIF coefficients from location is shown.

[0096] [Equation 5]

[0097]

[0098] Figure 4 illustrates the integer reference samples used to derive the 8-tap SIF coefficients. Referring to Figure 4, the black 8 integer samples to derive filter coefficients for location is displayed in gray. Also, can be the starting sample of the eight reference samples. The filter coefficients can be adjusted with integer implementation.

[0099]

[0100] In one embodiment of the present invention, the adaptive interpolation filter of the present invention determines the characteristics of a block using the size of the block and the absolute value of the difference between samples that are one sample apart from a reference sample, and the type of interpolation filter used for the block can be selected. In addition, the characteristics of the block can be determined based on the sum of the absolute values ​​of the reference samples.

[0101] Specifically, the mean value of SAD (MeanSAD) is obtained based on the absolute value of the difference between samples that are 1 sample apart from the reference sample, and the type of interpolation filter (4-tap DCT-IF, 4-tap SIF, 8-tap DCT-IF, 8-tap SIF, etc.) can be determined considering the obtained MeanSAD and the size of the block.

[0102] At this time, the characteristic that the smaller the block (CU) size, the larger the average value of SAD (MeanSAD), and the larger the block (CU) size, the smaller the MeanSAD can be utilized.

[0103] The location of the reference sample used in MeanSAD for the reference sample of the block can be any one of a pre-defined location, a location directly specified from the bitstream, and a location indirectly determined from the prediction mode within the picture.

[0104] The pre-defined position can be at least one of the left block, the bottom left block, the top left block, the top block, the top right block, the right block, the bottom right block, and the bottom block of the current block. For example, the pre-defined position can be the left block or the top block of the current block. For example, the pre-defined position can be the left block and the bottom left block of the current block. For example, the pre-defined position can be the top block and the top right block of the current block. For example, the pre-defined position can be the left block and the top block of the current block.

[0105] A position directly specified from the bitstream is a position specified by position information signaled from the bitstream, and may be at least one of the left block, the lower left block, the upper left block, the upper block, the upper right block, the right block, the lower right block, and the lower block of the current block.

[0106] The position indirectly determined from the intra-screen prediction mode is a position determined according to the directionality of the intra-screen prediction mode, and may be at least one of the left block, the lower left block, the upper left block, the upper block, the upper right block, the right block, the lower right block, and the lower block of the current block. For example, when the intra-screen prediction mode is a vertical prediction mode, the upper reference sample Ref[x, -1], x=0,1,…,N-1. in which N: the width of the current block (CU) of the current block may be used. For example, when the intra-screen prediction mode is a horizontal prediction mode, the left reference sample Ref[-1, y], y=0,1,…,N-1. in which N: the height of the current block (CU) may be used. In the above example, the first reference sample line is taken as an example, but the second or third reference sample line may be used, which may be implicitly determined by considering the size of the block or may be determined by information signaled from the bitstream. For example, if the product of the width and height of a block is 32 or less, the first reference sample line may be used, and if it exceeds 32, the second or third reference sample line may be used. This is one embodiment and does not mean that the embodiments of the present invention are limited to the above embodiments.

[0107] MeanSAD can be obtained by the following mathematical expression 6 when the first upper reference sample line is used.

[0108] [Equation 6]

[0109]

[0110] Here, N can be the width of the current block. x can have values ​​of 0, 1,…, N-1.

[0111] MeanSAD can be obtained by the following mathematical expression 7 when the first left reference sample line is used.

[0112] [Equation 7]

[0113]

[0114] Here, N can be the height of the current block. y can have values ​​of 0, 1,…, N-1.

[0115] The adaptive interpolation filter of the present invention can be determined by considering the frequency characteristics of the block determined based on MeanSAD and the size of the block.

[0116] For blocks with large MeanSAD, an 8-tap DCT-IF with strong high frequency restoration characteristics can be applied, and for blocks with small MeanSAD, an 8-tap SIF, which is a strong low pass filter (LPF), can be applied.

[0117] By utilizing the characteristic that the smaller the block size, the larger the Means AD (the characteristic that there are many high frequency components), when the block size is less than or equal to a certain value and the high frequency ratio is greater than or equal to a set threshold (MeanSAD≥THR), the 8-tap DCT-IF with high frequency restoration characteristics can be applied. On the other hand, when the high frequency is less than a set threshold (MeanSAD <THR), 약한 LPF인 4-tap SIF를 사용할 수 있다.

[0118] By taking advantage of the characteristic that the larger the block size, the smaller the MeanSAD (the characteristic that there are many low frequencies), if the block size is larger than a certain value and the high frequency is smaller than a set threshold (MeanSAD < THR), a strong LPF, 8-tap SIF, can be applied. However, if the high frequency is larger than a set threshold (MeanSAD ≥ THR), a weak HPF, 4-tap DCT-IF, can be used.

[0119] Here, the block size can be the width or height of the block. Additionally, the reference value for the block size can be any of 2, 3, 4, 5, or 6.

[0120]

[0121] The adaptive interpolation filter of the present invention can be determined by considering the resolution of the picture, the frequency characteristics of the block determined based on MeanSAD, and the size of the block.

[0122] Figure 5 is a diagram showing the average correlation values ​​of reference samples for various video resolutions and each nTbS.

[0123] Specifically, Fig. 5 shows the average correlation values ​​of reference samples for various video resolutions and each nTbS defined in Equation 8, which can be determined based on the size of the CU at each screen resolution. As shown in Fig. 5, the correlation may increase as the CU size increases and the video resolution increases. Here, the video resolutions A1, A2, B, C, and D may be indicated in parentheses.

[0124] [Equation 8]

[0125] nTbs = (( Log 2(W) + Log 2 (H)) >> 1

[0126] Intra-CU size partitioning in video coding can rely on prediction performance to improve coding in terms of bitrate and distortion. Prediction performance can vary depending on the prediction error between predicted samples and samples in the current CU. If the current block contains a large number of high-frequency regions and a large number of detailed regions, the CU size can be split into smaller CU sizes, considering bitrate and distortion, using boundary reference samples with smaller widths and heights. However, if the current block consists of homogeneous regions, the CU size can be split into larger CU sizes, considering bitrate and distortion, using boundary reference samples with larger widths and heights.

[0127] Referring to Figure 5, the correlation values ​​of reference samples according to the nTbS size and video resolution, represented by A1, A2, B, C, and D, can be seen. This may mean that small nTbS have high-frequency characteristics consistent with low correlation, and large nTbS have low-frequency characteristics consistent with high correlation.

[0128] Figure 6 shows the magnitude response at 16 / 32 pixel locations for 4-tap DCT-IF, 4-tap SIF, 8-tap DCT-IF, and 8-tap SIF.

[0129] The X-axis of Figure 6 represents normalized radian frequency, and the Y-axis can represent magnitude response.

[0130] 8-tap DCT-IF may have better HPF characteristics than 4-tap DCT-IF, and 8-tap SIF may have better LPF characteristics than 4-tap SIF. Therefore, 8-tap SIF may provide better interpolation than 4-tap SIF for low-frequency reference samples, and 8-tap DCT-IF may provide better interpolation than 4-tap DCT-IF for high-frequency reference samples.

[0131]

[0132] Figure 7 illustrates embodiments for 8-tap DCT interpolation filter coefficients.

[0133] Figure 8 illustrates an embodiment of 8-tap smoothing interpolation filter coefficients.

[0134] In Figs. 7 and 8, index=0~7 can represent 8 integer samples r[i_0]~r[i_0+7] for deriving filter coefficients.

[0135]

[0136] Figure 9 illustrates an example of a method for selecting an interpolation filter using SAD-based information when the resolution is less than 4K.

[0137] The adaptive interpolation filter of the present invention can be determined by considering the resolution of the picture, the frequency characteristics of the block determined based on MeanSAD, and the size of the block.

[0138] [Example 1]

[0139] When the resolution is less than 4K, if the MeanSAD of a block with nTbS 2 (small block) is less than a given threshold (THR), 4-tap SIF is applied, and if the MeanSAD is greater than or equal to the given threshold (THR), 8-tap DCT-IF can be applied (long-tap DCT-IF can perform better than 8-tab DCT-IF). Also, if the MeanSAD of a block with nTbS 5 or more (large block) is less than a given threshold (THR), 8-tap SIF is applied (long-tap SIF can perform better than 8-tab SIF), and if the MeanSAD is greater than or equal to the given threshold (THR), 4-tap DCT-IF can be applied.

[0140] In one embodiment, when the resolution is 4K or higher (video resolution ≥ 3840x2160), regardless of the value of MeanSAD, 8-tap SIF can be applied when the block size is large, and 4-tap DCT-IF can be applied when the block size is small (long-tap SIF can perform better than 8-tab SIF).

[0141] That is, when the resolution is less than 4K, MeanSAD is considered, but when the resolution is 4K or higher, MeanSAD may not be considered.

[0142] [Example 2]

[0143] In contrast, the second embodiment can consider MeanSAD even when the resolution is 4K or higher. When the resolution is less than 4K, the interpolation filter can be applied in the same manner as the first embodiment.

[0144] When the resolution is 4K or higher, if the MeanSAD of a block with nTbS 2 (small block) is less than a given threshold (THR), 4-tap SIF is applied, and if the MeanSAD is greater than or equal to the given threshold (THR), 4-tap DCT-IF can be applied. Also, if the MeanSAD of a block with nTbS 5 or more (large block) is less than a given threshold (THR), 8-tap SIF is applied, and if the MeanSAD is greater than or equal to the given threshold (THR), 4-tap DCT-IF can be applied.

[0145] [Example 3]

[0146] If the resolution is less than 4K, the interpolation filter can be applied in the same manner as in the first embodiment.

[0147] For resolutions higher than 4K, 4-tap SIF can be applied if the MeanSAD of a block with nTbS 2 (small block) is less than a given threshold (THR), and 4-tap DCT-IF can be applied if the MeanSAD is greater than or equal to the given threshold (THR). Additionally, 8-tap SIF can be applied regardless of the MeanSAD of a block with nTbS 5 or more (large block).

[0148] [Example 4]

[0149] If the resolution is less than 4K, the interpolation filter can be applied in the same manner as in the first embodiment.

[0150] For resolutions higher than 4K, 4-tap SIF can be applied regardless of MeanSAD for blocks with nTbS of 2 (small blocks). Also, for blocks with nTbS of 5 or more (large blocks), 8-tap SIF can be applied if MeanSAD is smaller than a given threshold (THR), and 4-tap DCT-IF can be applied if MeanSAD is greater than or equal to a given threshold (THR).

[0151]

[0152] Although the first to fourth embodiments compared the resolution with 4K, the resolution may be determined based on 8K rather than 4K. That is, cases less than 4K may be considered less than 8K, and cases greater than 4K may be considered 8K or more, and the like, may be included in the embodiments of the present invention.

[0153] The threshold values ​​of the first to fourth embodiments may be any one of 0.1, 0.5, 1, 2, 3, 4, and 5.

[0154] In determining the interpolation filter of the present invention, selection of the interpolation filter may be performed by considering correlation or gradient instead of MeanSAD.

[0155] The video decoding / encoding method of the present invention comprises the steps of: determining an intra-screen prediction mode used for intra-screen prediction of a current block; interpolating a reference pixel used for the intra-screen prediction; and performing intra-screen prediction of the current block based on the interpolated reference pixel and the intra-screen prediction mode, wherein the type of interpolation filter applied to the interpolation of the reference pixel includes a 4-tap filter and an 8-tap filter, and the interpolation filter can be determined based on the size of the current block and the mean value (MeanSAD) of absolute values ​​of difference values ​​of reference pixels of the current block.

[0156] The determination of the interpolation filter used for the intra-screen prediction and the intra-screen prediction of the present invention can be performed in the intra-screen prediction unit (102, 207).

[0157]

[0158] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.

[0159] The present invention has applicability in the video signal encoding / decoding industry.

Claims

1. A step of determining an on-screen prediction mode used for on-screen prediction of the current block; A step of interpolating reference pixels used for prediction within the above screen; A step of performing intra-screen prediction of the current block based on the interpolated reference pixel and the intra-screen prediction mode, The types of interpolation filters applied to the interpolation of the above reference pixels include 4-tap filters and 8-tap filters, An image decoding method, wherein the interpolation filter is determined based on the size of the current block and the average value (MeanSAD) of the absolute values ​​of the difference values ​​of the reference pixels of the current block.

2. In paragraph 1, An image decoding method, wherein the 8-tap filter includes an 8-tap DCT-IF (Discrete Cosine Transform based Interpolation Filter) and an 8-tap SIF (Smoothing Interpolation Filter).

3. In paragraph 1, A video decoding method, wherein the above interpolation filter is determined by further considering the result of comparing the resolution of the current picture including the current block with 4K resolution (3840x2160).

4. In paragraph 3, If the resolution of the current picture is less than the 4K resolution, the interpolation filter is determined based on the size of the current block and the MeanSAD, A video decoding method, wherein the interpolation filter is determined based on the size of the current block, regardless of the MeanSAD, when the resolution of the current picture is equal to or greater than the 4K resolution.

5. In paragraph 4, In case the resolution of the current picture is less than the 4K resolution, If the size of the current block is 2 and the MeanSAD is less than the threshold, the interpolation filter is determined as 4-tap SIF, A method for decoding an image, wherein the interpolation filter is determined as an 8-tap DCT-IF when the size of the current block is 2 and the MeanSAD is greater than or equal to a threshold value.

6. In paragraph 4, In case the resolution of the current picture is less than the 4K resolution, If the size of the current block is 5 or more and the MeanSAD is less than the threshold, the interpolation filter is determined as 8-tap SIF, An image decoding method, wherein the interpolation filter is determined as a 4-tap DCT-IF when the size of the current block is 5 or more and the MeanSAD is greater than or equal to a threshold value.

7. A step of determining an on-screen prediction mode used for on-screen prediction of the current block; A step of interpolating reference pixels used for prediction within the above screen; A step of performing intra-screen prediction of the current block based on the interpolated reference pixel and the intra-screen prediction mode, The types of interpolation filters applied to the interpolation of the above reference pixels include 4-tap filters and 8-tap filters, A method for encoding an image, wherein the interpolation filter is determined based on the size of the current block and the average value (MeanSAD) of the absolute values ​​of the differences of reference pixels of the current block.

8. In paragraph 7, An image encoding method, wherein the 8-tap filter includes an 8-tap DCT-IF (Discrete Cosine Transform based Interpolation Filter) and an 8-tap SIF (Smoothing Interpolation Filter).

9. In paragraph 7, A video encoding method wherein the above interpolation filter is determined by further considering the result of comparing the resolution of the current picture including the current block with 4K resolution (3840x2160).

10. In paragraph 9, If the resolution of the current picture is less than the 4K resolution, the interpolation filter is determined based on the size of the current block and the MeanSAD, A video encoding method, wherein the interpolation filter is determined based on the size of the current block, regardless of the MeanSAD, when the resolution of the current picture is equal to or greater than the 4K resolution.

11. In paragraph 10, In case the resolution of the current picture is less than the 4K resolution, If the size of the current block is 2 and the MeanSAD is less than the threshold, the interpolation filter is determined as 4-tap SIF, A method for encoding an image, wherein the interpolation filter is determined as an 8-tap DCT-IF when the size of the current block is 2 and the MeanSAD is greater than or equal to a threshold value.

12. In paragraph 10, In case the resolution of the current picture is less than the 4K resolution, If the size of the current block is 5 or more and the MeanSAD is less than the threshold, the interpolation filter is determined as 8-tap SIF, A video encoding method, wherein when the size of the current block is 5 or more and the MeanSAD is greater than or equal to a threshold value, the interpolation filter is determined as a 4-tap DCT-IF.

13. A method for transmitting a bitstream generated by a video encoding method, The above image encoding method comprises the steps of: determining an intra-screen prediction mode used for intra-screen prediction of a current block; A step of interpolating reference pixels used for prediction within the above screen; A step of performing intra-screen prediction of the current block based on the interpolated reference pixel and the intra-screen prediction mode, The types of interpolation filters applied to the interpolation of the above reference pixels include 4-tap filters and 8-tap filters, A method for transmitting a bitstream, wherein the interpolation filter is determined based on the size of the current block and the average value (MeanSAD) of the absolute values ​​of the differences of reference pixels of the current block.

Citation Information

Patent Citations

  • Method and apparatus for interpolating image using smoothed interpolation filter

    KR101707610B1

  • A system generating automatically workbook for coding robot

    KR1020240012928A

  • Camera module for vehicle

    KR1020250079601A

  • A novel compound having antioxidant activity and anti-inflammatory activity and a composition for antioxidant activity and anti-inflammatory according to thereof

    KR102657359B1

  • Harmonization of prediction-domain filters with interpolation filtering

    US20200252653A1