Video encoding and decoding apparatus and method
By deriving quantization blocks from the bitstream and performing inverse transformation, using appropriate transformation cores to improve the encoding efficiency of the video signal, the problem of low encoding efficiency of video signals in the prior art is solved, and more efficient signal processing is achieved.
Patent Information
- Application Number
- CN202380076947.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-21
- Filing Date
- 2023-10-04
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to improve the encoding efficiency of video signals, especially in the face of a rapidly growing volume of multimedia data.
By deriving the quantization block from the bitstream, inverse quantization is achieved to obtain the transform block, and performing inverse transformation according to the determined transform core, thereby improving the encoding efficiency of the video signal.
By applying appropriate cores to transform, the encoding efficiency of video signals is significantly improved and the performance of signal processing is improved.
Smart Images

Figure CN120202672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video encoding and decoding apparatus and method, and more particularly, to a video encoding and decoding apparatus and method for deriving a kernel of at least one of a primary transform and a secondary transform and applying the derived transform kernel to a corresponding transform. Background Art
[0002] Recently, on the Internet, the demand for multimedia data such as video has grown rapidly. However, the speed of development of channel bandwidth is difficult to keep up with the rapidly increasing amount of multimedia data. Summary of the Invention
[0003] Technical Problem
[0004] The present invention aims to improve the encoding efficiency of video signals.
[0005] Solution to the Technical Problem
[0006] The encoding / decoding method and computer-readable recording medium storing a bitstream of the present invention may include: dequantizing a quantized block obtained from the bitstream to obtain a secondary transform block; determining whether to perform a secondary inverse transform on the secondary transform block; when it is determined to perform the secondary inverse transform, performing the secondary inverse transform on the secondary transform block to obtain a primary transform block; and performing a primary inverse transform on the primary transform block.
[0007] In the encoding / decoding method and computer-readable recording medium storing a bitstream of the present invention, the transform kernel of the secondary inverse transform and the transform kernel of the primary inverse transform may be specified by an index signaled from the bitstream.
[0008] In the encoding / decoding method and computer-readable recording medium storing a bitstream of the present invention, the maximum value and configuration of the index may vary depending on whether the applied transform kernel is a one-dimensional transform kernel or a two-dimensional transform kernel.
[0009] In the encoding / decoding method and computer-readable recording medium storing a bitstream of the present invention, the transform kernel of the secondary inverse transform and the transform kernel of the primary inverse transform may include a Karhunen Loeve Transform (KLT).
[0010] In the encoding / decoding method and computer-readable recording medium storing a bitstream of the present invention, the size of the output block of at least one of the secondary inverse transform and the primary inverse transform may be smaller than the size of the input block.
[0011] In the encoding / decoding method and computer-readable recording medium storing a bitstream of the present invention, it is possible to determine whether to perform the secondary transform based on at least one of the type of transform kernel, the number of transform coefficients, and the size of the block.
[0012] The encoding / decoding method and computer-readable recording medium storing a bitstream of the present invention may include: dequantizing a quantized block obtained from the bitstream to obtain a primary transform block; determining a transform kernel for a primary inverse transform of the primary transform block; and performing the primary inverse transform on the primary transform block based on the determined transform kernel of the primary inverse transform to obtain a residual block.
[0013] The encoding / decoding method and computer-readable recording medium storing a bitstream of the present invention include: determining an intra prediction mode of a current block; and interpolating reference pixels for the intra prediction mode, where the reference pixels are included in adjacent reference blocks around the current block, and an interpolation filter applied to the interpolation includes an 8-tap filter.
[0014] Technical Effects
[0015] The video encoding and decoding apparatus and method according to the present invention can improve the encoding efficiency of a video signal by applying an appropriate kernel to a transform. Brief Description of the Drawings
[0016] Figure 1 is a block diagram showing an image encoding apparatus according to an embodiment of the present invention.
[0017] Figure 2 is a block diagram showing an image decoding apparatus (200) according to an embodiment of the present invention.
[0018] Figure 3 is a diagram showing an off-line training process for deriving a transform kernel.
[0019] Figure 4 shows representative blocks obtained by obtaining representative values and performing clustering.
[0020] Figure 5 shows an embodiment in which a 1D (one-dimensional) transform kernel is applied.
[0021] Figure 6 shows an embodiment in which a 2D transform kernel is applied.
[0022] Figure 7 is a diagram showing a transform process in an encoder.
[0023] Figure 8 is a diagram showing an inverse transform process in a decoder.
[0024] Figure 9 It is a diagram showing a first embodiment of applying a 1D transform kernel by using dimension reduction.
[0025] Figure 10 It is a diagram showing a second embodiment of applying a 1D transform kernel by using dimension reduction.
[0026] Figure 11 It is a diagram showing a third embodiment of applying a 1D transform kernel by using dimension reduction.
[0027] Figure 12 It is a diagram showing a fourth embodiment of applying a 1D transform kernel when transforming all data without dimension reduction.
[0028] Figure 13 It is a diagram showing a first embodiment of applying a 2D transform kernel by using dimension reduction.
[0029] Figure 14 It is a diagram showing a second embodiment of applying a 2D transform kernel by using dimension reduction.
[0030] Figure 15 It is a diagram showing a second embodiment of applying a 2D transform kernel by using dimension reduction.
[0031] Figure 16 It is a diagram showing a third embodiment of applying a 2D transform kernel by using dimension reduction.
[0032] Figure 17 An example of a vector being rearranged into a 2D block is shown.
[0033] Figure 18 An embodiment of scanning the coefficients of the rearranged block is shown.
[0034] Figure 19 It is a diagram showing an embodiment of signaling a 1D transform kernel for primary transform.
[0035] Figure 20 It is a diagram showing an embodiment of signaling a 2D transform kernel for primary transform.
[0036] Figure 21 It is a diagram showing an embodiment of signaling a 1D transform kernel for secondary transform.
[0037] Figure 22 It is a diagram showing an embodiment of signaling a 2D transform kernel for secondary transform.
[0038] Figure 23 An example of applying only the primary transform in an encoder is shown.
[0039] Figure 24 Shows an example of only applying the primary inverse transform in the decoder.
[0040] Figure 25 Is a diagram for describing an example of using only some columns out of MN (= M × N) columns.
[0041] Figure 26 Is a diagram showing the transform kernel size related to a signal and the corresponding inverse kernel.
[0042] Figure 27 Is a diagram showing the scanning of one-dimensional data.
[0043] Figure 28 Is a diagram showing h[n], y[n], and z[n] for obtaining 8-tap coefficients.
[0044] Figure 29 Shows the integer reference samples for deriving 8-tap SIF coefficients.
[0045] Figure 30 Is a diagram showing the directions and angles of intra prediction modes.
[0046] Figure 31 Is a diagram showing the average correlation values of reference samples for multiple video resolutions and each nTbS.
[0047] Figure 32 Shows an example of a method for selecting an interpolation filter by using frequency information.
[0048] Figure 33 Shows an embodiment of 8-tap DCT interpolation filter coefficients.
[0049] Figure 34 Shows an embodiment of 8-tap smoothing interpolation filter coefficients.
[0050] Figure 35 Shows the magnitude responses of 4-tap DCT-IF, 4-tap SIF, 8-tap DCT-IF, and 8-tap SIF at the 16 / 32 pixel positions.
[0051] Figure 36 Shows a diagram related to each threshold according to nTbS.
[0052] Figure 37 Shows the sequence name, screen size, screen rate, and bit depth of the CTC video sequences for each category.
[0053] Figure 38 Shows the interpolation filter selection method and the interpolation filter applied according to the selected method to test the efficiency of 8-tap / 4-tap interpolation filters.
[0054] Figure 39 Tables IX and X therein show the simulation results of Methods A, B, C, and D.
[0055] Figure 40 Table XI therein shows the ratio of CUs that apply 4-tap DCT-IF at the VVC anchor points and 8-tap DCT-IF based on high_freq_ratio among the methods proposed for all test sequences.
[0056] Figure 41 show the experimental results of the proposed filtering method. Detailed implementation
[0057] Optimal embodiments for implementing the present invention
[0058] The present invention can inverse-quantize the quantized blocks obtained from the bitstream to obtain secondary transform blocks, determine whether to perform a secondary inverse transform on the secondary transform blocks, when it is determined to perform the secondary inverse transform, perform the secondary inverse transform on the secondary transform blocks to obtain primary transform blocks, and perform a primary inverse transform on the primary transform blocks.
[0059] Since the present disclosure can make various changes and has multiple embodiments, specific embodiments are shown in the drawings and are described in detail in the detailed implementation. However, this does not limit the present disclosure to the specific embodiments, and it should be understood to include all changes, equivalent solutions, and alternative solutions included in the concept and technical scope of the present disclosure. When describing each drawing, similar reference numerals are used for similar elements.
[0060] Terms such as first, second, etc. can be used to describe various elements, but the elements should not be limited by these terms. These terms are only used to distinguish one element from other elements. For example, without departing from the scope of the rights of the present disclosure, the first element can be called the second element, and similarly, the second element can also be called the first element. The term and / or includes a combination of multiple related description items or any one of the multiple related description items.
[0061] When an element is referred to as "connected" or "linked" to another element, it should be understood that it can be directly connected or linked to that other element, but there may be another element between them. In addition, when an element is referred to as "directly connected" or "directly linked" to another element, it should be understood that there is no other element between them.
[0062] The terms used in this disclosure are only for describing specific embodiments and are not intended to limit this disclosure. Unless otherwise clearly specified in the context, singular expressions include plural expressions. In this disclosure, it should be understood that terms such as "including" or "having" are only intended to specify the existence of the features, numbers, steps, operations, elements, components, or combinations thereof described in this specification, and it does not preclude the possibility of the existence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations thereof.
[0063] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Hereinafter, the same reference numerals are used for the same elements in the drawings, and repeated descriptions of the same elements are omitted.
[0064] Different from the case of using separable horizontal or vertical one-dimensional transforms (such as signal-independent DCT-2 (Discrete Cosine Transform-2), DCT-8 (Discrete Cosine Transform-8), DST-7 (Discrete Sine Transform-7), etc.) used in video compression / reconstruction standards, the present invention proposes a method for obtaining data after transformation when using a signal-dependent KL (Karhunen-Loeve) or SVD (Singular Value Decomposition) transform that utilizes the covariance and correlation of each two-dimensional block (here, the block represents a residual signal block or a transform block), and a method for correspondingly rearranging and scanning transform coefficients.
[0065] Figure 1 is a block diagram showing an image coding apparatus according to an embodiment of the present invention.
[0066] Referring to Figure 1 , the image coding apparatus 100 may include an image partitioner 101, an intra predictor 102, an inter predictor 103, a subtractor 104, a transformer 105, a quantizer 106, an entropy encoder 107, an inverse quantizer 108, an inverse transformer 109, an adder 110, a filter 111, and a memory 112.
[0067] Figure 1 Each structural unit shown in is independently shown to represent different characteristic functions in the image coding apparatus, and it does not mean that each structural unit is composed of a separate hardware or a single software unit. In other words, for the sake of convenience of description, each structural unit is included by being listed as each structural unit, and at least two structural units in each structural unit may be combined to form one structural unit, or one structural unit may be divided into multiple structural units to perform functions, and the integrated embodiments and separate embodiments of each structural unit are also included within the scope of the claims of the present invention as long as they do not deviate from the essence of the present invention.
[0068] In addition, some elements are not essential elements for performing the necessary functions in the present invention and can be optional elements only for improving performance. The present invention can be implemented by including only the structural units necessary for realizing the essence of the present invention without including elements only for improving performance, and the structure including only the essential elements without including the optional elements only for improving performance is also included within the scope of the claims of the present invention.
[0069] The image partitioner 100 can partition an input image into at least one block. In this case, the input image can have various shapes and sizes, such as a picture, a strip, parallel blocks, segments, etc. The blocks can represent coding units (CUs), prediction units (PUs), or transform units (TUs). The partitioning can be performed based on at least one of a quadtree and a binary tree. A quadtree is a method for partitioning an upper-level block into four lower-level blocks (where the width and height are half of the width and height of the upper-level block). A binary tree is a method for partitioning an upper-level block into two lower-level blocks (where any one of the width and height is half of the width or height of the upper-level block). Through the above binary-tree-based partitioning, the blocks can have non-square shapes as well as square shapes.
[0070] The predictors 102 and 103 can include an inter-frame predictor 103 that performs inter-frame prediction and an intra-frame predictor 102 that performs intra-frame prediction. It can be determined whether to perform inter-frame prediction or intra-frame prediction on a prediction unit, and the specific information according to each prediction method (e.g., intra-frame prediction mode, motion vector, reference picture, etc.) can be determined. In this case, the processing unit that performs the prediction and the processing unit that determines the prediction method and details can be different. For example, the prediction method, prediction mode, etc. can be determined in the prediction unit, and the prediction can be performed in the transform unit.
[0071] The residual value (residual block) between the generated prediction block and the original block can be input to the transformer 105. In addition, prediction mode information, motion vector information, etc. for prediction can be encoded with the residual value in the entropy encoder 107 and sent to the decoder. When using a specific coding mode, it is also possible to encode the original block as it is without generating a prediction block through the predictors 102 and 103 and send it to the decoder.
[0072] The intra-frame predictor 102 can generate a prediction block based on the reference pixel information around the current block, which is pixel information within the current picture. When the prediction mode of an adjacent block of the current block for which intra-frame prediction is to be performed is inter-frame prediction, the reference pixels included in the adjacent block to which inter-frame prediction is applied can be replaced with the reference pixels within another adjacent block to which intra-frame prediction is applied. In other words, when the reference pixels are not available, the unavailable reference pixel information can be used by replacing it with at least one of the available reference pixels.
[0073] In intra prediction, the prediction mode may have a directional prediction mode that uses reference pixel information according to the prediction direction and a non-directional mode that does not use directional information when performing prediction. The mode for predicting luminance information and the mode for predicting chrominance information may be different, and the intra prediction mode information for predicting luminance information or luminance signal information may be used to predict chrominance information.
[0074] The intra predictor 102 may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolator, and a DC filter. The AIS filter is a filter that performs filtering on the reference pixels of the current block and may adaptively determine whether to apply the filter according to the prediction mode in the current prediction unit. When the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.
[0075] When the intra prediction mode in the prediction unit is a prediction unit that performs intra prediction based on the pixel values obtained by interpolating the reference pixels, the reference pixel interpolator of the intra predictor 102 may interpolate the reference pixels to generate reference pixels at fractional unit positions. When the prediction mode in the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixels, the reference pixels may not be interpolated. When the prediction mode of the current block is the DC mode, the DC filter may generate a prediction block through filtering.
[0076] The inter predictor 103 generates a prediction block by using the motion information stored in the memory 112 and the pre-reconstructed reference image. The motion information may include, for example, motion vectors, reference picture indices, list 1 prediction flags, list 0 prediction flags, etc.
[0077] A residual block including residual value information may be generated, where the residual value information is the difference between the prediction units generated in the predictors 102 and 103 and the original block in the prediction unit. The generated residual block may be input to the transformer 130 and be transformed.
[0078] The inter predictor 103 may derive a prediction block based on information of at least one of the previous picture and the subsequent picture of the current picture. Additionally, based on information of some regions that have been encoded within the current picture, a prediction block of the current block may be derived. The inter predictor 103 according to an embodiment of the present invention may include a reference picture interpolator, a motion predictor, and a motion compensator.
[0079] The reference picture interpolator can receive reference picture information from the memory 112 and generate pixel information less than or equal to integer pixels from the reference pictures. For luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than or equal to integer pixels in 1 / 4 pixel units. For chrominance signals, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than or equal to integer pixels in 1 / 8 pixel units.
[0080] The motion predictor can perform motion prediction based on the reference pictures interpolated by the reference picture interpolator. Various methods such as FBMA (Full Search-based Block Matching Algorithm), TSS (Three-Step Search), NTS (New Three-Step Search Algorithm), etc. can be used as methods for calculating the motion vector. The motion vector can have a motion vector value in 1 / 2 or 1 / 4 pixel units based on the interpolated pixels. The motion predictor can predict the predicted block of the current block by changing the motion prediction method. Various methods such as the skip block method, the merge method, the AMVP (Advanced Motion Vector Prediction) method, etc. can be used as motion prediction methods.
[0081] The subtractor 104 subtracts the current block to be encoded from the predicted block generated in the intra predictor 102 or the inter predictor 103 to generate the residual block of the current block.
[0082] The transformer 105 can transform the residual block including the residual data by using transformation methods such as DCT, DST, KLT (Karhunen Loeve Transform), SVD, etc. In this case, the transformation method (or transformation kernel) can be determined based on the intra prediction mode in the prediction unit used to generate the residual block. For example, according to the intra prediction mode, DCT can be used in the horizontal direction and DST can be used in the vertical direction.
[0083] The quantizer 106 can quantize the values transformed into the frequency domain in the transformer 105. The quantization coefficient can vary according to the importance of the block or the image. The values calculated in the quantizer 106 can be provided to the inverse quantizer 108 and the entropy encoder 107.
[0084] The transformer 105 and / or the quantizer 106 may be selectively included in the image coding device 100. In other words, the image coding device 100 may encode a residual block by performing at least one of transformation and quantization on the residual data of the residual block or by skipping both transformation and quantization. Even when neither transformation nor quantization is performed in the image coding device 100 or when both transformation and quantization are not performed, the block input as the input to the entropy encoder 107 is generally referred to as a transformed block. The entropy encoder 107 performs entropy coding on the input data. The entropy coding may use various coding methods such as exponential Golomb, CAVLC (context adaptive variable length coding), CABAC (context adaptive binary arithmetic coding), etc.
[0085] The entropy encoder 107 may encode various information such as coefficient information of the transformed block, block type information, prediction mode information, partition unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, filtering information, etc. The coefficients of the transformed block may be encoded in sub-block units within the transformed block.
[0086] To encode the coefficients of the transformed block, various syntax elements may be encoded, such as the syntax element Last_sig showing the position of the first non-zero coefficient according to the scan order, the flag Coded_sub_blk_flag showing whether there is at least one non-zero coefficient in the sub-block, the flag Sig_coeff_flag showing whether the coefficient is non-zero, the flag Abs_greaterN_flag showing whether the absolute value of the coefficient is greater than N (where N may be a natural number such as 1, 2, 3, 4, 5, etc.), and the flag Sign_flag showing the sign of the coefficient. The residual values of the coefficients not encoded separately by the syntax elements may be encoded by the syntax element remaining_coeff.
[0087] The inverse quantizer 108 and the inverse transformer 109 perform inverse quantization on the values quantized in the quantizer 106 and perform inverse transformation on the values transformed in the transformer 105. The residual values generated in the inverse quantizer 108 and the inverse transformer 109 may be combined with the prediction units predicted by the motion estimators, motion compensators, and intra-predictor 102 included in the predictors 102 and 103 to generate a reconstructed block. The adder 110 adds the prediction block generated in the predictors 102 and 103 and the residual block generated by the inverse transformer 109 to generate a reconstructed block.
[0088] The filter 111 may include at least one of a deblocking filter, an offset correction unit, and an ALF (adaptive loop filter).
[0089] The deblocking filter can remove block distortion generated by the boundaries between blocks in the reconstructed picture. To determine whether to perform deblocking processing, it can be determined whether to apply the deblocking filter to the current block based on the pixels included in several columns or rows included in the block. When the deblocking filter is applied to a block, a strong filter or a weak filter can be applied according to the required deblocking filtering strength. Additionally, when performing vertical filtering and horizontal filtering during the application of the deblocking filter, the horizontal filtering and the vertical filtering can be processed in parallel.
[0090] The offset correction unit can correct the offset relative to the original image in pixel units for the image on which deblocking processing is performed. To perform offset correction on a specific picture, a method for dividing the pixels included in the image into a certain number of regions, determining the regions where the offset will be performed, and applying the offset to the corresponding regions, or a method for applying the offset by considering the edge information of each pixel can be used.
[0091] ALF (Adaptive Loop Filtering) can be performed based on the value obtained by comparing the filtered reconstructed image with the original image. After dividing the pixels included in the image into predetermined groups, it can be determined one filter to be applied to the corresponding group to perform filtering differentially for each group. The information related to whether ALF is applied can be transmitted through the coding unit (CU) of the luminance signal, and the shape and filter coefficients of the ALF filter to be applied for each block can be different. Additionally, an ALF filter of the same form (fixed form) can be applied regardless of the characteristics of the block to which it is to be applied.
[0092] The memory 112 can store the reconstructed blocks or pictures calculated by the filter 111, and can provide the stored reconstructed blocks or pictures to the predictors 102 and 103 during the execution of inter-frame prediction.
[0093] Next, an image decoding apparatus according to an embodiment of the present invention will be described with reference to the accompanying drawings. Figure 2 FIG. is a block diagram showing an image decoding apparatus 200 according to an embodiment of the present invention.
[0094] Referring to Figure 2 , the image decoding apparatus 200 may include an entropy decoder 201, an inverse quantizer 202, an inverse transformer 203, an adder 204, a filter 205, a memory 206, and predictors 207 and 208.
[0095] When the image bitstream generated by the image encoding apparatus 100 is input to the image decoding apparatus 200, the input bitstream can be decoded according to the processing opposite to that performed in the image encoding apparatus 100.
[0096] The entropy decoder 201 can perform entropy decoding in a process opposite to the process of performing entropy encoding in the entropy encoder 107 of the image encoding apparatus 100. For example, in response to the method executed in the image encoder, various methods such as exponential Golomb, CAVLC (Context Adaptive Variable Length Coding), and CABAC (Context Adaptive Binary Arithmetic Coding) can be applied. The entropy decoder 201 can decode the above-mentioned syntax elements (i.e., Last_sig, Coded_sub_blk_flag, Sig_coeff_flag, Abs_greaterN_flag, Sign_flag, and remaining_coeff). Additionally, the entropy decoder 201 can decode information related to intra prediction and inter prediction performed in the image encoding apparatus 100.
[0097] The inverse quantizer 202 performs inverse quantization on the quantized transform block to generate a transform block. The inverse quantizer 202 operates in substantially the same manner as the inverse quantizer 108 in Figure 1 .
[0098] The inverse transformer 203 performs an inverse transform on the transform block to generate a residual block. In this case, the transform method can be determined based on information regarding the prediction method (inter prediction or intra prediction), the size and / or shape of the block, the intra prediction mode, etc. The inverse transformer 203 operates in substantially the same manner as the inverse transformer 109 in Figure 1 .
[0099] The adder 204 adds the prediction block generated in the intra predictor 207 or the inter predictor 208 and the residual block generated by the inverse transformer 203 to generate a reconstructed block. The adder 204 operates in substantially the same manner as the adder 110 in Figure 1 .
[0100] The filter 205 reduces various types of noise that appear in the reconstructed block.
[0101] The filter 205 may include a deblocking filter, an offset correction unit, and an ALF.
[0102] Information regarding whether the deblocking filter is applied to the corresponding block or picture and when the deblocking filter is applied, and information regarding whether a strong filter or a weak filter is applied can be provided from the image encoding apparatus 100. The deblocking filter of the image decoding apparatus 200 can receive information related to the deblocking filter provided from the image encoding apparatus 100 and perform deblocking filtering on the corresponding block in the image decoding apparatus 200.
[0103] The offset correction unit can perform offset correction on the reconstructed image based on the type of offset correction applied to the image during encoding, offset value information, etc.
[0104] The ALF can be applied to a coding unit based on information about whether to apply ALF, ALF coefficient information, etc., provided from the image coding apparatus 100. The ALF information can be included in a specific parameter set and provided. The filter 205 operates in substantially the same manner as Figure 1 the filter 111 in
[0105] The memory 206 stores the reconstructed block generated by the adder 204. The memory 206 operates in substantially the same manner as Figure 1 the memory 112 in
[0106] The predictors 207 and 208 can generate a prediction block based on correlation information provided from the entropy decoder 201 and pre-decoded block or picture information provided from the memory 206 for the prediction block.
[0107] The predictors 207 and 208 can include an intra predictor 207 and an inter predictor 208. Although not shown separately, the predictors 207 and 208 may further include a prediction unit determination unit. The prediction unit determination unit can receive various information (such as prediction unit information input from the entropy decoder 201, prediction mode information of an intra prediction method, motion prediction related information of an inter prediction method, etc.) to distinguish the prediction unit from the current coding unit and determine whether the prediction unit performs inter prediction or intra prediction. The inter predictor 208 can perform inter prediction on the current prediction unit by using information required for inter prediction in the current prediction unit provided from the image encoder 100, based on information included in at least one of a previous picture and a subsequent picture of the current picture including the current prediction unit. Alternatively, inter prediction can be performed based on information of a previously reconstructed partial region within the current picture including the current prediction unit.
[0108] To perform inter prediction, it can be determined based on the coding unit whether the motion prediction method in the prediction unit included in the corresponding coding unit is a skip mode, a merge mode, or an AMVP mode.
[0109] The intra predictor 207 generates a prediction block by using pixels located around the block to be currently coded and previously reconstructed.
[0110] The intra predictor 207 can include an adaptive intra smoothing (AIS) filter, a reference pixel interpolator, and a DC filter. The AIS filter is a filter that performs filtering on the reference pixels of the current block, and can adaptively determine whether to apply the filter according to the prediction mode in the current prediction unit. AIS filtering can be performed on the reference pixels of the current block by using AIS filtering information and the prediction mode in the prediction unit provided from the image coding apparatus 100. When the prediction mode of the current block is a mode in which AIS filtering is not performed, the AIS filter may not be applied.
[0111] When the prediction mode in the prediction unit is a prediction unit that performs intra prediction based on the pixel values obtained by interpolating reference pixels, the reference pixel interpolator of the intra predictor 207 can interpolate the reference pixels to generate reference pixels at fractional unit positions. The reference pixels generated at the fractional unit positions can be used as prediction pixels for the pixels within the current block. When the prediction mode in the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixels, the reference pixels may not be interpolated. When the prediction mode of the current block is the DC mode, the DC filter can generate a prediction block by filtering.
[0112] The intra predictor 207 operates in substantially the same manner as Figure 1 the intra predictor 102 in
[0113] The inter predictor 208 generates an inter prediction block by using the reference pictures and motion information stored in the memory 206. The inter predictor 208 operates in substantially the same manner as Figure 1 the inter predictor 103 in
[0114] Figure 3 is a diagram showing an offline training process for deriving a transform kernel.
[0115] The method for deriving the transform kernel can be described in Figure 3 . First, in order to derive the clustering and the representative value of each cluster, the information extracted from the encoding process can be used. Here, the information may include residual blocks, primary transform blocks, etc. This information can be used as the input information in the steps for deriving the transform kernel. As an example, when deriving the primary transform kernel, the residual block can be used as the input information for offline training. When deriving the secondary transform kernel, the primary transform block can be used as the input information for offline training.
[0116] In addition, in order to quickly obtain training data, the primary transform kernel can be derived by using the residual blocks directly reconstructed in the decoder, and the secondary transform kernel can be derived by using the reconstructed primary transform blocks.
[0117] Hereinafter, each step for deriving the Figure 3 transform kernel in
[0118] [1] Steps for clustering and setting the representative value of each cluster
[0119] Multiple blocks can be clustered by using the K-means algorithm (or the Isodata algorithm, etc.), and the representative value of each cluster can be set.
[0120] For the primary transformation, a residual block with dimensions of M×N (M: height, N: width) can be used. For the secondary transformation, a primary transformation block with dimensions of M×N (M: height, N: width) can be used.
[0121] In this case, for both the primary / secondary transformation, the number of clusters and the number of representative values can be set to any K. Here, K can be an integer such as 1, 2, 3, 4, etc.
[0122] All or a part of the K representative values can be used to derive the transformation kernel. As an example, x of the K representative values can be selected and used to derive the transformation kernel. Here, x can be a natural number less than K.
[0123] As an example, when obtaining representative values of 30,000 residual blocks with a size of 4×4 blocks, K can be set to 3. Since K is set to 3, a total of 3 representative values can be obtained. In this case, 2 of the 3 representative values can be selected and used to derive the transformation kernel. Therefore, 2 transformation kernels can be derived by 2 representative values in the step [2] for deriving the transformation kernel.
[0124] [2] Step for deriving the transformation kernel
[0125] The KL or SVD transformation kernel is derived by using the covariance matrix or correlation matrix of the blocks clustered in the step [1] for clustering and setting representative values for each cluster. In this case, the size of the blocks used for deriving the kernel can be M×N (M×N can be represented using matrix notation).
[0126] [2]-A. Method for deriving a 2D KL transformation kernel through the covariance matrix
[0127] First, Cor (MN×MN) 's eigenvalue λ can be derived.
[0128] (As a reference, when deriving the SVD transformation kernel, the correlation Cor (MN×MN) = μμ T can be used instead of the covariance.)
[0129] First, as a condition for deriving the eigenvalue λ, |Cov MN×MN - λI| = 0, and λ can have a range from λ1 to λ MN . λ1 > λ2 >... > λ MN can show the order of high energy.
[0130] The covariance matrix can be derived by Equation 1 below.
[0131] [Equation 1]
[0132]
[0133] Here, x i can represent a cluster block in the form of an MN×1 vector. Additionally, it can represent the MN×1 average vector of L sample data.
[0134] After that, the orthonormal φ i vector can be calculated according to the value of λ i (φ MN×MN )(i = 1, 2, 3, …, MN). After that, it can be used to derive the (KL transform kernel) as in Equation 2. can refer to the transpose of φ MN×MN . MN refers to M×N. As an example, if M = 4 and N = 8, then MN can be 32. Equation 2 can show the definition of φ MN×MN in the equation.
[0135] [Equation 2]
[0136]
[0137] [2]-B. Method for Deriving 1D KL Transform Kernel (Using Vertical Covariance Matrix and Horizontal Covariance Matrix)
[0138] [Method for Deriving 1D KL Transform Kernel by Using Vertical Covariance Matrix]
[0139] To derive the vertical kernel, the eigenvalues λ of V Cov (M×N) can be derived.
[0140] (As a reference, when deriving the SVD transform kernel, correlation can be used instead of covariance.)
[0141] First, as a condition for deriving the eigenvalue λ, |V Cov M×M -λI| = 0, and λ can have a range from λ1 to λ M . λ1 > λ2 >... > λ M can show the order of high energy.
[0142] After calculating the orthonormal φ i vector according to the value of λ1, the (vertical KL transform kernel) can be derived as in Equation 3. can be the transpose of φ M×M .
[0143] The vertical covariance matrix can be derived by the following Equation 3.
[0144] [Equation 3]
[0145]
[0146] Here, i can be 1, 2, 3, …, N, x ij can be an M×1 vector, i can be a column, j can be a sample block number, and μ i can be the M×1 mean vector of the i-th column of x ij .
[0147] Here, in Figure 4 is shown x obtained by "[1] Steps for clustering and setting representative values for each cluster" ij and μ i .
[0148] After that, it can be used to derive φ Figure 4 as M×M (vertical KL transform kernel). Equation 4 can show the definition of φ M×M in the equation.
[0149] [Equation 4]
[0150]
[0151] [Method for deriving 1D KL transform kernel by using horizontal covariance matrix]
[0152] To derive the horizontal kernel, the eigenvalues λ of HCov N×N can be derived.
[0153] (As a reference, when deriving the SVD transform kernel, correlation can be used instead of covariance.)
[0154] First, as a condition for deriving the eigenvalue λ, |HCov N×N -λI| = 0, and λ can have a range from λ1 to λ N . λ1 > λ2 >... > λ N can show the order of high energy.
[0155] After calculating the orthonormal φ i vector according to the value of λ1, can be derived as in Equation 5 (horizontal KL transform kernel). can be the transpose of φ N×N .
[0156] [Equation 5]
[0157]
[0158] Here, i can be 1, 2, 3, …, N, x ij can be an M×1 vector, i can be a row, j can be a sample block number, and μi It can be x ij The M×1 average vector of the i-th row of
[0159] After that, it can be used to derive φ as Figure 6 to derive φ N×N (horizontal KL transform kernel).
[0160] [Equation 6]
[0161]
[0162] An embodiment of applying the kernel derived in the steps for deriving the transform kernel in [2] is described below. In this case, it is mainly described based on the KL transform kernel, but it can also be applied to the SVD transform kernel in the same way.
[0163] Figure 5 An embodiment of applying a 1D (one-dimensional) transform kernel is shown.
[0164] Figure 5 The embodiment in Figure 5 is an example of applying a 1D transform kernel (separable KLT), and examples of applying the 1D transform kernel to the primary transform or the secondary transform are described. The input value in
[0165] Figure 6 An embodiment of applying a 2D transform kernel is shown. Figure 6 The embodiment in Figure 6 is an example of applying a 2D transform kernel (non-separable KLT). The input value in
[0166] Figure 7 is a diagram showing the transform processing in the encoder.
[0167] First, a first transform can be performed on the residual block to obtain a primary transform block. It can be determined whether to perform a secondary transform on the primary transform block. When it is determined to perform the secondary transform, the secondary transform can be performed to obtain a secondary transform block. The quantization block obtained by quantizing the secondary transform block can be encoded into the bitstream.
[0168] The transform kernels applied to the primary transform using residual blocks may include kernels such as DST and DCT, and may also include kernels such as KLT. The KLT kernel may be selectively included. Whether to include the KLT kernel may be determined according to the characteristics of the residual block. Here, the characteristics of the residual block may include the width, height, size, shape, partition depth, etc. of the residual block. In contrast, the information indicating whether the KLT kernel is selectively included may be encoded into the bitstream.
[0169] When applying the primary transform, the information indicating the type of kernel applied to the primary transform may be encoded into the bitstream.
[0170] Next, the secondary transform may be performed by considering the conditions of the secondary transform. Examples of the conditions of the secondary transform are as follows. The secondary transform may be performed for at least one of the following conditions.
[0171] 1) When the primary transform kernel is DCT 2, DCT 2 (vertical kernel, horizontal kernel)
[0172] 2) When there are at least 4 non-zero transform coefficients after the primary transform
[0173] 3) When the block size is 4×4 to 16×16 or smaller
[0174] For any of these conditions, a flag indicating whether the secondary transform is applied may be encoded.
[0175] The secondary transform may be performed by using the primary transform block as the input. In addition to kernels such as DST and DCT, the KLT kernel may be applied to the secondary transform. The KLT kernel may be selectively applied. Whether to apply the KLT kernel may be determined according to the characteristics of the residual block. Here, the characteristics of the residual block may include the width, height, size, shape, partition depth, etc. of the residual block. In contrast, the information indicating whether the KLT kernel is selectively applied may be encoded into the bitstream.
[0176] When applying the secondary transform, the information indicating the type of kernel applied to the secondary transform may be encoded into the bitstream.
[0177] Examples of encoding the primary transform kernel or the secondary transform kernel are as follows.
[0178] 1) When the intra prediction direction mode is DC or planar, it is encoded using kernel index 1.
[0179] 2) When it is the horizontal mode ±5, it is encoded using kernel index 2.
[0180] 3) When it is the vertical mode ±5, it is encoded using kernel index 3.
[0181] 4) When it is in diagonal mode 0, it is encoded using kernel index 4.
[0182] Four index coding examples are shown, but a part of them may be omitted.
[0183] The vectors of the secondary transform can arrange the frequency information with high energy in the form of blocks in the diagonal scan direction starting from the two-dimensional coordinates (0, 0). Additionally, the blocks can be the input values under the quantization step.
[0184] Figure 8 It is a diagram showing the inverse transform process in the decoder.
[0185] First, the quantized blocks obtained from the bitstream can be dequantized to obtain secondary transform blocks. It can be determined whether to perform the secondary inverse transform on the secondary transform blocks. When it is determined to perform the secondary inverse transform, the secondary inverse transform can be performed on the secondary transform blocks to obtain primary transform blocks. It can include performing the primary inverse transform on the primary transform blocks.
[0186] The dequantization step can be executed, and the coefficients can be rearranged by diagonal scanning. When the secondary inverse transform is applied, it can be rearranged in the form of vectors.
[0187] Next, it can be determined whether to perform the secondary inverse transform based on the information signaled from the bitstream regarding whether to apply the secondary inverse transform. In contrast, the secondary inverse transform can also be performed by considering the secondary inverse transform conditions. The conditions can be the same as those considered in the encoding step.
[0188] When it is determined to apply the secondary inverse transform, the kernel to be applied to the secondary inverse transform can be determined by the secondary transform kernel information signaled from the bitstream. The secondary inverse transform can be performed based on the determined secondary inverse transform kernel.
[0189] Next, the primary transform kernel information for the primary inverse transform can be signaled from the bitstream. The transform kernel to be applied to the primary inverse transform can be determined based on the primary transform kernel information. The primary inverse transform can be performed using the determined transform kernel. The residual block can be obtained through the primary inverse transform.
[0190] In the step for applying the transform kernel, the input and output can be defined as follows.
[0191] For the primary transform, the input and output are defined
[0192] - Input: Residual block of size M×N (M: height, N: width)
[0193] - Output: Primary transform block that performs the KLT, from which the basis (basis) with low energy (corresponding to small λ values) has been removed (the size changes according to dimensionality reduction)
[0194] For the primary inverse transform, define the input and output
[0195] - Input: Primary transform block (size varies according to dimensionality reduction)
[0196] - Output: Residual block of size M×N
[0197] For the secondary transform, define the input and output
[0198] - Input: Primary transform block of size P×Q (P≤M, Q≤N)
[0199] - Output: Secondary transform block on which KLT is performed, from which the basis (corresponding to small λ values) with low energy is removed (size varies according to dimensionality reduction)
[0200] For the secondary inverse transform, define the input and output
[0201] - Input: Secondary transform block
[0202] - Output: Primary transform block of size P×Q
[0203] Figure 9 is a diagram showing a first embodiment of applying a 1D transform kernel by using dimensionality reduction.
[0204] Figure 9 The first embodiment in is an example of applying a 1D transform kernel (separable KLT), and can be an example of applying the 1D transform kernel to the primary transform or the secondary transform. Since dimensionality reduction is performed in the embodiment of Figure 9 , when the input block is M×N, the size of the output block can be M / 2×N / 2, which is different from Figure 5 (for Figure 5 , the output block is M×N). Specifically, it shows an example of using a vertical KL transform kernel and a horizontal KL transform kernel, and shows a dimensionality reduction of 1 / 2 in the vertical direction and the horizontal direction respectively.
[0205] Figure 10 is a diagram showing a second embodiment of applying a 1D transform kernel by using dimensionality reduction.
[0206] Figure 10 The second embodiment in is an example of applying a 1D transform kernel (separable KLT), and the 1D transform kernel can be applied to the primary transform or the secondary transform. Since dimensionality reduction is performed in the embodiment of Figure 10 , when the input block is M×N, the size of the output block can be M / 4×N / 2. Specifically, a vertical KL transform kernel and a horizontal KL transform kernel are used, and a dimensionality reduction of 1 / 4 can be shown in the vertical direction and a dimensionality reduction of 1 / 2 can be shown in the horizontal direction.
[0207] Figure 11 FIG. is a diagram showing a third embodiment of applying a 1D transform kernel by using dimensionality reduction.
[0208] Figure 11 The third embodiment in is an example of applying a 1D transform kernel (separable KLT), and the 1D kernel can be applied to the primary transform or the secondary transform. Since dimensionality reduction is performed in the Figure 11 embodiment, when the input block is M×N, the size of the output block can be M / 4×N / 4. Specifically, a vertical KL transform kernel and a horizontal KL transform kernel are used, and a dimensionality reduction of 1 / 4 can be shown in the vertical and horizontal directions respectively.
[0209] Figure 12 FIG. is a diagram showing a fourth embodiment of applying a 1D transform kernel when transforming all data without performing dimensionality reduction.
[0210] Figure 12 The fourth embodiment in is an example of applying a 1D transform kernel (separable KLT), and the 1D kernel can be applied to the primary transform or the secondary transform. Since all data is transformed without performing dimensionality reduction in the Figure 11 embodiment, both the size of the input block and the size of the output block can be M×N. Specifically, it can be shown that a vertical KL transform kernel and a horizontal KL transform kernel are used without using dimensionality reduction.
[0211] Figure 13 FIG. is a diagram showing a first embodiment of applying a 2D transform kernel by using dimensionality reduction.
[0212] Since dimensionality reduction is performed in the Figure 13 first embodiment of applying a 2D transform kernel, the output vector can be MN / 2 instead of MN. Specifically, a 2D KL transform kernel is used, and a dimensionality reduction of 1 / 2 can be shown.
[0213] Figure 14 FIG. is a diagram showing a second embodiment of applying a 2D transform kernel by using dimensionality reduction.
[0214] Since dimensionality reduction is performed in the second embodiment of applying a 2D transform kernel, when the input block is M×N, the output vector can be MN / 4 instead of MN. Specifically, a 2D KL transform kernel is used, and a dimensionality reduction of 1 / 4 can be shown.
[0215] Figure 15 FIG. is a diagram showing a second embodiment of applying a 2D transform kernel by using dimensionality reduction.
[0216] Since dimensionality reduction is performed in the second embodiment of applying the 2D transform kernel, when the input block is M×N, the output vector can be MN / 4 instead of MN. Specifically, a 2D KL transform kernel is used, and a dimensionality reduction of 1 / 4 can be shown.
[0217] Figure 16 FIG. is a diagram showing a third embodiment of applying a 2D transform kernel by utilizing dimensionality reduction.
[0218] Since dimensionality reduction is not performed in the third embodiment of applying the 2D transform kernel, the output vector can be MN. Specifically, it can be expressed as applying a 2D transform kernel of MN×MN size.
[0219] Whether to utilize dimensionality reduction can be determined by the characteristics of the block described above or the information signaled through the bitstream. This can be signaled in the higher unit of the block as well as in the block unit.
[0220] Figure 17 An example of rearranging the vector into a 2D block is shown.
[0221] Specifically, it shows an example where when the MN×1 vector derived in the step of applying the transform kernel is constructed into a 2D M×N transform block, they are rearranged in 2D in the order of high energy.
[0222] The rearrangement can include bottom-up diagonal arrangement, horizontal arrangement, zigzag arrangement, vertical arrangement, etc. According to each inclusion method, the coefficients in the vector can be rearranged into the 2D block in the order of high energy. This can be equally applied to primary or secondary transforms and can also be equally applied to SVD.
[0223] Figure 18 An embodiment of scanning the coefficients of the rearranged block is shown.
[0224] The scanning order of the quantization coefficients can vary according to the rearrangement method in the step of applying the transform kernel. The scanning order can include the scanning order for bottom-up diagonal arrangement, the scanning order for horizontal arrangement, the scanning order for zigzag arrangement, and the scanning order for vertical arrangement. Information about the determined scanning order or information about the rearrangement can be encoded / decoded through the bitstream. When the scanning order is determined by the information about the rearrangement, the rearrangement method and the scanning order can have a 1:1 relationship.
[0225] Figure 19 FIG. is a diagram showing an embodiment of signaling a 1D transform kernel for primary transform.
[0226] First, H-KLT and V-KLT may respectively refer to the horizontal KL / SVD transform kernel and the vertical KL / SVD transform kernel derived in the present invention. For the primary transform, the transform kernel is signaled through mts_idx.
[0227] When the proposed 1D transform kernel is applied to the primary transform, mts_idx may be defined and signaled as shown in Figure 19 A, shown in Figure 19 B, and shown in Figure 19 C.
[0228] The maximum value of mts_idx may increase according to the number of kernels to be derived and applied, or it may be signaled by replacing the DCT and DST kernels. Examples such as DCT-2 and V-KLT are also feasible.
[0229] Figure 20 is a diagram showing an example of signaling the 2D transform kernel for the primary transform.
[0230] When the proposed 2D transform kernel is applied to the primary transform, mts_idx may be defined and signaled as shown in Figure 20 A, shown in Figure 20 B, and shown in Figure 20 C.
[0231] The maximum value of mts_idx may increase according to the number of kernels to be derived and applied, or it may be signaled by replacing the DCT and DST kernels. KLT may refer to the KL / SVD transform kernel derived in the present invention. As an example, for Figure 20 A, the maximum value of mts_idx may be 4, but for Figure 20 B, the maximum value of mts_idx may be 5.
[0232] Figure 21 is a diagram showing an example of signaling the 1D transform kernel for the secondary transform.
[0233] For the secondary transform, the transform kernel may be signaled through secondary_idx.
[0234] When the proposed 1D transform kernel is applied to the secondary transform, secondary_idx may be defined and signaled as shown in Figure 21 A, shown in Figure 21 B, and shown in Figure 21 C.
[0235] The maximum value of secondary_idx can be increased according to the number of kernels to be derived and applied, or can be signaled by replacing the secondary kernels. H-KLT and V-KLT may respectively refer to the horizontal KL / SVD transform kernels and vertical KL / SVD transform kernels derived in the present invention.
[0236] Figure 22 It is a diagram showing an example of signaling 2D transform kernels for secondary transform.
[0237] When the proposed 2D transform kernels are applied to secondary transform, secondary_idx can be defined and signaled as shown in Figure 22 A, shown in Figure 22 B, and shown in Figure 22 C.
[0238] The maximum value of secondary_idx can be increased according to the number of kernels to be derived and applied, or can be signaled by replacing the secondary kernels. KLT may refer to the KL transform / SVD transform kernels derived in the present invention.
[0239] Figure 23 An example of only applying primary transform in the encoder is shown.
[0240] The transform kernels for primary transform of the residual block can be determined. Based on the determined transform kernels for primary transform, the primary transform can be performed on the residual block to obtain a primary transform block. The quantized block obtained by quantizing the primary transform block can be encoded into the bitstream.
[0241] Applying the processing of the transform encoder according to Figure 23 can be a processing that omits the secondary transform step differently from Figure 2 In other words, it can be a processing that only performs primary transform. It can correspond to the case where the condition of whether to perform secondary transform described above is no, but can also include the case of omission regardless of the condition of whether to perform secondary transform.
[0242] Whether to perform the transform processing in Figure 23 can be determined based on at least one of the prediction mode and the size of the block (or the product of the width and height of the block).
[0243] As an example, the transform processing in Figure 23 can be performed only in the intra prediction mode. In contrast, the transform processing in Figure 23 can be performed only in the inter prediction mode.
[0244] As another example, the transform processing in Figure 23The transform processing in. As another example, it can be applied only when the block size is 4×4, 4×8, 8×4, 8×8, 4×16, and 16×4 Figure 23 The transform processing in.
[0245] As another example, it can be applied only when the product of the width and height of the block is less than 64 Figure 23 The transform processing in.
[0246] As another example, it can be applied only when the prediction mode of the current block is the intra prediction mode and the block size is 4×4, 4×8, 8×4, 8×8, 4×16, and 16×4 Figure 23 The transform processing in.
[0247] As another example, it can be applied only when the prediction mode of the current block is the intra prediction mode and the block size is 4×4, 4×8, 8×4, and 8×8 Figure 23 The transform processing in.
[0248] As another example, it can be applied only when the prediction mode of the current block is the intra prediction mode and the product of the width and height of the block is less than 64 Figure 23 The transform processing in.
[0249] In Figure 23 the processing of, it can be performed in the same manner as Figure 2 in the way and the same steps as Figure 2 the steps in.
[0250] As an example, as Figure 2 shown, it can be signaled as follows Figure 23 the steps for signaling the primary transform kernel in.
[0251] 1) Kernel 1 when the intra prediction direction mode is DC or planar
[0252] 2) Kernel 2 when it is the horizontal mode ±5
[0253] 3) Kernel 3 when it is the vertical mode ±5
[0254] 4) Kernel 4 when it is the diagonal mode 0
[0255] Figure 23 The scanning in may be different from the scanning order for sending quantization coefficients according to the arrangement method in the step of applying the transform kernel. (Even when scanning is performed from the lowest energy, the results may be similar)
[0256] As Figure 23 an example of the scanning in, Figure 17 may represent the scanning order for the bottom-up diagonal arrangement, Figure 18 may represent the scanning order for the horizontal arrangement,Figure 19 can represent a scanning order for vertical arrangement, and Figure 20 can represent a scanning order for zigzag arrangement.
[0257] Figure 24 An example of only applying a primary inverse transform in a decoder is shown.
[0258] The quantized block obtained from the bitstream can be dequantized to obtain a primary transform block. The transform kernel of the primary inverse transform for the primary transform block can be determined. Based on the determined transform kernel of the primary inverse transform, the primary inverse transform can be performed on the primary transform block to obtain a residual block.
[0259] Applying the processing of the transform decoder according to Figure 24 can be a processing that omits the secondary inverse transform step differently from Figure 3 In other words, it can be a processing that only performs the primary inverse transform. It can correspond to the case where the condition of whether to perform the secondary transform described above is no, but can also include the case of omission regardless of the condition of whether to perform the secondary transform.
[0260] Whether to perform the inverse transform processing in Figure 24 can be determined based on at least one of the prediction mode and the size of the block (or the product of the width and height of the block).
[0261] As an example, the inverse transform processing in Figure 24 can be performed only in the intra prediction mode. In contrast, the inverse transform processing in Figure 24 can be performed only in the inter prediction mode.
[0262] As another example, the inverse transform processing in Figure 24 can be applied only when the block size is 4×4, 4×8, 8×4, and 8×8. As another example, the inverse transform processing in Figure 24 can be applied only when the block size is 4×4, 4×8, 8×4, 8×8, 4×16, and 16×4.
[0263] As another example, the inverse transform processing in Figure 24 can be applied only when the product of the width and height of the block is less than 64.
[0264] As another example, the inverse transform processing in Figure 24 can be applied only when the prediction mode of the current block is the intra prediction mode and the block size is 4×4, 4×8, 8×4, 8×8, 4×16, and 16×4.
[0265] As another example, the inverse transform processing in Figure 24 can be applied only when the prediction mode of the current block is the intra prediction mode and the block size is 4×4, 4×8, 8×4, and 8×8.
[0266] As another example, the inverse transform process in Figure 24 may be applied only when the prediction mode of the current block is an intra prediction mode and the product of the width and height of the block is less than 64. Figure 24
[0267] In Figure 24 's processing, the same steps as those in Figure 3 can be executed in the same manner as the way in Figure 3
[0268] Figure 24 The scanning in
[0269] can be different in the scanning order for sending quantization coefficients according to the arrangement method in the step of applying the transform kernel. (Even when scanning is performed from the lowest energy, the results may be similar) Figure 24 As an example of the scanning in Figure 17 can represent the scanning order for a bottom - up diagonal arrangement, Figure 18 can represent the scanning order for a horizontal arrangement, Figure 19 can represent the scanning order for a vertical arrangement, and Figure 20 can represent the scanning order for a zig - zag arrangement.
[0270] Figure 25 is a diagram for describing an example of using only some of the MN ( = M×N) columns. Figure 26 is a diagram showing the signal - related transform kernel size and its corresponding inverse kernel.
[0271] First, referring to Figure 25 , φ can be obtained through the same processing as the processing in Equation 2 above MN×MN to have MN columns. Each column is expressed as λ, and λ1, λ2,... λ MN columns can be generated. Only some columns can be used as the transform kernel. This takes into account the energy characteristics of the signal.
[0272] As a reference, the signal - related transform (SDT) of all data without considering energy can be expressed as Equation 7 and Equation 8.
[0273] [Equation 7]
[0274]
[0275] Forward transform
[0276] [Equation 8]
[0277] φ MN×MN ×SDT (MN×1) = T MN×1
[0278] Inverse transform
[0279] In contrast, a signal correlation transform can be performed by reducing the size of the transform kernel by considering only a part of the overall energies λ1, λ2, ... λ MN of the transform kernel.
[0280] Referring to Figure 26 , an example of reducing the size of the transform kernel by considering only a part of the energy can be confirmed. Additionally, the size of the transform kernel can be further reduced compared to the Figure 26 embodiment in
[0281] As an example, since MN is 128 when M is 16 and N is 8, the size of the kernel considering the overall energy is 128×128 (lossless coding). However, when the is configured as 32×128 to select λ1, λ2, …, λ that contribute to high energy 32 and perform a transform on them, 128 data are represented as 32 data while the energy is retained to some extent, so the compression performance can be improved.
[0282] As another example, since MN is 128 when N is 8, the size of the kernel considering the overall energy is 128×128 (lossless coding). However, when the is configured as 16×128 to select λ1, λ2, …, λ that contribute to high energy 16 and perform a transform on them, 128 data are represented as 16 data while the energy is retained to some extent, so the compression performance can be improved.
[0283] To encode the quantized transform coefficients (quantizing the transform data), the following items can be encoded / decoded: a flag indicating whether there is a non-zero coefficient defined in units of a coefficient group, a flag indicating whether a coefficient defined in units of a coefficient is non-zero, a flag indicating whether the absolute value of a coefficient defined in units of a coefficient is greater than a specific value, information about the remaining absolute value of a coefficient defined in units of a coefficient, etc.
[0284] Figure 27 is a view showing the scanning of one-dimensional data.
[0285] As in Figure 27 , when applying a 2D non-separable KLT or SVD, the transform result of a two-dimensional block can be one-dimensional data with reduced data dimension (except for 4×4). Therefore, entropy coding (CABAC, VLC, etc.) can be performed while scanning the one-dimensional data as it is after quantization as shown in Figure 27 .
[0286] When constructing a signal-dependent transform, for intra prediction, three to four transform kernels can be trained and constructed according to the intra mode information. In this case, the index of the specific transform kernel used for the transform must also be sent to the decoder. Similarly, for inter coding, three to four transform kernels can be trained and constructed for each residual block size after AMVP (Advanced Motion Vector Prediction) or MV (Motion Vector) merge. In this case, the index of the specific transform kernel used for the transform must also be sent to the decoder.
[0287] In addition, the signal-dependent transform can be pre-trained and applied by using the sub-block partitioning transform in the inter prediction of the existing signal-independent transform and the residual signal of the intra sub-partitioning.
[0288] The coefficient group can be used to encode / decide whether there are coefficients through the flag of the bitstream.
[0289] When applying a signal-adaptive transform, the size of the transform kernel is reduced. Therefore, for lossless coding of video, the transform step must be skipped and compression must be performed directly.
[0290] After the signal-adaptive transform (primary transform), a signal-adaptive transform (secondary transform) can be performed again to further compress the data. In this case, the proposed method is applied by arranging the one-dimensional data of the primary transform coefficients two-dimensionally (in horizontal, vertical, diagonal, or zigzag order) before applying the secondary transform. After that, the transform result of the two-dimensional block after the secondary transform can be one-dimensional data with a reduced number of data (dimensions) (except for 4×4). Thus, after quantization, entropy coding (CABAC, VLC, etc.) is performed while scanning the one-dimensional data as it is.
[0291] The interpolation filter using the frequency of the present disclosure can be applied to at least one of the encoding / decoding steps. As an example, the interpolation filter can be used to interpolate reference samples, can be used to adjust prediction values, can be used to adjust residual values, can be used to improve encoding / decoding efficiency after prediction is completed, and can be performed as an encoding / decoding preprocessing step.
[0292] By replacing the previously used 4-tap discrete cosine transform-based interpolation filter (DCT-IF) and 4-tap smoothing interpolation filter (SIF) in VVC intra prediction, an 8-tap DCT-IF and an 8-tap SIF using more reference samples are applied. In some cases, filters such as 10-tap, 12-tap, 14-tap, 16-tap, etc. can be used instead of the 8-tap filter. As an example, the characteristics of a block can be determined by using the size of the block and the frequency characteristics of the reference samples, and the type of interpolation filter applied to the block can be selected.
[0293] The 8-tap DCT-IF can be obtained by Equation 9 below.
[0294] [Equation 9]
[0295]
[0296] The p / 32 pixel interpolation filter (when using p = 0, 1, 2, 3, …, 31, 1 / 32 fractional samples) can be obtained by replacing n = 3 + p / 32 with a linear combination of discrete cosine coefficients and x(m) (m = 0, 1, 2, …, 7) in Equation 9 above.
[0297] As an example, the 8-tap DCT-IF coefficients derived for the (0 / 32, 1 / 32, 2 / 32, …, 16 / 32) fractional sample positions can be obtained by using n = 3+(0 / 32, 1 / 32, 2 / 32, …, 16 / 32). The 8-tap DCT-IF coefficients for (17 / 32, 18 / 32, 19 / 32, …, 31 / 32) can also be obtained in the same manner as described above.
[0298] The 8-tap SIF coefficients can be obtained from the convolution of z[n] and the 1 / 32 fractional linear filter. Here, Figure 28 z[n] in can be obtained from the convolution of h[n] and y[n] in Equation 10 and Equation 11. Here, h[n] can be a 3-tap [1, 2, 1] LPF (low-pass filter).
[0299] Equation 10 and Equation 11 show the process for deriving y[n] and z[n]. Figure 28 Denote h[n], y[n] and z[n], and the 8-tap SIF coefficients can be obtained by linear interpolation of z[n] and the 1 / 32 fractional linear filter.
[0300] [Equation 10]
[0301]
[0302] [Equation 11]
[0303]
[0304] The following Equation 12 and Equation 13 describe the process for calculating the 8-tap SIF coefficients. Here, it can be g[n] = z[n - 3], n = 0, 1, 2, …, 6.
[0305] [Equation 12]
[0306]
[0307] Equation 13 shows an example of deriving the SIF coefficients at the position of i0 + 3 + 16 / 32 pixel when p is 16.
[0308] [Equation 13]
[0309]
[0310] Figure 29 Shows the integer reference samples for deriving the 8-tap SIF coefficients. Refer to Figure 29 , eight integer samples r[i0] to r[i0 + 7] for deriving the filter coefficients at the black i0 + 3 + 16 / 32 pixel position are indicated in gray. Additionally, r[i0] can be the starting sample of the eight reference samples. The filter coefficients can be adjusted for integer implementation.
[0311] Since the 8-tap DCT-IF has higher frequency characteristics than the 4-tap DCT-IF, and the 8-tap SIF has lower frequency characteristics than the 4-tap SIF, the 8-tap interpolation filter type can be selected and used according to the characteristics of the block.
[0312] The characteristics of the block can be determined by using the size of the block and the frequency characteristics of the reference samples, and the type of interpolation filter for the corresponding block can be selected.
[0313] The following characteristics can be used: as the size of the block becomes smaller, the correlation becomes lower and more frequencies become higher; as the size of the block becomes larger, the correlation becomes higher and more frequencies become lower.
[0314] [Obtaining Correlation]
[0315] To determine the reference sample characteristics based on the CU size, the correlation in Equation 14 is calculated from the top or left reference samples of the current CU according to the intra prediction mode. Here, N can be the width or height of the current CU.
[0316] [Equation 14]
[0317]
[0318] Figure 30 Is a diagram showing the direction and angle of the intra prediction mode. When the prediction mode of the current CU is greater than Figure 30 the diagonal mode 34 in, the reference samples at the top position of the current CU can be used in Equation 14. Otherwise, the reference samples at the left position of the current CU can be used in Equation 14.
[0319] Figure 31 Is a diagram showing the average correlation values of the reference samples for various video resolutions and each nTbS. Specifically, Figure 31 shows the average correlation values of the reference samples for each nTbS and various video resolutions defined in Equation 15, which can be determined according to the CU size at each screen resolution. As Figure 31As shown, the correlation can increase with the increase in CU size and video resolution. Here, the video resolutions A1, A2, B, C, and D can be shown in parentheses.
[0320] [Equation 15]
[0321] nTbS = ((Log2(W) + Log2(H)) >> 1
[0322] The intra-CU size partitioning in video coding can depend on the prediction performance for improving coding in terms of bitrate and distortion. The prediction performance can vary according to the prediction error between the predicted sample and the samples of the current CU. When the current block has many detailed regions including high frequencies, the CU size can be partitioned into small parts by considering bitrate and distortion using boundary reference samples with small width and height. However, when the current block consists of uniform regions, the CU size can be partitioned into large parts by considering bitrate and distortion using boundary reference samples with large width and height.
[0323] Referring to Figure 31 , the correlation values of the reference samples indicated by A1, A2, B, C, and D depending on the video resolution and nTbS size can be known. This can respectively mean that small nTbS has high-frequency characteristics consistent with low correlation, and large nTbS has low-frequency characteristics consistent with high correlation.
[0324] The frequency characteristics of the reference samples can be obtained by applying a transform to the reference samples of the block using DCT-II. According to the intra prediction mode, the reference samples for obtaining the frequency characteristics of the reference samples can be determined. As an example, when the direction of the intra prediction mode is vertical, the top reference sample of the current coding block (or sub-block) is used, and when the direction of the intra prediction mode is horizontal, the left reference sample of the current coding block (or sub-block) is used. When the direction of the intra prediction mode is diagonal, at least one of the left and top reference samples of the current coding block (or sub-block) is used. Here, the reference samples can be adjacent to the current coding block (or sub-block), or can be k pixels away from the current coding block (or sub-block). Here, k can be a natural number such as 1, 2, 3, 4, etc.
[0325] As the high-frequency energy percentage is higher, the block can have high-frequency characteristics. The frequency characteristics of the block can be determined by comparing the high-frequency energy percentage with a threshold depending on the block size, and the interpolation filter to be applied to the block can be selected.
[0326] According to the frequency information, the 8-tap DCT-IF as a strong high-pass filter (HPF) can be applied to the blocks with more high frequencies, and the 8-tap SIF as a strong low-pass filter (LPF) can be applied to the blocks with more low frequencies.
[0327] When the block size is small, an 8-tap DCT-IF, which is a strong HPF, can be applied by using a method for applying a strong HPF to a block with more high frequencies according to frequency information and the characteristic of lower correlation as the block size is smaller. When there are fewer high frequencies, a 4-tap SIF, which is a weak LPF, can be applied.
[0328] When the block size is large, an 8-tap SIF, which is a strong LPF, can be applied by using a method for applying a strong LPF to a block with more low frequencies according to frequency information and the characteristic of higher correlation as the block size is larger. When there are many high frequencies, a 4-tap DCT-IF, which is a weak HPF, can be applied.
[0329] An example for obtaining the high-frequency energy percentage can be as follows.
[0330] When the intra prediction mode is horizontal, N is the height of the block, and when the intra prediction mode is vertical, N is the width of the block. When fewer or more reference samples are used, the value of N can be smaller or larger. X can refer to the reference samples. In this case, the high-frequency domain uses reference samples with a length of 1 / 4 of N, and when fewer reference samples are used to obtain the high-frequency energy or when more reference samples are used, the length of this domain can be decreased or increased. Equation 16 can represent an example for obtaining the high-frequency energy percentage.
[0331] [Equation 16]
[0332]
[0333] Figure 32 An example of a method for selecting an interpolation filter by using frequency information is shown.
[0334] When the high-frequency energy percentage of a block with TbS of 2 is less than the threshold, a 4-tap SIF is applied, otherwise, an 8-tap DCT-IF is applied.
[0335] When the high-frequency energy percentage of a block with nTbS of 5 or greater is less than the threshold, an 8-tap SIF is applied, otherwise, a 4-tap DCT-IF is applied.
[0336] Instead of an 8-tap filter, filters such as 10-tap, 12-tap, 14-tap, 16-tap, etc. can be used.
[0337] In addition, in Figure 32 filters for each of nTbS and high_freq_ratio are described, but multiple filters can also be used.
[0338] As an example, when nTbS is 2, a 4-tap SIF and an 8-tap SIF can be used instead of the 4-tap SIF. As an example, an 8-tap DCT-IF and a 16-tap DCT-IF can be used instead of the 8-tap DCT-IF.
[0339] When using multiple filters in this way, the encoding device can encode information (index) specifying any one of them, and the decoding device can signal information through the bitstream to specify any one of the multiple filters.
[0340] Optionally, when using multiple filters, any one of the multiple filters can be implicitly specified by the intra prediction mode.
[0341] In addition, when determining the filter type, in addition to nTbS and high_freq_ratio, the shape of the block being square or rectangular can also be considered additionally.
[0342] Figure 33 An embodiment of the 8-tap DCT interpolation filter coefficients is shown. Figure 34 An embodiment of the 8-tap smoothing interpolation filter coefficients is shown. Here, index = 0 and index = 7 can correspond to 8 integer samples r[i0] to r[i0 + 7] used to derive the filter coefficients.
[0343] Figure 35 The magnitude responses of the 4-tap DCT-IF, 4-tap SIF, 8-tap DCT-IF, and 8-tap SIF at the 16 / 32 pixel positions are shown. Here, the X-axis can represent the normalized angular frequency, and the Y-axis can represent the magnitude response.
[0344] The 8-tap DCT-IF has better HPF characteristics than the 4-tap DCT-IF, and the 8-tap SIF has better LPF characteristics than the 4-tap SIF. Therefore, the 8-tap SIF provides better interpolation in the low-frequency reference samples than the 4-tap SIF, and the 8-tap DCT-IF provides better interpolation in the high-frequency reference samples than the 4-tap DCT-IF.
[0345] VVC uses two interpolation filters. When nTbS = 2, the 4-tap DCT-IF is used for all blocks. When nTbS = 3, 4, the 4-tap DCT-IF or 4-tap SIF is used based on minDistVerHor and intraHorVerDistThres[nTbS], and when nTbS ≥ 5, the 4-tap SIF is used for all blocks.
[0346] In addition to the proposed interpolation filters, the present disclosure provides a method for selecting an interpolation filter for generating accurate fractional boundary prediction samples by using the frequency information of integer reference samples.
[0347] Although the CU reference samples have low-frequency characteristics, DCT-IF is used for the CU reference samples with nTbS = 2 in the VVC standard. However, as Figure 35 shown, regardless of the nTbS size, using SIF according to the low-frequency characteristics of the reference samples is more effective than DCT-IF. Similarly, although the CU reference samples have high-frequency characteristics, SIF is used for the CUs with nTbS > 4 in the VVC standard. However, regardless of the nTbS size, using Figure 35 the DCT-IF in is more effective than SIF according to the high-frequency characteristics of the reference samples. To solve this problem, a method has been developed to select two different filters composed of SIF and DCT-IF according to the frequency characteristics of the reference samples. The reference samples can be transformed by using a scaled integer one-dimensional (1-D) DCT-II kernel to detect the high-frequency energy of the reference samples. The scaled DCT-II coefficients X[k] (k = 0, 1, 2, …, N - 1) are derived from Equations 17 and 18 as follows.
[0348] [Equation 17]
[0349]
[0350] [Equation 18]
[0351] M = log2(N), shift = log2(N) + 1
[0352]
[0353] Here, N is the number of reference samples required for X[k]. After the one-dimensional transformation, high-frequency energy is observed in the transformed region. Since the energy is concentrated on the low-frequency components, the reference samples consist of uniform samples. However, since the energy exists in the high-frequency components, the reference samples include high-frequency samples, indicating that the samples of the CU have high-frequency components. The transformation size of the reference samples can be set according to the intra prediction mode of the current block. When the intra prediction mode is greater than mode 34 (diagonal mode), the top CU reference samples can be transformed into N = CU width in Equations 17 and 18. Moreover, when the intra prediction mode is less than mode 34, the left reference samples can be used as N = CU height in Equations 17 and 18. X[k] can be used to measure the energy ratio of the high-frequency coefficients. When the high-frequency data has energy, the reference samples consist of high-frequency data, so DCT-IF can be used. In contrast, SIF can be used for the high-energy reference samples of low-frequency data. high_freq_ratio (the percentage of energy of the high-frequency coefficients) can be calculated in Equation 19.
[0354] [Equation 19]
[0355]
[0356] In Figure 32 , the threshold (THR) of high_freq_ratio can be determined through experiments.
[0357] Figure 36 Shows each threshold THR1, THR2,... depending on nTbS, and Figure 36 (a) of Figure 36 (b) of Figure 36 (a) of Figure 36 are the results of nTbS = 2 and nTbS = 5 respectively. In
[0358] In Figure 36 (a) of Figure 36 (a) of Figure 35 Figure 35 Figure 35 Figure 35
[0359] Similarly, when the nTbS size of the CU is greater than 4, if high_freq_ratio < THR, an 8-tap SIF with strong LPF characteristics can be applied to the reference sample as shown in Figure 35 Figure 35 Figure 36As shown, when nTbS > 4, the threshold can be relatively higher compared to when nTbS = 2. Otherwise, if high_freq_ratio ≥ THR, then a 4-tap DCT-IF with weak HPF characteristics can be applied to the reference sample as shown in Figure 35 shown.
[0360] Figure 37 The sequence name, screen size, screen frame rate, and bit depth of the CTC video sequences for each category are shown.
[0361] The proposed method was implemented in the VVC reference software VTM-14.2
[37] and tested in the all intra (AI) configuration under the JVET common test conditions (CTC)
[38] . Each sequence of categories A1, A2, B, C, and D was tested with quantization parameter (QP) values of 22, 27, 32, and 37, respectively.
[0362] Figure 38 The interpolation filter selection method and the interpolation filter applied according to the selected method are shown to test the efficiency of the 8-tap / 4-tap interpolation filter.
[0363] Method A uses an 8-tap DCT-IF for nTbS = 2 and a 4-tap SIF for nTbS > 4, and method B uses an 8-tap SIF for nTbS > 4 and a 4-tap DCT-IF for nTbS = 2, and the DCT-IF or SIF is selected in the same way as the VVC anchor. The difference between method A and the VVC method is that method A uses an 8-tap DCT-IF only for nTbS = 2 instead of a 4-tap DCT-IF. The difference between method B and the VVC method is that method B uses an 8-tap SIF only for nTbS > 4 instead of a 4-tap SIF.
[0364] Figure 39 Tables IX and X in show the simulation results of methods A, B, C, and D.
[0365] Method C uses an 8-tap DCT-IF or a 4-tap SIF according to high_freq_ratio in Equation 19 for nTbS = 2 and a 4-tap SIF for nTbS > 4. Method D uses an 8-tap SIF or a 4-tap DCT-IF. This depends on high_freq_ratio for nTbS > 4 and on a 4-tap DCT-IF for nTbS = 2. Here, the filter selection method and the interpolation filter are used for CUs with nTbS = 3 and nTbS = 4 under the VVC anchor.
[0366] For Method A, for the Y, Cb, and Cr components, the overall BD-rate increases observed are -0.13%, -0.12%, and -0.08% respectively, where the sign (-) indicates bit savings. For Method C, for the Y, Cb, and Cr components, the overall BD-rate increases observed are -0.14%, -0.09%, and -0.11% respectively.
[0367] Specifically, for Category C with a resolution of 832×480 and Category D with a resolution of 416×240, component gains (-0.40%, -0.30%) from Method A and Y, and Categories C and D are achieved by Method C. For Methods A and C, an 8-tap DCT-IF with nTbS = 2 and a SIF with nTbS>4 are applied to each CU, regardless of the filter selection method.
[0368] For Method B, all BD-rate gains for the Y, Cb, and Cr components are -0.02%, -0.03%, and -0.02% respectively. For Method D, all BD-rate gains for the Y, Cb, and Cr components are -0.01%, -0.01%, and -0.03% respectively.
[0369] Method B uses an 8-tap SIF for nTbS>4 and a 4-tap DCT-IF for nTbS = 2, and Method D uses an 8-tap SIF or a 4-tap DCT-IF according to the proposed high_freq_ratio for nTbS>4 and a 4-tap DCT-IF for nTbS = 2.
[0370] In Methods B and D, there is little overall BD-rate increase, but in Methods B and D, BD-rate increases (-0.08%, -0.09%) are obtained from the Y component of Class A1 with a resolution of 3840×2160 respectively. The proposed frequency-based adaptive interpolation filter using high_freq_ratio and nTbS and the existing VVC methods have been developed to utilize Methods C and D.
[0371] Figure 40 Table XI in shows the ratio of CUs that apply a 4-tap DCT-IF at the VVC anchor point and an 8-tap DCT-IF based on high_freq_ratio among the proposed methods for all test sequences.
[0372] For nTbS = 2, in the proposed high_freq_ratio-based adaptive filtering method, 4-tap DCT-IF was selected 100% in 4×4 CUs, 4×8 CUs, and 8×4 CUs of VVC anchors, while 8-tap DCT-IF was selected 97.16% in 4×4 CUs, 95.80% in 4×8 CUs, and 96.77% in 8×4 CUs. The percentage of 4-tap SIF selection for nTbS = 2 can be inferred from the DCT-IF selection percentages in Table XI.
[0373] Figure 41 The experimental results of the proposed filtering method are shown.
[0374] The increased percentage selections under 4-tap SIF and 8-tap DCT-IF brought BD-rate gains from Table XII. When using 4-tap SIF with low LPF characteristics and 8-tap DCT-IF with strong HPF characteristics according to the high_freq_ratio of small CUs, it helps to increase the BD-rate in the proposed method. Moreover, when using 8-tap SIF with strong LPF characteristics and 4-tap DCT-IF with poor HPF characteristics according to the high_freq_ratio of large CUs, it helps to slightly increase the BD-rate in the proposed method.
[0375] Except for nTbS > 4, 4-tap DCT-IF is only applied to CUs using MRL or ISP tools under VVC anchors.
[0376] However, in the proposed method, 8-tap DCT-IF and 4-tap SIF based on high_freq_ratio are applied to CUs using MRL or ISP. Thus, 8-tap DCT-IF was selected 0.07% in 32×32 CUs, 0.04% in 16×64 CUs, 0.07% in 64×16 CUs, and 0.07% in 64×64 CUs. This is compared with the cases of 10.59% selection in 32×32 CUs, 100% selection in 16×64 CUs, 100% selection in 64×16 CUs, and 5.56% selection in 64×64 CUs under VVC anchors.
[0377] Table XII shows the results of the proposed high_freq_ratio-based adaptive filtering method, and in the AIMain 10 configuration (Main 10 configuration within a full frame), EncT and DecT represent the total encoding and decoding time rates compared to the VVC anchors for various test sequences in classes A1 to D, respectively.
[0378] The proposed method can achieve an overall BD-rate increase of -0.16%, -0.13%, and -0.09% for the Y, Cb, and Cr components, respectively, and the computational complexity increases by 2% and 5% on average compared to the VVC anchors in the encoder and decoder, respectively.
[0379] With a slight increase in computational complexity, the proposed method can be used to reduce the BD-rate compared to the VVC anchor. The sequence showing the highest BD-rate reduction is the BasketballDrill sequence from class C, and the proposed method generates a gain of -1.20% for the Y component.
[0380] In summary, the present disclosure proposes an adaptive filtering method for generating partial reference samples for intra prediction of directional VVC frames. To improve the accuracy of fractional reference samples by using high_freq_ratio derived from 1D-scaled DCT, in addition to the 4-tap DCT-IF and 4-tap SIF, 8-tap DCT-IF and 8-tap SIF are also proposed. According to the high_freq_ratio for the block size, an interpolation filter is applied to the reference samples. It is concluded that when the correlation between samples is high, the 8-tap interpolation filter with a strong HPF or strong LPF has the least impact on the BD-rate gain, but when the correlation between samples is low, a strong 8-tap interpolation filter is used. The HPF or strong LPF characteristics affect the BD-rate improvement. For the proposed high_freq_ratio-based adaptive filtering method, overall BD-rate gains of -0.16%, -0.13%, and -0.09% are observed for the Y, Cb, and Cr components, respectively, compared to the VVC anchor. The method for searching for high-frequency terms in the frequency domain contributes to video coding modules that require strong / weak HPF and strong / weak LPF for the next-generation video coding standard.
[0381] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) and non-transitory computer-readable media, which cause the operations of the methods according to various embodiments to be executed on a device or computer. In the non-transitory computer-readable media, such software or instructions, etc., are stored on the device or computer and can be executed on the device or computer.
[0382] [Industrial Applicability]
[0383] The present invention can be used as a video encoding and decoding apparatus and method.
Claims
1. An image decoding method, comprising: Inverse quantizing a quantized block obtained from a bitstream to obtain a secondary transform block; Determining whether to perform a secondary inverse transform on the secondary transform block; When it is determined to perform the secondary inverse transform, performing the secondary inverse transform on the secondary transform block to obtain a primary transform block; and Performing a primary inverse transform on the primary transform block.
2. The method according to claim 1, wherein: The transform kernel of the secondary inverse transform and the transform kernel of the primary inverse transform are specified by an index signaled from the bitstream.
3. The method according to claim 1, wherein: The maximum value and configuration of the index differ according to whether the applied transform kernel is a one-dimensional transform kernel or a two-dimensional transform kernel.
4. The method according to claim 1, wherein: The transform kernel of the secondary inverse transform and the transform kernel of the primary inverse transform include the Karhunen-Loeve transform (KLT).
5. The method according to claim 1, wherein: The size of the output block of at least one of the secondary inverse transform and the primary inverse transform is smaller than the size of the input block.
6. An image encoding method, comprising: Performing a primary transform on a residual block to obtain a primary transform block; Determining whether to perform a secondary transform on the primary transform block; When it is determined to perform the secondary transform, performing the secondary transform to obtain a secondary transform block; and Encoding the quantized block obtained by quantizing the secondary transform block as a bitstream.
7. An image encoding method, wherein, Determining whether to perform the secondary transform based on at least one of the type of the transform kernel, the number of transform coefficients, and the size of the block.
8. A computer-readable recording medium storing a bitstream generated by an encoding method, wherein, The encoding method comprises: Performing a primary transform on a residual block to obtain a primary transform block; Determining whether to perform a secondary transform on the primary transform block; When it is determined to perform the secondary transform, performing the secondary transform to obtain a secondary transform block; and Encoding the quantized block obtained by quantizing the secondary transform block as the bitstream.
9. An image decoding method, comprising: Inverse quantizing a quantized block obtained from a bitstream to obtain a primary transform block; Determining the transform kernel for the primary inverse transform of the primary transform block; And Based on the determined transform kernel of the primary inverse transform, performing the primary inverse transform on the primary transform block to obtain a residual block.
10. An image encoding method, comprising: Determining the transform kernel for the primary transform of a residual block; Based on the determined transform kernel of the primary transform, performing the primary transform on the residual block to obtain a primary transform block; And Encoding the quantized block obtained by quantizing the primary transform block as a bitstream.
11. A computer-readable recording medium storing a bitstream generated by an encoding method, wherein, The encoding method comprises: Determining the transform kernel for the primary transform of a residual block; Based on the determined transform kernel of the primary transform, performing the primary transform on the residual block to obtain a primary transform block; and Encoding the quantized block obtained by quantizing the primary transform block as a bitstream.
12. An image decoding method, comprising: Determining the intra prediction mode of a current block; And Interpolating reference pixels for the intra prediction mode, wherein the reference pixels are included in adjacent reference blocks around the current block. Among them, the interpolation filter applied to the interpolation includes an 8-tap filter.
13. An image coding method, comprising: Determining an intra prediction mode of a current block; And Interpolating reference pixels for the intra prediction mode, wherein the reference pixels are included in adjacent reference blocks around the current block, and the interpolation filter applied to the interpolation includes an 8-tap filter.
14. A computer-readable recording medium storing a bitstream generated by an encoding method, the computer-readable recording medium comprising: Determining an intra prediction mode of a current block; And Interpolating reference pixels for the intra prediction mode, wherein the reference pixels are included in adjacent reference blocks around the current block, and the interpolation filter applied to the interpolation includes an 8-tap filter.