Method for encoding and decoding video signals and non-transitory readable storage medium
By configuring the secondary transformation set and hybrid transformation core, the storage and processing requirements for efficient processing of high-resolution video signals are solved, and the encoding efficiency and quality are improved.
Patent Information
- Application Number
- CN202310313452.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-02
- Filing Date
- 2019-07-02
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-07-02
AI Technical Summary
The prior art is difficult to efficiently handle the problem of the sharp increase in memory storage, memory access rate and processing capabilities caused by the high spatial resolution, high frame rate and high-dimensional scene presentation of next-generation video content, especially in terms of coding efficiency and complexity.
By configuring the secondary transformation set, using a hybrid secondary transformation set and intra prediction mode configuration, the transformation kernel is derived, and the inverse quantization blocks applied to the video signal are used for secondary transformation, including the mixed use of 8×8 and 4×4 transformation kernels, and efficient encoding and decoding is performed through binarized transformation indexes.
It improves the encoding efficiency of video signal processing, reduces storage and processing requirements, and improves the encoding quality and efficiency of video signals.
Smart Images

Figure CN116320414B_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application for invention with the original application number 201980045151.X (International Application Number: PCT / KR2019 / 008060, filing date: July 2, 2019, invention title: Method and apparatus for processing video signals based on secondary transformation). Technical Field
[0002] The present disclosure relates to a method and apparatus for processing video signals, and more particularly, to a method of designing and applying a secondary transformation. Background Art
[0003] Next-generation video content will have characteristics such as high spatial resolution, high frame rate, and high-dimensional scene rendering. To process the content, there will be a sharp increase in memory storage, memory access rate, and processing power.
[0004] Therefore, there is a need to design new coding tools for more efficiently processing next-generation video content. Specifically, when applying a transformation, there is a need to design a more efficient transformation in terms of coding efficiency and complexity. Summary of the Invention
[0005] Technical Problem
[0006] Embodiments of the present disclosure provide a method of configuring a secondary transformation set by considering various conditions for applying a secondary transformation.
[0007] In addition, embodiments of the present disclosure also provide a method of efficiently encoding / decoding a secondary transformation index signaled from an encoder when configuring a second transformation set based on the size of an encoding block (or transformation block) and / or an intra prediction mode.
[0008] The technical objects of the present disclosure are not limited to the above-mentioned technical objects, and for those of ordinary skill in the art, other technical objects not mentioned above will become apparent from the following description.
[0009] Technical Solution
[0010] In one aspect of the present disclosure, there is provided a method of decoding a video signal, the method including the steps of: generating an inverse quantized transformation block by performing inverse quantization on a current block; obtaining an intra prediction mode of the current block; determining a secondary transformation set applied to the current block among a plurality of secondary transformation sets based on the intra prediction mode; deriving a transformation kernel applied to the current block in the determined secondary transformation set; and performing a secondary transformation on a specific upper left region of the inverse quantized transformation block by using the derived transformation kernel, wherein the plurality of secondary transformation sets may include at least one hybrid secondary transformation set.
[0011] Preferably, the hybrid second transformation set may include at least one 8×8 transformation kernel applied to a region of 8×8 size and at least one 4×4 transformation kernel applied to a region of 4×4 size.
[0012] Preferably, when the plurality of second transformation sets include a plurality of hybrid second transformation sets, the plurality of hybrid second transformation sets may respectively include different numbers of transformation kernels.
[0013] Preferably, deriving the transformation kernel may further include obtaining a second transformation index indicating the transformation kernel applied to the current block in the determined second transformation set, and binarizing the second transformation index by a truncated unary scheme based on the maximum number of available transformation kernels in the determined second transformation set.
[0014] Preferably, determining the second transformation set may include: determining whether to use a hybrid second transformation set based on the intra prediction mode, and when it is determined to use a hybrid second transformation set, determining the second transformation set applied to the current block among the plurality of second transformation sets including the at least one hybrid second transformation set based on the size of the current block, and when it is determined not to use a hybrid second transformation set, determining the second transformation set applied to the current block among the remaining second transformation sets other than the at least one hybrid second transformation set based on the size of the current block.
[0015] In another aspect of the present disclosure, there is provided an apparatus for decoding a video signal, the apparatus including: an inverse quantization unit that generates an inverse quantized transformed block by performing inverse quantization on a current block; a prediction mode acquisition unit that acquires the intra prediction mode of the current block; a second transformation set determination unit that determines the second transformation set applied to the current block among a plurality of second transformation sets based on the intra prediction mode; a transformation kernel derivation unit that derives the transformation kernel applied to the current block in the determined second transformation set; and a second inverse transformation unit that performs a second inverse transformation on a specific upper left region of the inverse quantized transformed block by using the derived transformation kernel, wherein the plurality of second transformation sets may include at least one hybrid second transformation set.
[0016] Preferably, the hybrid second transformation set may include at least one 8×8 transformation kernel applied to a region of 8×8 size and at least one 4×4 transformation kernel applied to a region of 4×4 size.
[0017] Preferably, when the plurality of second transformation sets include a plurality of hybrid second transformation sets, the plurality of hybrid second transformation sets may respectively include different numbers of transformation kernels.
[0018] Preferably, the transform kernel derivation unit may obtain a secondary transform index indicating the transform kernel applied to the current block in the determined set of secondary transforms, and may binarize the secondary transform index by a truncated unary scheme based on the maximum number of available transform kernels in the determined set of secondary transforms.
[0019] Preferably, the secondary transform set determination unit may determine whether to use a hybrid secondary transform set based on the intra prediction mode, and when it is determined to use the hybrid secondary transform set, determine the secondary transform set applied to the current block among the multiple secondary transform sets including the at least one hybrid secondary transform set based on the size of the current block, and when it is determined not to use the hybrid secondary transform set, determine the secondary transform set applied to the current block among the remaining secondary transform sets other than the at least one hybrid secondary transform set based on the size of the current block.
[0020] Advantageous Effects
[0021] According to an embodiment of the present disclosure, by considering various conditions to configure a secondary transform set to apply a secondary transform, a transform for the secondary transform can be efficiently selected.
[0022] The effects achievable in the present disclosure are not limited to the effects mentioned above, and those skilled in the art will clearly understand other unmentioned effects from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a schematic block diagram of an encoder that performs video signal encoding by applying an embodiment of the present disclosure.
[0024] Figure 2 is a schematic block diagram of a decoder that performs video signal decoding by applying an embodiment of the present disclosure.
[0025] Figures 3a to 3d illustrates embodiments to which the present disclosure can be applied, Figure 3a is a diagram describing a quadtree (QT) (hereinafter referred to as "QT") block partitioning structure, Figure 3b is a diagram describing a binary tree (BT) (hereinafter referred to as "BT") block partitioning structure, Figure 3c is a diagram describing a ternary tree (TT) (hereinafter referred to as "TT") block partitioning structure, and Figure 3d is a diagram describing an asymmetric tree (AT) (hereinafter referred to as "AT") block partitioning structure.
[0026] Figure 4It is a schematic block diagram of a transform unit 120, a quantization unit 130, an inverse quantization unit 140, and an inverse transform unit 150 in an encoder to which embodiments of the present disclosure are applied.
[0027] Figure 5 It is a schematic block diagram of an inverse quantization unit 220 and an inverse transform unit 230 in a decoder to which embodiments of the present disclosure are applied.
[0028] Figure 6 It is a table showing a transform configuration group that applies multiple transform selection (MTS) to which embodiments of the present disclosure are applied.
[0029] Figure 7 It is a flowchart showing an encoding process that performs multiple transform selection (MTS) to which embodiments of the present disclosure are applied.
[0030] Figure 8 It is a flowchart showing a decoding process that performs multiple transform selection (MTS) to which embodiments of the present disclosure are applied.
[0031] Figure 9 It is a flowchart for describing an encoding process of an MTS flag and an MTS index to which embodiments of the present disclosure are applied.
[0032] Figure 10 It is a flowchart for describing a decoding process of applying a horizontal transform or a vertical transform to a row or a column based on an MTS flag and an MTS index to which embodiments of the present disclosure are applied.
[0033] Figure 11 It is a flowchart for performing an inverse transform based on transform-related parameters to which embodiments of the present disclosure are applied.
[0034] Figure 12 It is a table showing the assignment of a transform set for each intra prediction mode in NSST to which embodiments of the present disclosure are applied.
[0035] Figure 13 It is a calculation flowchart of Givens rotation to which embodiments of the present disclosure are applied.
[0036] Figure 14 It illustrates a round configuration in a 4×4 NSST composed of a Givens rotation layer and a permutation to which embodiments of the present disclosure are applied.
[0037] Figure 15 It is a block diagram for describing operations of a forward simplification transform and a reverse simplification transform to which embodiments of the present disclosure are applied.
[0038] Figure 16It is a diagram illustrating a process of performing a backward scan from the 64th to the 17th in a backward scan order as an embodiment of applying the present disclosure.
[0039] Figure 17 Illustrates three forward scan orders of a transform coefficient block (transform block) as an embodiment of applying the present disclosure.
[0040] Figure 18 Illustrates the positions and forward scan orders of valid transform coefficients of each 4×4 block when applying diagonal scan and 4×4 RST in the upper left 4×8 block as an embodiment of applying the present disclosure.
[0041] Figure 19 Illustrates a case where the valid transform coefficients of two 4×4 blocks are combined into one 4×4 block when applying diagonal scan and 4×4 RST in the upper left 4×8 block as an embodiment of applying the present disclosure.
[0042] Figure 20 It is a flowchart for encoding a video signal based on a simplified quadratic transform as an embodiment of applying the present disclosure.
[0043] Figure 21 It is a flowchart for decoding a video signal based on a simplified quadratic transform as an embodiment of applying the present disclosure.
[0044] Figure 22 It is a diagram illustrating a method for determining a transform type applied to a quadratic transform according to an embodiment of applying the present disclosure.
[0045] Figure 23 It is a diagram illustrating a simplified transform structure based on a simplification factor that can be applied as an embodiment of the present disclosure.
[0046] Figure 24 It is a diagram illustrating a method for performing decoding by adaptively applying a simplified transform as an embodiment that can be applied to the present disclosure.
[0047] Figure 25 It is a diagram illustrating a method for performing decoding by adaptively applying a simplified transform as an embodiment that can be applied to the present disclosure.
[0048] Figure 26 and Figure 27 It is a diagram illustrating an example of a forward simplified quadratic transform and a backward simplified quadratic transform applied as an embodiment of the present disclosure and the pseudocode for deriving them.
[0049] Figure 28 It is a diagram illustrating a method for applying a simplified quadratic transform to a non-square region as an embodiment of applying the present disclosure.
[0050] Figure 29 is a diagram illustrating a simplified transform controlled by a simplification factor as an embodiment of applying the present disclosure.
[0051] Figure 30 is a diagram illustrating an inverse transform method according to an embodiment of applying the present disclosure.
[0052] Figure 31 is a diagram illustrating an inverse transform unit according to an embodiment of applying the present disclosure.
[0053] Figure 32 Illustrates a video coding system to which the present disclosure is applied.
[0054] Figure 33 is a structural diagram of a content streaming system as an embodiment of applying the present disclosure. Detailed Embodiments
[0055] Hereinafter, the configuration and operation of an embodiment of the present disclosure will be described with reference to the drawings. The configuration and operation of the present disclosure described through the drawings will be described as one embodiment, and thus the technical spirit, core configuration, and operation of the present disclosure are not limited.
[0056] In addition, the terms used in the present disclosure are selected as general terms that are as widely used as possible at present. In specific cases, terms arbitrarily selected by the applicant will be used for description. In this case, since its meaning is clearly described in the detailed description of this part, it should not be simply interpreted only by the name of the term used in the description of the present disclosure. It should be understood that the meaning of the term should be interpreted.
[0057] In addition, when there are general terms selected for describing the present invention or other terms with similar meanings, the terms used in the present disclosure can be replaced with more appropriate interpretations. For example, signals, data, samples, pictures, frames, blocks, etc. can be appropriately replaced and interpreted in each coding process. In addition, separation, decomposition, segmentation, and division can be appropriately replaced and interpreted in each coding process.
[0058] In this document, multiple transform selection (MTS) may refer to a method for performing a transform using at least two transform types. This may also be expressed as adaptive multiple transform (AMT) or explicit multiple transform (EMT). Similarly, mts_idx may also be expressed as AMT_idx, EMT_idx, tu_mts_idx, AMT_TU_idx, EMT_TU_idx, transform index, or transform combination index, and the present disclosure is not limited to these expressions.
[0059] Figure 1It is a schematic block diagram of an encoder that executes video signal encoding as an embodiment of the present disclosure.
[0060] Referring to Figure 1 , the encoder 100 may be configured to include an image partitioning unit 110, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, a filtering unit 160, a decoded picture buffer (DPB) 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190.
[0061] The image partitioning unit 110 may partition an input image (or picture or frame) input to the encoder 100 into one or more processing units. For example, the processing unit may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transform unit (TU).
[0062] However, these terms are only for convenience in describing the present disclosure, and the present disclosure is not limited to the definitions of these terms. Additionally, in the present disclosure, for convenience of description, the term "coding unit" is used as a unit used when encoding or decoding a video signal, but the present disclosure is not limited thereto and may be appropriately interpreted according to the present disclosure.
[0063] The encoder 100 subtracts a prediction signal (or prediction block) output from the inter prediction unit 180 or the intra prediction unit 185 from the input image signal to generate a residual signal (or residual block), and the generated residual signal is sent to the transform unit 120.
[0064] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. The transform process may be applied to a quadtree-structured square block and blocks (square or rectangular) divided according to a binary tree structure, a ternary tree structure, or an asymmetric tree structure.
[0065] The transform unit 120 may perform a transform based on multiple transforms (or transform combinations), and the transform scheme may be referred to as multiple transform selection (MTS). MTS may also be referred to as adaptive multiple transform (AMT) or enhanced multiple transform (EMT).
[0066] MTS (or AMT or EMT) may refer to a transform scheme that performs a transform based on a transform (or transform combination) adaptively selected from multiple transforms (or transform combinations).
[0067] Multiple transforms (or transform combinations) may include the transforms (or transform combinations) described in the Figure 6 of the present disclosure. In the present disclosure, a transform or transform type may be represented as, for example, DCT type 2, DCT-II, DCT2, or DCT-2.
[0068] The transform unit 120 may perform the following embodiments.
[0069] The present disclosure provides methods for designing RSTs that can be applied to 4×4 blocks.
[0070] The present disclosure provides methods for configuring the regions to which 4×4 RSTs are applied, arranging the transform coefficients generated after applying the 4×4 RSTs, the scanning order of the arranged transform coefficients, methods for sorting and combining the transform coefficients generated for each block, and so on.
[0071] The present disclosure provides methods for encoding the transform indices that specify 4×4 RSTs.
[0072] The present disclosure provides methods for conditionally encoding the corresponding transform indices by checking for the presence of non-zero transform coefficients in unacceptable regions when applying 4×4 RSTs.
[0073] The present disclosure provides methods for conditionally encoding the corresponding transform indices after encoding the position of the last non-zero transform coefficient and then omitting the relevant residual encoding for unacceptable positions.
[0074] The present disclosure provides methods for applying different transform index encodings and residual encodings to luminance blocks and chrominance blocks when applying 4×4 RSTs.
[0075] Its specific embodiments will be described in more detail in the present disclosure.
[0076] The quantization unit 130 may quantize the transform coefficients and send the quantized transform coefficients to the entropy encoding unit 190, and the entropy encoding unit 190 may perform entropy encoding on the quantized signal and output the entropy-encoded quantized signal as a bitstream.
[0077] Although the transform unit 120 and the quantization unit 130 are described as separate functional units, the present disclosure is not limited thereto, and they may be combined into one functional unit. The dequantization unit 140 and the inverse transform unit 150 may also be similarly combined into one functional unit.
[0078] The quantized signal output from the quantization unit 130 may be used to generate a prediction signal. For example, the dequantization and inverse transform are cyclically applied to the quantized signal by the dequantization unit 140 and the inverse transform unit 150 to reconstruct the residual signal. The reconstructed residual signal is added to the prediction signal output from the inter-frame prediction unit 180 or the intra-frame prediction unit 185 to generate a reconstructed signal.
[0079] In addition, due to the quantization error that occurs during such compression processing, deterioration showing block boundaries may occur. This phenomenon is called blocking artifact, which is one of the key elements for evaluating image quality. To reduce the deterioration, filtering processing can be performed. Through the filtering processing, the blocking artifact is eliminated and the error of the current picture is reduced to enhance the image quality.
[0080] The filtering unit 160 applies filtering to the reconstructed signal and outputs the applied reconstructed signal to the reproduction device, or sends the output reconstructed signal to the decoded picture buffer 170. The inter-frame prediction unit 180 can use the filtered signal sent to the decoded picture buffer 170 as a reference picture. In this way, in the inter-picture prediction mode, the filtered picture is used as a reference picture to enhance the image quality and coding efficiency.
[0081] The decoded picture buffer 170 can store the filtered picture so as to use the filtered picture as a reference picture in the inter-frame prediction unit 180.
[0082] The inter-frame prediction unit 180 performs temporal prediction and / or spatial prediction in order to remove temporal redundancy and / or spatial redundancy by referring to the reconstructed pictures. Here, since the reference pictures used for prediction are transform signals that are quantized and dequantized in units of blocks during encoding / decoding at a previous time, there may be blocking artifacts or ringing artifacts.
[0083] Therefore, the inter-frame prediction unit 180 can interpolate signals between pixels in units of sub-pixels by applying a low-pass filter in order to solve the performance degradation caused by the discontinuity or quantization of this signal. Here, sub-pixels mean virtual pixels generated by applying an interpolation filter, and integer pixels mean actual pixels existing in the reconstructed picture. As an interpolation method, linear interpolation, bilinear interpolation, Wiener filter, etc. can be adopted.
[0084] An interpolation filter is applied to the reconstructed picture to enhance the prediction accuracy. For example, the inter-frame prediction unit 180 applies an interpolation filter to integer pixels to generate interpolated pixels, and can perform prediction by using the interpolated block composed of the interpolated pixels as a prediction block.
[0085] In addition, the intra prediction unit 185 can predict the current block by referring to samples near the block to be currently encoded. The intra prediction unit 185 can perform the following processes to perform intra prediction. First, reference samples, which are required to generate a prediction signal, can be prepared. Additionally, a prediction signal can be generated by using the prepared reference samples. Thereafter, the prediction mode is encoded. In this case, the reference samples can be prepared by zero-padding the reference samples and / or filtering the reference samples. Since the reference samples have undergone prediction and reconstruction processes, there may be quantization errors. Therefore, a reference sample filtering process can be performed for each prediction mode used for intra prediction to reduce this error.
[0086] The prediction signal generated by the inter prediction unit 180 or the intra prediction unit 185 can be used to generate a reconstructed signal or to generate a residual signal.
[0087] Figure 2 is a schematic block diagram of a decoder that executes video signal decoding as an embodiment of applying the present disclosure.
[0088] Referring to Figure 2 , the decoder 200 can be configured to include a parsing unit (not illustrated), an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, a filtering unit 240, a decoded picture buffer (DPB) unit 250, an inter prediction unit 260, and an intra prediction unit 265.
[0089] Additionally, the reconstructed video signal output by the decoder can be reproduced by a reproducing device.
[0090] The decoder 200 can receive a signal output from the Figure 1 encoder 100 and can perform entropy decoding on the received signal by the entropy decoding unit 210.
[0091] The inverse quantization unit 220 obtains transform coefficients from the entropy-decoded signal by using quantization step information.
[0092] The inverse transform unit 230 performs an inverse transform on the transform coefficients to obtain a residual signal.
[0093] Here, the present disclosure provides a method for configuring a transform combination for each transform configuration group divided by at least one of a prediction mode, a block size, or a block shape, and the inverse transform unit 230 can perform an inverse transform based on the transform combination configured by the present disclosure. Additionally, the embodiments described in the present disclosure can be applied.
[0094] The inverse transform unit 230 can perform the following embodiments.
[0095] The present disclosure provides a method for reconstructing a video signal based on a simplified quadratic transform.
[0096] The inverse transform unit 230 may derive an inverse secondary transform corresponding to a secondary transform index, perform an inverse secondary transform on a transform coefficient block by using the secondary transform, and perform an inverse primary transform on the block on which the inverse secondary transform is performed. Here, the secondary transform refers to a simplified secondary transform, and the simplified secondary transform represents a transform that inputs N residual data (N×1 residual vector) to output L (L < N) transform coefficient data (L×1 transform coefficient vector).
[0097] The present disclosure is characterized in that a simplified secondary transform is applied to a specific region of a current block, and the specific region is the upper left M×M (M≤N) in the current block.
[0098] The present disclosure is characterized in that when performing an inverse secondary transform, a 4×4 simplified secondary transform is applied to each of the 4×4 blocks divided in the current block.
[0099] The present disclosure is characterized in that it is determined whether a secondary transform index is obtained based on the position of the last non-zero transform coefficient in a transform coefficient block.
[0100] The present disclosure is characterized in that when the last non-zero transform coefficient is not in a specific region, a secondary transform index is obtained, and the specific region indicates the remaining region except for the positions where non-zero transform coefficients may exist when arranging transform coefficients according to a scan order in the case of applying a simplified secondary transform.
[0101] The inverse transform unit 230 may derive a transform combination corresponding to a primary transform index, and perform an inverse primary transform by using the transform combination. Here, the primary transform index corresponds to any one of a plurality of transform combinations formed by a combination of DST7 and / or DCT8, and the transform combination includes a horizontal transform and a vertical transform. In this case, the horizontal transform and the vertical transform correspond to DST7 or DCT8.
[0102] Although the inverse quantization unit 220 and the inverse transform unit 230 are described as separate functional units, the present disclosure is not limited thereto, and they may be combined into one functional unit.
[0103] The obtained residual signal is added to a prediction signal output from the inter-frame prediction unit 260 or the intra-frame prediction unit 265 to generate a reconstructed signal.
[0104] The filtering unit 240 applies filtering to the reconstructed signal, and outputs the applied reconstructed signal to a generation device, or sends the output reconstructed signal to the decoded picture buffer unit 250. The inter-frame prediction unit 260 may use the filtered signal sent to the decoded picture buffer unit 250 as a reference picture.
[0105] In the present disclosure, the embodiments described in the respective functional units of the transform unit 120 and the encoder 100 can be equivalently applied to the corresponding functional units of the inverse transform unit 230 and the decoder, respectively.
[0106] Figures 3a to 3d Examples of embodiments to which the present disclosure can be applied are Figure 3a a diagram illustrating a quadtree (QT) (hereinafter referred to as "QT") block partitioning structure, Figure 3b a diagram illustrating a binary tree (BT) (hereinafter referred to as "BT") block partitioning structure, Figure 3c a diagram illustrating a ternary tree (TT) (hereinafter referred to as "TT") block partitioning structure, and Figure 3d a diagram illustrating an asymmetric tree (AT) (hereinafter referred to as "AT") block partitioning structure.
[0107] During video encoding, a block can be partitioned based on a quadtree (QT). Additionally, a sub-block partitioned by QT can be further recursively partitioned using QT. A leaf block that is no longer partitioned by QT can be partitioned by at least one of a binary tree (BT), a ternary tree (TT), and an asymmetric tree (AT). The BT can have two partitioning types: horizontal BT (2N×N, 2N×N) and vertical BT (N×2N, N×2N). The TT can have two partitioning types: horizontal TT (2N×1 / 2N, 2N×N, 2N×1 / 2N) and vertical TT (1 / 2N×2N, N×2N, 1 / 2N×2N). The AT can have four partitioning types: horizontal - upper AT (2N×1 / 2N, 2N×3 / 2N), horizontal - lower AT (2N×3 / 2N, 2N×1 / 2N), vertical - left AT (1 / 2N×2N, 3 / 2N×2N), vertical - right AT (3 / 2N×2N, 1 / 2N×2N). Each of the BT, TT, and AT can be further recursively partitioned by using the BT, TT, and AT.
[0108] Figure 3a Examples of QT partitioning are illustrated. Block A can be partitioned into four sub - blocks A0, A1, A2, and A3 by QT. Sub - block A1 can be further partitioned into four sub - blocks B0, B1, B2, and B3 by QT again.
[0109] Figure 3b Examples of BT partitioning are illustrated. Block B3, which is no longer partitioned by QT, can be partitioned into a vertical BT (C0, C1) or a horizontal BT (D0, D1). Each sub - block, like block C0, can be further recursively partitioned in the form of a horizontal BT (E0, E1) or a vertical BT (F0, F1).
[0110] Figure 3cAn example of the TT partition is illustrated. The block B3 that is no longer partitioned by the QT can be partitioned into a vertical TT (C0, C1, C2) or a horizontal TT (D0, D1, D2). Similar to the block C1, each sub-block can be further recursively partitioned in the form of a horizontal TT (E0, E1, E2) or a vertical TT (F0, F1, F2).
[0111] Figure 3d An example of the AT partition is illustrated. The block B3 that is no longer partitioned by the QT can be partitioned into a vertical AT (C0, C1) or a horizontal AT (D0, D1). Similar to the block C1, each sub-block can be further recursively partitioned in the form of a horizontal AT (E0, E1) or a vertical TT (F0, F1).
[0112] In addition, the BT, TT, and AT partitions can be used and performed together. For example, a sub-block partitioned by the BT can be partitioned by the TT or the AT. Additionally, a sub-block partitioned by the TT can be partitioned by the BT or the AT. Additionally, a sub-block partitioned by the AT can be partitioned by the BT or the TT. For example, after a horizontal BT partition, each sub-block can be partitioned into a vertical BT, or after a vertical BT partition, each sub-block can be partitioned into a horizontal BT. The two types of partitioning methods are different from each other in terms of the partitioning order, but the same in terms of the final partitioning shape.
[0113] In addition, when a block is partitioned, the order of searching the block can be defined in various ways. Generally, the search can be performed from left to right and from top to bottom, and the search of the block can mean the order of deciding whether to further partition each partitioned sub-block, the encoding order of each sub-block when each sub-block is no longer partitioned, or the search order when referring to the information of another adjacent block in the sub-block.
[0114] Figure 4 and Figure 5 illustrates an embodiment applying the present disclosure, Figure 4 is a schematic block diagram of a transform unit 120, a quantization unit 130, an inverse quantization unit 140, and an inverse transform unit 150 in an encoder, and Figure 5 is a schematic block diagram of an inverse quantization unit 220 and an inverse transform unit 230 in a decoder.
[0115] Referring to Figure 4 the transform unit 120 and the quantization unit 130 may include a primary transform unit 121, a secondary transform unit 122, and a quantization unit 130. The inverse quantization unit 140 and the inverse transform unit 150 may include an inverse quantization unit 140, an inverse secondary transform unit 151, and an inverse primary transform unit 152.
[0116] Referring to Figure 5, the dequantization unit 220 and the inverse transformation unit 230 may include a dequantization unit 220, an inverse secondary transformation unit 231, and an inverse primary transformation unit 232.
[0117] In the present disclosure, when performing a transformation, the transformation may be performed in multiple steps. For example, two steps of a primary transformation and a secondary transformation may be applied as exemplified in Figure 4 , or more transformation steps may be used according to the algorithm. Here, the primary transformation may also be referred to as a core transformation.
[0118] The primary transformation unit 121 may apply a primary transformation to the residual signal, and here, the primary transformation may be defined in a table of an encoder and / or a decoder.
[0119] The primary transformation may employ a discrete cosine transform type 2 (hereinafter, referred to as "DCT2").
[0120] Alternatively, only in specific cases, a discrete sine transform type 7 (hereinafter, referred to as "DST7") may be employed. For example, DST7 may be applied to a 4×4 block in an intra prediction mode.
[0121] In addition, the primary transformation may employ a combination of various transformations of multiple transform selection (MTS), DST 7, DCT 8, DST 1, and DCT 5. For example, Figure 6 .
[0122] The secondary transformation unit 122 may apply a secondary transformation to the signal after the primary transformation, and here, the secondary transformation may be defined in a table of an encoder and / or a decoder.
[0123] As an implementation, the secondary transformation may conditionally employ an inseparable secondary transformation (hereinafter, referred to as "NSST"). For example, NSST may be applied only to intra prediction blocks and may have a set of transformations suitable for each group of prediction modes.
[0124] Here, the group of prediction modes may be configured based on symmetry with respect to the prediction direction. For example, since prediction mode 52 and prediction mode 16 are symmetric based on prediction mode 34 (diagonal direction), the same set of transformations may be applied by forming a group. In this case, when applying the transformation for prediction mode 52, since prediction mode 52 has the same set of transformations as prediction mode 16, the input data is transposed and then applied.
[0125] In addition, since there is no symmetry in the direction in the case of the planar mode and the DC mode, each mode has a different set of transformations, and the corresponding set of transformations may be composed of two transformations. With respect to the remaining direction modes, each set of transformations may be composed of three transformations.
[0126] As another embodiment, the secondary transform may employ a combination of various transforms DST 7, DCT 8, DST 1, and DCT 5 of multiple transform selection (MTS). For example, it may employ Figure 6 .
[0127] As another embodiment, DST 7 may be applied to the secondary transform.
[0128] As another embodiment, the secondary transform may not be applied to the entire primary transform block but may be applied only to a specific upper left region. For example, when the block size is 8×8 or larger, 8×8 NSST is applied, and when the block size is less than 8×8, a 4×4 secondary transform may be applied. In this case, the block may be divided into 4×4 blocks, and then the 4×4 secondary transform may be applied to each divided block.
[0129] As another embodiment, even in the case of 4×N / N×4 (N >= 16), a 4×4 secondary transform may be applied.
[0130] The secondary transform (e.g., NSST), the 4×4 secondary transform, and the 8×8 secondary transform will be described in more detail with reference to Figures 12 to 15 and other embodiments in the specification.
[0131] The quantization unit 130 may perform quantization on the signal after the secondary transform.
[0132] The dequantization unit 140 and the inverse transform unit 150 perform the above processing conversely, and redundant descriptions thereof will be omitted.
[0133] Figure 5 is a schematic block diagram of the dequantization unit 220 and the inverse transform unit 230 in the decoder.
[0134] Referring to Figure 5 , the dequantization unit 220 and the inverse transform unit 230 may include a dequantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.
[0135] The dequantization unit 220 obtains transform coefficients from the entropy-decoded signal by using quantization step information.
[0136] The inverse secondary transform unit 231 performs an inverse secondary transform on the transform coefficients. Here, the inverse secondary transform represents the inverse transform of the secondary transform described in Figure 4 .
[0137] As another embodiment, the secondary transform may employ a combination of various transforms DST 7, DCT 8, DST 1, and DCT 5 of multiple transform selection (MTS). For example, it may employ Figure 6 .
[0138] The inverse primary transform unit 232 performs an inverse primary transform on the signal (or block) after the inverse secondary transform and obtains a residual signal. Here, the inverse primary transform represents the inverse transform of the primary transform described in Figure 4 .
[0139] As an implementation, the primary transform may employ a combination of various transforms DST 7, DCT 8, DST 1, and DCT 5 of multiple transform selection (MTS). For example, Figure 6 may be employed.
[0140] As an implementation of the present disclosure, DST 7 may be applied to the primary transform.
[0141] As an implementation of the present disclosure, DST 8 may be applied to the primary transform.
[0142] The present disclosure provides a method for configuring a transform combination for each transform configuration group divided according to at least one of a prediction mode, a block size, or a block shape, and the inverse primary transform unit 232 may perform an inverse transform based on the transform combination configured by the present disclosure. Additionally, the implementations described in the present disclosure may be applied.
[0143] Figure 6 is a table showing a transform configuration group applying multiple transform selection (MTS) as an implementation of applying the present disclosure.
[0144] Transformation configuration group applying multiple transform selection (MTS)
[0145] In the present disclosure, the j-th transform combination candidate for the transform configuration group G i is represented by the pair shown in Equation 1 below.
[0146] [Equation 1]
[0147] (H(G i , j), V(G i , j))
[0148] Here, H(Gi,j) indicates the horizontal transform of the j-th candidate, and V(Gi,j) indicates the vertical transform of the j-th candidate. For example, in Figure 6 , it may be represented that H(G3,2) = DST7 and V(G3,2) = DCT8. Depending on the context, the value assigned to H(G i , j) or V(G i , j) may be a nominal value as in the above example to distinguish the transform, or may be an index value indicating the transform, or may be a two-dimensional (2D) matrix for the transform.
[0149] Additionally, in the present disclosure, the 2D matrix values of DCT and DST may be represented as shown in Equation 2 and Equation 3 below.
[0150] [Formula 2]
[0151] DCT Type 2: DCT Type 8:
[0152] [Formula 3]
[0153] DST Type 7: DST Type 4:
[0154] Here, regardless of whether DST or DCT is represented by S or C, the type number is represented in Roman numerals in the superscript, and N in the subscript represents an N×N transform. Additionally, 2D matrices such as and are assumed to have column vectors forming the transform basis.
[0155] Referring to Figure 6 , the transform configuration groups can be determined based on the prediction mode, and the number of groups can be a total of six groups G0 to G5. Additionally, G0 to G4 correspond to the cases where intra-frame prediction is applied, and G5 represents the transform combination (or transform set and transform combination set) applied to the residual block generated by inter-frame prediction.
[0156] A transform combination can be composed of a horizontal transform (or row transform) applied to the rows of the corresponding 2D block and a vertical transform (or column transform) applied to the columns.
[0157] Here, each of all the transform configuration groups can have four transform combination candidates. The four transform combinations can be selected or determined by transform combination indices 0 to 3, and are sent from the encoder to the decoder by encoding the transform combination indices.
[0158] As an implementation, according to the intra-frame prediction mode, the residual data (or residual signal) obtained by intra-frame prediction can have different statistical characteristics. Therefore, as Figure 6 illustrated in, transforms other than the general cosine transform can be applied to each intra-frame prediction mode.
[0159] Referring to Figure 6 , the cases of using 35 intra-frame prediction modes and the cases of using 67 intra-frame prediction modes are illustrated. Multiple transform combinations can be applied to each transform configuration group divided in each intra-frame prediction mode column. For example, these multiple transform combinations can be composed of four (row direction transform and column direction transform) combinations. As a specific example, DST-7 and DST-5 can be applied in the row (horizontal) direction and column (vertical) direction in group 0, and as a result, a total of four combinations are available.
[0160] Since a total of four transform kernel combinations can be applied to each intra prediction mode, a transform combination index for selecting one of the transform kernel combinations can be sent for each transform unit. In the present disclosure, the transform combination index may be referred to as an MTS index and denoted as mts_idx.
[0161] In addition, in addition to the transform kernels represented above Figure 6 it is also possible that, due to the characteristics of the residual signal, DCT2 is optimal for both the row direction and the column direction. Therefore, an MTS flag is defined for each coding unit to perform the transform adaptively. Here, when the MTS flag is 0, DCT 2 can be applied to both the row direction and the column direction, and when the MTS flag is 1, one of the four combinations can be selected or determined by the MTS index.
[0162] As an implementation, when the MTS flag is 1, if the number of non-zero transform coefficients for a transform unit is not greater than a threshold, DST-7 can be applied to both the row direction and the column direction without applying Figure 6 the transform kernel. For example, the threshold can be set to 2, and this threshold can be set differently based on the block size or the size of the transform unit. This also applies to other implementations in the specification.
[0163] As an implementation, if the number of non-zero transform coefficients is not greater than the threshold, by first parsing the transform coefficient values, the amount of additional information transmission can be reduced by applying DST-7 without parsing the MTS index.
[0164] As an implementation, when the MTS flag is 1, if the number of non-zero transform coefficients is greater than the threshold for a transform unit, the MTS index can be parsed, and the horizontal transform and the vertical transform can be determined based on the MTS index.
[0165] As an implementation, MTS can be applied only when both the width and height of the transform unit are equal to or less than 32.
[0166] As an implementation, it can be pre-configured through offline training Figure 6 .
[0167] As an implementation, the MTS index can be defined as an index that can simultaneously indicate the horizontal transform and the vertical transform. Alternatively, the MTS index can be defined separately as a horizontal transform index and a vertical transform index.
[0168] As an implementation manner, the MTS flag or MTS index can be defined at at least one level among a sequence, a picture, a slice, a block, a coding unit, a transform unit, or a prediction unit. For example, the MTS flag or MTS index can be defined at at least one level among a sequence parameter set (SPS), a coding unit, or a transform unit. Additionally, as an example, a syntax flag for enabling / disabling MTS can be defined at at least one level among a sequence parameter set (SPS), a picture parameter set (PPS), or a slice header.
[0169] As another implementation manner, the transform combination corresponding to the transform index (horizontal transform or vertical transform) can be configured without depending on the MTS flag, the prediction mode, and / or the block shape. For example, the transform combination can be composed of at least one of DCT2, DST7, and / or DCT8. As a specific example, when the transform index is 0, 1, 2, 3, or 4, each transform combination can be (DCT2, DCT2), (DST7, DST7), (DCT8, DST7), (DST7, DCT8), or (DCT8, DCT8).
[0170] Figure 7 is a flowchart showing an encoding process for performing multiple transform selection (MTS) as an implementation manner of applying the present disclosure.
[0171] In the present disclosure, an implementation manner in which transforms are respectively applied to the horizontal direction and the vertical direction is basically described, but the transform combination can be configured as an inseparable transform.
[0172] Alternatively, the transform combination can be composed of a mixture of a separable transform and an inseparable transform. In this case, when an inseparable transform is used, row / column transform selection or horizontal / vertical direction selection may not be required, and only when a separable transform is selected, the Figure 6 transform combination can be used.
[0173] Additionally, the solution proposed by the present disclosure can be applied regardless of whether it is a primary transform or a secondary transform. That is, there is no restriction that the solution should be applied only to either the primary transform or the secondary transform, and the solution can be applied to both the primary transform and the secondary transform. Here, the primary transform can mean a transform for first transforming a residual block, and the secondary transform can mean a transform for applying a transform to the block generated as a result of the primary transform.
[0174] First, the encoder can determine a transform configuration group corresponding to the current block. Here, the transform configuration group can mean Figure 6 the transform configuration group of, the present disclosure is not limited thereto, and the transform configuration group can be composed of other transform combinations.
[0175] The encoder may perform a transform (S720) on candidate transform combinations available in a transform configuration group.
[0176] As a result of performing the transform, the encoder may determine or select a transform combination having the minimum rate distortion (RD) cost (S730).
[0177] The encoder may encode a transform combination index corresponding to the selected transform combination (S740).
[0178] Figure 8 FIG. is a flowchart showing a decoding process for performing multiple transform selection (MTS) as an embodiment of applying the present disclosure.
[0179] First, the decoder may determine a transform configuration group for a current block (S810).
[0180] The decoder may parse (or obtain) a transform combination index from a video signal, and here, the transform combination index may correspond to any one of multiple transform combinations in the transform configuration group (S820). For example, the transform configuration group may include discrete sine transform type (DST) 7 and discrete cosine transform type (DCT) 8. The transform combination index may be referred to as an MTS index.
[0181] As an embodiment, the transform configuration group may be configured based on at least one of a prediction mode, a block size, or a block shape of the current block.
[0182] The decoder may derive a transform combination corresponding to the transform combination index (S830). Here, the transform combination may be composed of a horizontal transform and a vertical transform, and may include at least one of DST-7 or DCT-8.
[0183] In addition, the transform combination may mean the transform combination described in Table 6, but the present disclosure is not limited thereto. That is, depending on other embodiments in the present disclosure, the transform combination may be composed of other transform combinations.
[0184] The decoder may perform an inverse transform on the current block based on the transform combination (S840). When the transform combination is composed of a row (horizontal) transform and a column (vertical) transform, the column (vertical) transform may be applied after first applying the row (horizontal) transform. However, the present disclosure is not limited thereto, and the transform order may be reversed, or when the transform combination is composed of an inseparable transform, the inseparable transform may be applied immediately.
[0185] As an embodiment, when the vertical transform or the horizontal transform is DST-7 or DCT-8, the inverse transform of DST-7 or DCT-8 may be applied to each column and then to each column.
[0186] As an implementation, for vertical transformation or horizontal transformation, different transformations can be applied to each row and / or each column.
[0187] As an implementation, a transformation combination index can be obtained based on an MTS flag indicating whether to perform MTS. That is, when MTS is performed according to the MTS flag, a transformation combination index can be obtained.
[0188] As an implementation, the decoder can check whether the number of non-zero transformation coefficients is greater than a threshold. In this case, a transformation can be obtained when the number of non-zero transformation coefficients is greater than the threshold.
[0189] As an implementation, the MTS flag or MTS index can be defined at at least one level of sequence, picture, slice, block, coding unit, transformation unit, or prediction unit.
[0190] As an implementation, an inverse transformation can be applied only when both the width and height of the transformation unit are equal to or less than 32.
[0191] On the other hand, as another implementation, the process of determining a transformation configuration group and the process of parsing a transformation combination index can be performed simultaneously. Alternatively, step S810 can be pre-configured and omitted in the encoder and / or decoder.
[0192] Figure 9 is a flowchart for describing the encoding process of the MTS flag and MTS index as an implementation of applying the present disclosure.
[0193] The encoder can determine whether to apply Multiple Transform Selection (MTS) to the current block (S910).
[0194] When applying Multiple Transform Selection (MTS), the encoder can encode MTS flag = 1 (S920).
[0195] In addition, the encoder can determine the MTS index based on at least one of the prediction mode, horizontal transformation, and vertical transformation of the current block (S930). Here, the MTS index can mean an index indicating any one of multiple transformation combinations for each intra prediction mode, and the MTS index can be sent for each transformation unit.
[0196] When the MTS index is determined, the encoder can encode the MTS index (S940).
[0197] On the other hand, when not applying Multiple Transform Selection (MTS), the encoder can encode MTS flag = 0 (S920).
[0198] Figure 10It is a flowchart for describing a decoding process of applying a horizontal transformation or a vertical transformation to a row or a column based on an MTS flag and an MTS index as an embodiment of the present disclosure.
[0199] The decoder can parse the MTS flag from the bitstream (S1010). Here, the MTS flag can indicate whether to apply Multiple Transform Selection (MTS) to the current block.
[0200] The decoder can determine whether to apply Multiple Transform Selection (MTS) to the current block based on the MTS flag (S1020). For example, it can be checked whether the MTS flag is 1.
[0201] When the MTS flag is 1, the decoder can check whether the number of non-zero transform coefficients is greater than (or equal to or greater than) a threshold (S1030). For example, the threshold can be set to 2, and this threshold can be set differently based on the block size or the size of the transform unit.
[0202] When the number of non-zero transform coefficients is greater than the threshold, the decoder can parse the MTS index (S1040). Here, the MTS index can mean any one of multiple transform combinations for each intra prediction mode, and the MTS index can be sent for each transform unit. Alternatively, the MTS index can mean an index indicating any one of the transform combinations defined in a pre-configured transform combination table, and here, the pre-configured transform combination table can mean Figure 6 , but the present disclosure is not limited thereto.
[0203] The decoder can derive or determine a horizontal transformation and a vertical transformation based on at least one of the MTS index and the prediction mode (S1050).
[0204] Alternatively, the decoder can derive the transform combination corresponding to the MTS index. For example, the decoder can derive or determine the horizontal transformation and the vertical transformation corresponding to the MTS index.
[0205] When the number of non-zero transform coefficients is not greater than the threshold, the decoder can apply a pre-configured vertical inverse transformation (S1060). For example, the vertical inverse transformation can be the inverse transformation of DST7.
[0206] In addition, the decoder can apply a pre-configured horizontal inverse transformation to each row (S1070). For example, the horizontal inverse transformation can be the inverse transformation of DST7. That is, when the number of non-zero transform coefficients is not greater than the threshold, a transform kernel pre-configured by the encoder or the decoder can be used. For example, a transform kernel that is not defined in the Figure 6 illustrated transform combination table but is widely used (such as DCT-2, DST-7, or DCT-8) can be used.
[0207] In addition, when the MTS flag is 0, the decoder can apply a pre-configured inverse vertical transform to each column (S1080). For example, the inverse vertical transform can be the inverse transform of DCT2.
[0208] In addition, the decoder can apply a pre-configured inverse horizontal transform to each row (S1090). For example, the inverse horizontal transform can be the inverse transform of DCT2. That is, when the MTS flag is 0, a transform kernel pre-configured by the encoder or the decoder can be used. For example, a transform kernel that is not defined in the transform combination table illustrated in Figure 6 but is widely used can be used.
[0209] Figure 11 is a flowchart for performing an inverse transform based on transform-related parameters in an embodiment of applying the present disclosure.
[0210] A decoder applying the present disclosure can obtain sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag (S1110). Here, sps_mts_intra_enabled_flag indicates whether there is a tu_mts_flag in the residual coding syntax of an intra-coded unit. For example, when sps_mts_intra_enabled_flag = 0, there is no tu_mts_flag in the residual coding syntax of the intra-coded unit, and when sps_mts_intra_enabled_flag = 0, there is a tu_mts_flag in the residual coding syntax of the intra-coded unit. In addition, sps_mts_inter_enabled_flag indicates whether there is a tu_mts_flag in the residual coding syntax of an inter-coded unit. For example, when sps_mts_inter_enabled_flag = 0, there is no tu_mts_flag in the residual coding syntax of the intra-coded unit, and when sps_mts_inter_enabled_flag = 0, there is a tu_mts_flag in the residual coding syntax of the inter-coded unit.
[0211] The decoder can obtain tu_mts_flag based on sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag (S1120). For example, when sps_mts_intra_enabled_flag = 1 or sps_mts_inter_enabled_flag = 1, the decoder can obtain tu_mts_flag. Here, tu_mts_flag indicates whether multiple transform selection (hereinafter referred to as "MTS") is applied to the residual samples of the luma transform block. For example, when tu_mts_flag = 0, MTS is not applied to the residual samples of the luma transform block, and when tu_mts_flag = 1, MTS is applied to the residual samples of the luma transform block.
[0212] As another example, at least one of the embodiments in this document can be applied to tu_mts_flag.
[0213] The decoder can obtain mts_idx based on tu_mts_flag (S1130). For example, when tu_mts_flag = 1, the decoder can obtain mts_idx. Here, mts_idx indicates which transform kernel is applied to the luma residual samples along the horizontal and / or vertical directions of the current transform block.
[0214] For example, at least one of the embodiments in this document can be applied to mts_idx. As a specific example, the embodiments that can be applied Figure 6 are at least one of the embodiments.
[0215] The decoder can derive the transform kernel corresponding to mts_idx (S1140). For example, it can be defined by dividing the transform kernel corresponding to mts_idx into horizontal transform and vertical transform.
[0216] As another example, different transform kernels can be applied to the horizontal transform and the vertical transform. However, the present disclosure is not limited thereto, and the same transform kernel can be applied to the horizontal transform and the vertical transform.
[0217] As an embodiment, mts_idx can be defined as shown in Table 1 below.
[0218] [Table 1]
[0219] mts_idx[x0][y0] trTypeHor trTypeVer 0 0 0 1 1 1 2 2 1 3 1 2 4 2 2
[0220] In addition, the decoder can perform inverse transform based on the transform kernel (S1150).
[0221] Above Figure 11Among them, the following implementation manners are mainly described: obtaining tu_mts_flag to determine whether to apply MTS, and obtaining mts_idx according to the subsequently obtained value of tu_mts_flag to determine the transform kernel, but the present disclosure is not limited thereto. As an example, the decoder directly parses mts_idx without parsing tu_mts_flag to determine the transform kernel. In this case, Table 1 above can be used. That is, when the mts_idx value indicates 0, DCT-2 can be applied in the horizontal / vertical direction, and when the mts_idx value indicates a value other than 0, DST-7 and / or DCT-8 can be applied according to the mts_idx value.
[0222] As another implementation manner of the present disclosure, a decoding process for performing transform processing is described.
[0223] The decoder may check the transform size nTbS (S10). Here, the transform size nTbS may be a variable representing the horizontal sample size of the scaled transform coefficients.
[0224] The decoder may check the transform kernel type trType (S20). Here, the transform kernel type trType may be a variable representing the transform kernel type, and various implementation manners of this document may be applied. The transform kernel type trType may include a horizontal transform kernel type trTypeHor and a vertical transform kernel type trTypeVer.
[0225] Referring to Table 1, when the transform kernel type trType is 0, the transform kernel type may represent DCT2, when the transform kernel type trType is 1, the transform kernel type may represent DST7, and when the transform kernel type trType is 2, the transform kernel type may represent DCT8.
[0226] The decoder may perform transform matrix multiplication (S30) based on at least one of the transform size nTbS or the transform kernel type.
[0227] As another example, when the transform kernel type is 1 and the transform size is 4, a predetermined transform matrix 1 may be applied when performing transform matrix multiplication.
[0228] As another example, when the transform kernel type is 1 and the transform size is 8, a predetermined transform matrix 2 may be applied when performing transform matrix multiplication.
[0229] As another example, when the transform kernel type is 1 and the transform size is 16, a predetermined transform matrix 3 may be applied when performing transform matrix multiplication.
[0230] As another example, when the transform kernel type is 1 and the transform size is 32, a predefined transform matrix 4 can be applied when performing the transform matrix multiplication.
[0231] Similarly, when the transform kernel type is 2 and the transform sizes are 4, 8, 16, or 32, predefined transform matrices 5, 6, 7, and 8 can be applied respectively.
[0232] Here, each of the predefined transform matrices 1 to 8 can correspond to any one of various types of transform matrices. As an example, a transform matrix of the type exemplified in Figure 6 can be applied.
[0233] The decoder can derive transformed samples (or transform coefficients) based on the transform matrix multiplication (S40).
[0234] Each of the above embodiments can be used, but the present disclosure is not limited thereto, and it can be used in combination with the above embodiments and other embodiments of the present disclosure.
[0235] Figure 12 is a table showing the assignment of transform sets for each intra prediction mode in NSST as an embodiment of applying the present disclosure.
[0236] Non-separable second-order transform (NSST)
[0237] The secondary transform unit can apply a secondary transform to the signal after the primary transform, and here, the secondary transform can be defined in a table of the encoder and / or decoder.
[0238] As an embodiment, the secondary transform can conditionally employ an inseparable secondary transform (hereinafter referred to as "NSST"). For example, NSST can be applied only to intra prediction blocks and can have transform sets suitable for each group of prediction modes.
[0239] Here, the groups of prediction modes can be configured based on symmetry with respect to the prediction direction. For example, since prediction mode 52 and prediction mode 16 are symmetric based on prediction mode 34 (diagonal direction), the same transform set can be applied by forming one group. In such a case, when applying the transform for prediction mode 52, since prediction mode 52 has the same transform set as prediction mode 16, the input data is transposed and then applied.
[0240] In addition, since there is no symmetry in the direction in the case of the planar mode and the DC mode, each mode has a different transform set, and the corresponding transform set can be composed of two transforms. With respect to the remaining direction modes, each transform set can be composed of three transforms. However, the present disclosure is not limited thereto, and each transform set can be composed of multiple transforms.
[0241] In an embodiment, a transform set table other than the transform set table exemplified in Figure 12 may be defined. For example, as shown in Table 2 below, a transform set may be determined from a predefined table according to an intra prediction mode (or a group of intra prediction modes). Syntax indicating a specific transform in the transform set determined according to the intra prediction mode may be signaled from an encoder to a decoder.
[0242] [Table 2]
[0243] IntraPredMode Transform set index IntraPredMode<0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= IntraPredMode 1
[0244] Referring to Table 2, a predefined transform set may be assigned to a group of intra prediction modes (or a group of intra prediction modes). Here, the IntraPredMode value may be a mode value transformed in consideration of wide-angle intra prediction (WAIP).
[0245] Figure 13 is a flowchart of the calculation of Givens rotation as an embodiment applying the present disclosure.
[0246] As another embodiment, instead of applying a secondary transform to the entire primary transform block, the secondary transform may be applied only to the upper left 8×8 region. For example, when the block size is 8×8 or larger, 8×8 NSST is applied, and when the block size is less than 8×8, 4×4 NSST is applied, and in this case, the block is divided into 4×4 blocks, and then 4×4 NSST is applied to each of the divided blocks.
[0247] As another embodiment, even in the case of 4×N / N×4 (N >= 16), 4×4 NSST may be applied.
[0248] Since both 8×8 NSST and 4×4 NSST follow the transform combination configuration described in this document and are non-separable transforms, 8×8 NSST receives 64 data and outputs 64 data, and 4×4 NSST has 16 inputs and 16 outputs.
[0249] Both 8×8 NSST and 4×4 NSST are configured by a hierarchical combination of Givens rotations. The matrix corresponding to one Givens rotation is shown in Equation 4, and the matrix product is shown in Equation 5 below.
[0250] [Equation 4]
[0251]
[0252] [Equation 5]
[0253] t m = x m cosθ - x n Sinθ
[0254] t n = x m sinθ + x n cosθ
[0255] As Figure 13 illustrated in, since one Givens rotation rotates two data, in order to process 64 data (for 8×8 NSST) or 16 data (for 4×4 NSST), a total of 32 or 8 Givens rotations are required.
[0256] Therefore, a bundle of 32 or 8 is used to form the Givens rotation layer. The output data of one Givens rotation layer is transmitted as the input data of the next Givens rotation layer through the determined permutation.
[0257] Figure 14 Illustrated is one round of configuration in 4×4 NSST composed of a Givens rotation layer and a permutation as an embodiment of applying the present disclosure.
[0258] Referring to Figure 14 , illustrated is processing four Givens rotation layers in sequence in the case of 4×4 NSST. As Figure 14 illustrated in, the output data of one Givens rotation layer is transmitted as the input data of the next Givens rotation layer through the determined permutation (or shuffling).
[0259] As Figure 14 illustrated in, the pattern to be permuted is determined regularly, and in the case of 4×4 NSST, four Givens rotation layers and their corresponding permutations are combined to form one round.
[0260] In the case of 8×8 NSST, six Givens rotation layers and corresponding permutations form one round. 4×4 NSST goes through two rounds, and 8×8 NSST goes through four rounds. Different rounds use the same permutation pattern, but the Givens rotation angles applied are different. Therefore, it is necessary to store the angle data of all Givens rotations constituting each transformation.
[0261] As the last step, finally, a permutation is further performed on the data output by the Givens rotation layer, and the corresponding permutation information is stored separately for each transformation. In the forward NSST, the corresponding permutation is performed finally, while on the contrary, in the reverse NSST, the corresponding reverse permutation is applied first.
[0262] In the case of the reverse NSST, the Givens rotation layer and the permutation applied to the forward NSST are performed in the reverse order, and the angle of each Givens rotation is even rotated by taking the negative value.
[0263] Figure 15It is a block diagram for describing operations of forward and inverse simplification transforms which are embodiments of applying the present disclosure.
[0264] Reduced second-order transform (RST)
[0265] When assuming that the orthogonal matrix representing a transform has an N×N form, among N transform basis vectors (R < N), the simplification transform (hereinafter referred to as "RT") only leaves R transform basis vectors. The matrix of the forward RT for generating transform coefficients is given by Equation 6 below.
[0266] [Equation 6]
[0267]
[0268] Since the matrix of the inverse RT becomes the transpose matrix of the forward RT matrix, the applications of the forward RT and the inverse RT are exemplified as in Figure 15 . Here, the simplification factor is defined as R / N (R < N).
[0269] The number of elements of the simplification transform matrix is R×N, which is smaller than the size of the entire matrix (N×N). In other words, the required matrix is R / N of the entire matrix. Additionally, the required number of multiplications is R×N, which is lower by R / N than the original N×N. When applying the simplification transform, R coefficients are provided. As a result, only R coefficient values can be sent instead of N coefficients.
[0270] Assume a case where RT is applied to the upper left 8×8 block of a transform block that has undergone a primary transform. This RT can be referred to as an 8×8 reduced secondary transform (8×8RST).
[0271] When the R value in Equation 6 above is 16, the forward 8×8RST has a 16×64 matrix form, and the inverse 8×8RST has a 64×16 matrix form.
[0272] In addition, a transform set configuration identical to the one exemplified in Figure 12 can even be applied to the 8×8RST. That is, the corresponding 8×8RST can be applied according to the transform set in Figure 12 .
[0273] As an embodiment, when, according to the intra prediction mode in Figure 12 , a transform set is composed of two or three transforms, it can be configured to select one of up to 4 transforms including the case where no secondary transform is applied. Here, one transform can be regarded as an identity matrix.
[0274] When indices 0, 1, 2, and 3 are assigned to the four transforms respectively, a syntax element called the NSST index can be signaled for each transform block, thereby specifying the corresponding transform. That is, in the case of NSST, an 8×8 NSST can be specified for the 8×8 upper-left block through the NSST index, and an 8×8 RST can be specified in the RST configuration. Additionally, in this case, index 0 can be assigned to the case where the identity matrix, i.e., the quadratic transform, is not applied.
[0275] When the forward 8×8 RST shown in Equation 6 is applied, 16 valid transform coefficients are generated. As a result, it can be considered that the 64 input data constituting the 8×8 region are reduced to 16 output data. From the perspective of a two-dimensional region, only one-fourth of the region is filled with valid transform coefficients. Therefore, Figure 16 the 4×4 upper-left region in
[0276] Figure 16 is a diagram illustrating the process of performing a backward scan from the 64th to the 17th in the backward scan order as an embodiment of applying the present disclosure.
[0277] Figure 16 It illustrates the scan from the 17th coefficient to the 64th coefficient (in the forward scan order) when the forward scan order starts from 1. However, Figure 16 illustrates the backward scan, and it illustrates performing the backward scan from the 64th coefficient to the 17th coefficient.
[0278] Referring to Figure 16 , the upper-left 4×4 region is the region of interest (ROI), valid transform coefficients are assigned to the ROI, and the remaining regions are empty. That is, by default, the value 0 can be assigned to the remaining regions.
[0279] If there are valid transform coefficients other than 0 in the regions of Figure 16 except for the ROI region, this means that the 8x8 RST has not been applied. As a result, in this case, the corresponding NSST index coding can be omitted.
[0280] Conversely, if there are no non-zero transform coefficients outside the ROI region of Figure 16 (when applying the 8×8 RST and assigning 0 to the regions other than the ROI), it is possible that the 8×8 RST has been applied. As a result, the NSST index can be encoded.
[0281] Thus, since the presence of non-zero transform coefficients must be checked, conditional NSST index coding can be performed after the residual coding process.
[0282] The present disclosure provides an RST design method and an associated optimization method that can be applied to 4×4 blocks from an RST structure. In addition to 4×4 RSTs, the embodiments disclosed in the present disclosure can also be applied to 8×8 RSTs or another type of transform.
[0283] Figure 17 Illustrated are three forward scan orders for a transform coefficient block (transform block) that is an embodiment applying the present disclosure.
[0284] Embodiment 1: RST applicable to 4×4 blocks
[0285] An inseparable transform that can be applied to a 4×4 block is a 16×16 transform. That is, when data elements constituting a 4×4 block are arranged in row-major or column-major order, a 16×1 vector is used to apply the inseparable transform.
[0286] The forward 16×16 transform consists of 16 row-direction transform basis vectors, and when an inner product is applied to the 16×1 vector and each transform basis vector, transform coefficients for the transform basis vectors are obtained. The process of obtaining transform coefficients corresponding to all 16 transform basis vectors is equivalent to multiplying the input 16×1 vector by a 16×16 inseparable transform matrix.
[0287] The transform coefficients obtained by the matrix product have the form of a 16×1 vector, and the statistical characteristics may be different for each transform coefficient. For example, when a 16×1 transform coefficient vector is composed of the 0th element to the 15th element, the variance of the 0th element may be greater than the variance of the 15th element. In other words, as the element position is earlier, the corresponding variance value of the element is larger, such that the element can have a larger energy value.
[0288] When applying an inverse 16×16 inseparable transform to the 16×1 transform coefficients, the original 4×4 block signal can be restored. When the forward 16×16 inseparable transform is an orthogonal transform, the corresponding inverse 16×16 transform can be obtained by the transpose matrix for the forward 16×16 transform.
[0289] When the 16×1 transform coefficient vector is multiplied by the inverse 16×16 inseparable transform matrix, data in the form of a 16×1 vector can be obtained, and when the obtained data is arranged in row-major or column-major order as first applied, the 4×4 block signal can be restored.
[0290] As described above, the elements constituting the 16×1 transform coefficient vector can have different statistical characteristics.
[0291] If the transform coefficients arranged on the front side (close to the 0th element) have large energy, a signal very close to the original signal can be restored even by applying an inverse transform to some of the first-occurring transform coefficients without using all the transform coefficients. For example, when the inverse 16×16 non-separable transform is composed of 16 column basis vectors, only L column basis vectors are left to form a 16×L matrix. In addition, when the 16×L matrix and L×1 are multiplied by each other after only L important transform coefficients are left among the transform coefficients (L×1 vector), a 16×1 vector with a small error from the original input 16×1 vector data can be restored.
[0292] As a result, since only L coefficients are used for data restoration, when obtaining the transform coefficients, an L×1 transform coefficient vector is obtained instead of a 16×1 transform coefficient vector. That is, when an L×16 transform is configured by selecting L corresponding row direction vectors in the forward 16×16 non-separable transform matrix and the configured L×16 transform is multiplied by a 16×1 input vector, L important transform coefficients can be obtained.
[0293] The range of the L value is 1≤L<16, and generally, L vectors can be selected from among the 16 transform basis vectors by any method, but from the perspective of encoding and decoding, selecting the transform basis vectors with high importance in terms of signal energy can be advantageous in terms of encoding efficiency.
[0294] Embodiment 2: Configuration of the application area of 4×4 RST and arrangement of transform coefficients
[0295] 4×4 RST can be applied as a secondary transform and can be additionally applied to blocks to which a primary transform such as DCT type 2 is applied. When the size of the block to which the primary transform is applied is N×N, the size of the block to which the primary transform is applied is generally larger than 4×4. Therefore, when applying 4×4 RST to an N×N block, there are the following two methods.
[0296] Embodiment 2-1) 4×4 RST is not applied to all N×N regions, but can be applied only to certain regions. For example, 4×4 RST can be applied only to the upper-left M×M region (M≤N).
[0297] Embodiment 2-2) The region to which the secondary transform is to be applied can be divided into 4×4 blocks, and then 4×4 RST can be applied to each divided block.
[0298] As an embodiment, Embodiment 2-1) and 2-2) can be mixed and applied. For example, only the upper-left M×M region can be divided into 4×4 blocks, and then 4×4 RST can be applied.
[0299] As an implementation manner, the secondary transformation may be applied only to the upper left 8×8 region. When the N×N block is equal to or larger than 8×8, 8×8 RST may be applied. When the N×N block is smaller than 8×8 (4×4, 8×4, and 4×8), the N×N block may be divided into 4×4 blocks, and then 4×4 RST may be applied to each of the 4×4 blocks, as in Embodiment 2-2. Additionally, even in the case of 4×N / N×4 (N >= 16), 4×4 NSST may be applied.
[0300] When L (1 ≤ L < 16) transform coefficients are generated after applying 4×4 RST, degrees of freedom for arranging the L transform coefficients are generated. However, since there will be a predetermined order when processing the transform coefficients in the residual coding step, the coding performance may vary depending on the arrangement of the L transform coefficients in the 2D block.
[0301] For example, in the case of residual coding in HEVC, the coding starts from the position farthest from the DC position. This is to enhance the coding performance by taking advantage of the fact that the quantization coefficient values become zero or approach zero as they move away from the DC position.
[0302] Therefore, in terms of coding performance, it may be advantageous to arrange the more important coefficients with high energy for the L transform coefficients so that the L transform coefficients are subsequently coded in the order of residual coding.
[0303] Figure 17 Three forward scan orders in units of 4×4 transform blocks (coefficient groups (CG)) applied in HEVC are illustrated. Residual coding follows Figure 17 the reverse order of the scan order (i.e., coding is performed in the order from 16 to 1).
[0304] Since the three scan orders presented in Figure 17 are selected according to the intra prediction mode, the present disclosure may be configured to similarly determine the scan order for even the L transform coefficients according to the intra prediction mode.
[0305] Figure 18 The positions and forward scan order of the valid transform coefficients in each 4×4 block when applying diagonal scan and 4×4 RST in the upper left 4×8 block as an embodiment of applying the present disclosure are illustrated.
[0306] When following the diagonal scan order in Figure 17 and dividing the upper left 4×8 block into 4×4 blocks and applying 4×4 RST to each 4×4 block, if the L value is 8 (i.e., if only 8 out of 16 transform coefficients are left), the transform coefficients may be set as in Figure 18 .
[0307] Only half of each 4×4 block can have transform coefficients, and by default, a value of 0 can be applied to the positions marked with X.
[0308] Therefore, residual coding is performed by arranging L transform coefficients for each 4×4 block according to the Figure 17 scanning order illustrated in and assuming that the remaining (16–L) positions of each 4×4 block are filled with zeros.
[0309] Figure 19 Illustrated is a case where, when applying diagonal scanning and 4×4 RST in the upper left 4×8 block as an embodiment of applying the present disclosure, the valid transform coefficients of two 4×4 blocks are combined into one 4×4 block.
[0310] Referring to Figure 19 , the L transform coefficients arranged in two 4×4 blocks can be combined into one. Specifically, when the value of L is 8, since the transform coefficients of the two 4×4 blocks are combined while completely filling one 4×4 block, no transform coefficients are left in the other 4×4 block.
[0311] Therefore, since most residual coding is not required for an empty 4×4 block, the corresponding coded_sub_block_flag can be coded as 0.
[0312] In addition, as an embodiment of the present disclosure, various schemes can even be applied to how to mix the transform coefficients of two 4×4 blocks. The transform coefficients can be combined according to a random order, but the present disclosure can provide the following methods.
[0313] 1) The transform coefficients of two 4×4 blocks are alternately mixed in the scanning order. That is, in Figure 18 , when the transform coefficients of the top block are and and the transform coefficients of the bottom block are and , the transform coefficients can be alternately mixed one by one like and . Alternatively, the order of and can be changed. In other words, the order can be configured such that appears first.
[0314] 2) The transform coefficients for the first 4×4 block can be arranged first, and then the transform coefficients of the second 4×4 block can be arranged. In other words, the transform coefficients can be connected and arranged like and . Alternatively, the order can be changed like and .
[0315] Embodiment 3: Method for encoding the NSST index for 4×4 RST
[0316] When, as Figure 18 illustrated in the example, 4×4 RST is applied, the 16th to the (L + 1)th can be filled with 0 values according to the transform coefficient scan order of each 4×4 block.
[0317] Therefore, when non-zero values are generated at the 16th to the (L + 1)th positions even in one of two 4×4 blocks, it can be recognized that this is a case where 4×4 RST is not applied.
[0318] When 4×4 RST also has a structure in which one of the transform sets prepared for NSST is selected and applied, a transform index (which can be referred to as an NSST index in this embodiment) to which the transform is to be applied can be signaled.
[0319] It is assumed that any decoder can know the NSST index by bitstream parsing, and the parsing is performed after residual decoding.
[0320] When residual decoding is performed and it is confirmed that there is at least one non-zero transform coefficient between the (L + 1)th and the 16th, 4×4 RST is not applied, so the NSST index can be configured not to be parsed.
[0321] Therefore, the NSST index is selectively parsed only when signaling cost needs to be reduced.
[0322] If, as Figure 18 illustrated in the example, 4×4 RST is applied to a plurality of 4×4 blocks in a specific region (for example, the same 4×4 RST can be applied to all of a plurality of 4×4 blocks, or different 4×4 RSTs can be applied), the 4×4 RST applied to all 4×4 blocks can be specified by one NSST index. In this case, the same 4×4 RST can be specified, or the 4×4 RST applied to each of all 4×4 blocks can be specified.
[0323] Since it is determined whether 4×4 RST is applied to all 4×4 blocks by one NSST index, it is possible to check whether there are non-zero transform coefficients at the 16th to the (L + 1)th positions in all 4×4 blocks during the residual decoding process. As a result of the check, when non-zero transform coefficients exist at positions (the 16th to the (L + 1)th positions) that are not acceptable even in one 4×4 block, the NSST index can be configured not to be encoded.
[0324] NSST indices can be signaled separately for luminance blocks and chrominance blocks, and in the case of chrominance blocks, separate NSST indices can be signaled for Cb and Cr, and a single NSST index can be shared.
[0325] When Cb and Cr share a single NSST index, the 4×4 RST specified by the same NSST index can be applied. In this case, the 4×4 RSTs for Cb and Cr can be the same, or the NSST indices can be the same, but separate 4×4 RSTs can be provided.
[0326] To apply conditional signaling to the shared NSST index, it is checked whether there are non-zero transform coefficients at positions from the (L + 1)-th position to the 16-th position in all 4×4 blocks of Cb and Cr, and when there are non-zero transform coefficients, the NSST index can be configured not to be signaled.
[0327] As Figure 19 illustrated, even for the case where the transform coefficients of two 4×4 blocks are combined, when applying the 4×4 RST, it is checked whether there are non-zero transform coefficients at positions where there are no valid transform coefficients, and then it can be determined whether to signal the NSST.
[0328] For example, as Figure 19 illustrated in (b) of, when the L value is 8 and there are no valid transform coefficients in a 4×4 block (the block marked with X) when applying the 4×4 RST, the coded_sub_block_flag of the block where there are no valid transform coefficients can be checked. In this case, when the coded_sub_block_flag is 1, the NSST index can be configured not to be signaled.
[0329] Embodiment 4: Optimization method for the case of encoding the NSST index before residual encoding
[0330] When encoding the NSST index is performed before residual coding, whether to apply the 4×4 RST is predetermined, and as a result, residual coding can be omitted for positions where 0 is assigned to the transform coefficients.
[0331] Here, whether to apply the 4×4 RST can be configured to be known from the NSST index. For example, when the NSST index is 0, the 4×4 RST is not applied.
[0332] Alternatively, the NSST index can be signaled by a separate syntax element (e.g., NSST flag). For example, if the separate syntax element is called the NSST flag, the NSST flag is first parsed to determine whether to apply the 4×4 RST, and if the NSST flag value is 1, residual coding can be omitted for positions where valid transform coefficients cannot exist.
[0333] As an implementation, when performing residual coding, first encode the position of the last non-zero transform coefficient on the TU. When encoding the NSST index is performed after encoding the position of the last non-zero transform coefficient and assuming that 4×4RST is applied to the position of the last non-zero transform coefficient, if the position of the last non-zero transform coefficient is determined to be a position where no non-zero transform coefficient can be generated, 4×4RST can be configured not to be applied to the position of the last non-zero transform coefficient without encoding the NSST index.
[0334] For example, since in the case of the positions marked with X in Figure 18 , no valid transform coefficients are set when 4×4RST is applied (for example, these positions can be filled with zero values), so when the last non-zero transform coefficient is in the area marked with X, encoding of the NSST index can be omitted. When the last non-zero transform coefficient is not in the area marked with X, encoding of the NSST index can be performed.
[0335] As an implementation, when checking whether to apply 4×4RST by conditionally encoding the NSST index after encoding the position of the last non-zero transform coefficient, the remaining residual coding part can be processed by the following two schemes.
[0336] 1) Without applying 4×4RST, the general residual coding is kept as it is. That is, encoding is performed assuming that non-zero transform coefficients may exist at any position from the non-zero transform coefficient position to DC.
[0337] 2) When 4×4RST is applied, since there are no transform coefficients at specific positions or specific 4×4 blocks (for example, the X positions in Figure 18 , which can be filled with 0 by default), the residual for the corresponding position or block can be not performed.
[0338] For example, in the case of reaching the positions marked with X in Figure 18 , encoding of sig_coeff_flag can be omitted. Here, sig_coeff_flag means a flag indicating whether there is a non-zero transform coefficient at the corresponding position.
[0339] When, as exemplified in Figure 19 , the transform coefficients of two blocks are combined, encoding of coded_sub_block_flag can be omitted for the 4×4 blocks assigned 0, and the corresponding value can be deduced as 0, and all corresponding 4×4 blocks can be deduced as 0 values without separate encoding.
[0340] In the case of encoding the NSST index after encoding the positions of non-zero transform coefficients, when the x-position P x and y-position P y of the last non-zero transform coefficient are less than T x and T y respectively, the NSST index encoding can be omitted and the 4×4 RST may not be applied.
[0341] For example, the case where T x = 1 and T y = 1 means that for the case where there is a non-zero transform coefficient at the DC position, the NSST index encoding is omitted.
[0342] The scheme for determining whether to encode the NSST index by comparing with a threshold can be applied differently to luminance and chrominance. For example, different Ts x and Ts y can be applied to luminance and chrominance, and a threshold can be applied to luminance but not to chrominance. Or vice versa.
[0343] The above two methods can be applied simultaneously, that is, the first method of omitting the NSST index encoding when the non-zero transform coefficient is in a region where there is no valid transform coefficient, and the second method of omitting the NSST index encoding when each of the X coordinate and Y coordinate of the non-zero transform coefficient is less than a predetermined threshold.
[0344] For example, the threshold of the position coordinates of the last non-zero transform coefficient can be checked first, and then it can be checked whether the last non-zero transform coefficient is in a region where there is no valid transform coefficient. Alternatively, the order can be changed.
[0345] The method proposed in Embodiment 4 can even be applied to 8×8 RST. That is, when the last non-zero transform coefficient is in a region of the upper left 8×8 area other than the upper left 4×4, the encoding of the NSST index can be omitted; otherwise, the NSST index encoding can be performed.
[0346] In addition, when both the X and Y coordinate values for the non-zero transform coefficient are less than the threshold, the encoding of the NSST index can be omitted. Alternatively, the two methods can be applied together.
[0347] Embodiment 5: Applying different NSST index encoding and residual encoding schemes to luminance and chrominance when applying RST
[0348] The schemes described in Embodiments 3 and 4 can be applied differently to luminance and chrominance respectively. That is, the NSST index encoding and residual encoding schemes for luminance and chrominance can be applied differently.
[0349] For example, the luminance can adopt the solution of Embodiment 4 and the chrominance can adopt the solution of Embodiment 3. Alternatively, the luminance can adopt the conditional NSST index coding proposed in Embodiment 3 or 4, and the chrominance can not adopt the conditional NSST index coding. Or vice versa.
[0350] Figure 20 is a flowchart for encoding a video signal based on a simplified quadratic transform as an embodiment of applying the present disclosure.
[0351] The encoder can determine (or select) a forward quadratic transform (S2010) based on at least one of the prediction mode, block shape, and / or block size of the current block. In this case, candidates for the forward quadratic transform can include Figure 6 and / or Figure 12 at least one of the embodiments of.
[0352] The encoder can determine the optimal forward quadratic transform through rate-distortion optimization. The optimal forward quadratic transform can correspond to one of a plurality of transform combinations, and these plurality of transform combinations can be defined by a transform index. For example, for RD optimization, the overall results of performing the forward quadratic transform, quantization, residual coding, etc. can be compared for each candidate. In this case, formulas such as cost = rate + λ·distortion or cost = distortion + λ·rate can be used, but the present disclosure is not limited thereto.
[0353] The encoder can signal the quadratic transform index corresponding to the optimal forward quadratic transform (S2020). Here, the quadratic transform index can adopt other embodiments described in the present disclosure.
[0354] For example, the quadratic transform index can adopt Figure 12 the transform set configuration of. Since, according to the intra prediction mode, one transform set consists of two or three transforms, in addition to the case where no quadratic transform is applied, one of up to four transforms can be configured to be selected. When indices 0, 1, 2, and 3 are assigned to the four transforms respectively, the applied transform can be specified by signaling the quadratic transform index for each transform coefficient block. In this case, index 0 can be assigned to the case where no unit matrix, i.e., quadratic transform, is applied.
[0355] As another embodiment, signaling of the quadratic transform index can be performed in any one of the following steps: 1) before residual coding; 2) in the middle of residual coding (after coding the positions of non-zero transform coefficients); or 3) after residual coding. Hereinafter, the embodiments will be described in detail.
[0356] 1) Method of signaling the quadratic transform index before residual coding
[0357] The encoder may determine a forward secondary transform.
[0358] The encoder may signal a secondary transform index corresponding to the forward secondary transform.
[0359] The encoder may encode the position of the last non-zero transform coefficient.
[0360] The encoder may perform residual coding on syntax elements other than the position of the last non-zero transform coefficient.
[0361] 2) Method for Signaling a Secondary Transform Index in the Middle of Residual Coding
[0362] The encoder may determine a forward secondary transform.
[0363] The encoder may encode the position of the last non-zero transform coefficient.
[0364] When the non-zero transform coefficients are not in a specific region, the encoder may encode a secondary transform index corresponding to the forward secondary transform. Here, in the case of applying a simplified secondary transform, when the transform coefficients are arranged according to the scan order, the specific region represents the remaining region except for the positions where non-zero transform coefficients may exist. However, the present disclosure is not limited thereto.
[0365] The encoder may perform residual coding on syntax elements other than the position of the last non-zero transform coefficient.
[0366] 3) Method for Signaling a Secondary Transform Index Before Residual Coding
[0367] The encoder may determine a forward secondary transform.
[0368] The encoder may encode the position of the last non-zero transform coefficient.
[0369] When the non-zero transform coefficients are not in a specific region, the encoder may perform residual coding on syntax elements other than the position of the last non-zero transform coefficient. Here, in the case of applying a simplified secondary transform, when the transform coefficients are arranged according to the scan order, the specific region represents the remaining region except for the positions where non-zero transform coefficients may exist. However, the present disclosure is not limited thereto.
[0370] The encoder may encode a secondary transform index corresponding to the forward secondary transform.
[0371] In addition, the encoder may perform a forward primary transform (S2030) on the current block (residual block). Here, step S2010 and / or step S2020 may be similarly applied to the forward primary transform.
[0372] The encoder can perform a forward quadratic transform on the current block by using an optimal forward quadratic transform (S2040). For example, the optimal forward quadratic transform can be a simplified quadratic transform. The simplified quadratic transform refers to a transform that takes in N residual data (N×1 residual vector) and outputs L (L < N) transform coefficient data (L×1 transform coefficient vector).
[0373] As an implementation, the simplified quadratic transform can be applied to a specific region of the current block. For example, when the current block is N×N, the specific region can refer to the upper left N / 2×N / 2 region. However, the present disclosure is not limited thereto, and different configurations can be made according to at least one of the prediction mode, block shape, or block size. For example, when the current block is N×N, the specific region can refer to the upper left M×M region (M ≤ N).
[0374] In addition, the encoder performs quantization on the current block to generate a transform coefficient block (S2050).
[0375] The encoder performs entropy coding on the transform coefficient block to generate a bitstream.
[0376] Figure 21 is a flowchart for decoding a video signal based on a simplified quadratic transform as an implementation of applying the present disclosure.
[0377] The decoder can obtain a quadratic transform index from the bitstream (S2110). Here, the quadratic transform index can adopt other implementations described in the present disclosure. For example, the quadratic transform index can include Figure 6 and / or Figure 12 at least one of the implementations of.
[0378] As another implementation, the obtaining of the quadratic transform index can be performed in any of the following steps: 1) before residual coding; 2) in the middle of residual coding (after decoding the non-zero transform coefficient positions); or 3) after residual coding.
[0379] The decoder can derive the quadratic transform corresponding to the quadratic transform index (S2120). In this case, the candidates for the forward quadratic transform can include Figure 6 and / or Figure 12 at least one of the implementations of.
[0380] However, steps S2110 and S2120 are implementations, and the present disclosure is not limited thereto. For example, the decoder cannot obtain the quadratic transform index, but instead derives the quadratic transform based on at least one of the prediction mode, block shape, and / or block size of the current block.
[0381] In addition, the decoder can obtain the transform coefficient block by performing entropy decoding on the bitstream, and can perform inverse quantization on the transform coefficient block (S2130).
[0382] The decoder may perform an inverse quadratic transform (S2140) on the quantized inverse transform coefficient block. For example, the inverse quadratic transform may be a simplified quadratic transform. The simplified quadratic transform refers to a transform that inputs N residual data (N×1 residual vector) and outputs L (L < N) transform coefficient data (L×1 transform coefficient vector).
[0383] As an implementation, the simplified quadratic transform may be applied to a specific region of the current block. For example, when the current block is N×N, the specific region may refer to the upper left N / 2×N / 2 region. However, the present disclosure is not limited thereto, and different configurations may be made according to at least one of the prediction mode, block shape, or block size. For example, when the current block is N×N, the specific region may refer to the upper left M×M (M≤N) or M×L (M≤N, L≤N) region.
[0384] In addition, the decoder may perform an inverse primary transform on the inverse quadratic transform result (S2150).
[0385] The decoder generates a residual block through step S2150, and the residual block and the prediction block are added to generate a reconstructed block.
[0386] Embodiment 6: Method for configuring a hybrid second-order transform set
[0387] In an implementation, a method for configuring a set of quadratic transforms by considering various conditions when applying the quadratic transform is proposed.
[0388] In the present disclosure, the quadratic transform indicates a transform performed on all or some of the primary transform coefficients after applying the primary transform on the encoder side as described above, and the quadratic transform may be referred to as an inseparable quadratic transform (NSST), a low-frequency inseparable transform (LFNST), etc. After applying the quadratic transform to all or some of the quantized inverse transform coefficients, the decoder may apply the primary transform to all or some of the quantized inverse transform coefficients.
[0389] In addition, in the present disclosure, the set of hybrid quadratic transforms represents a set of transforms applicable to the quadratic transform, and the present disclosure is not limited to such a name. For example, the set of hybrid quadratic transforms (or transform kernels or transform types) may be referred to as a quadratic transform group, a quadratic transform table, a quadratic transform candidate, a quadratic transform candidate list, a hybrid quadratic transform group, a hybrid quadratic transform table, a hybrid quadratic transform candidate, a hybrid quadratic transform candidate list, etc., and the set of hybrid quadratic transforms may include multiple transform kernels (or transform types).
[0390] A secondary transformation can be applied to a sub-block with a specific size in the upper left side in the current block according to predefined conditions. In conventional video compression techniques, a 4×4 secondary transformation set or an 8×8 secondary transformation set is used according to the size of the selected sub-block. In this case, the 4×4 secondary transformation set only includes transformation kernels (hereinafter referred to as 4×4 transformation kernels) applied to a 4×4-sized area (or block), and the 8×8 secondary transformation set only includes transformation kernels (hereinafter referred to as 8×8 transformation kernels) applied to an 8×8-sized area (or block). In other words, in conventional video compression techniques, the secondary transformation set is composed of transformation kernels with a limited size only according to the size of the area to which the secondary transformation is applied.
[0391] Therefore, the present disclosure proposes a hybrid secondary transformation set including transformation kernels that can be applied to areas with various sizes.
[0392] As an implementation, the size of the transformation kernels included in the hybrid secondary transformation set (i.e., the size of the area to which the corresponding transformation kernel is applied) is not fixed, but can be variably determined (or set). For example, the hybrid secondary transformation set can include 4×4 transformation kernels and 8×8 transformation kernels.
[0393] In addition, in an implementation, the number of transformation kernels included in the hybrid secondary transformation set is not fixed, but can be variably determined (or set). In other words, the hybrid secondary transformation set can include multiple transformation sets, and each transformation set can include a different number of transformation kernels. For example, the first transformation set can include three transformation kernels, and the second transformation set can include four transformation kernels.
[0394] In addition, in an implementation, the order (or priority) between the transformation kernels included in the hybrid secondary transformation set is not fixed, but can be variably determined (or set). In other words, the hybrid secondary transformation set can include multiple transformation sets, and the order between the transformation kernels in each transformation set can be independently defined. In addition, different indexes can be mapped (or assigned) between the transformation kernels in each transformation set. For example, assuming that the first transformation set and the second transformation set each include a first transformation kernel, a second transformation kernel, and a third transformation kernel, the index values 1, 2, and 3 can be mapped to the first transformation kernel, the second transformation kernel, and the third transformation kernel in the first transformation set respectively, and the index values 3, 2, and 1 can be mapped to the first transformation kernel, the second transformation kernel, and the third transformation kernel in the second transformation set respectively.
[0395] Hereinafter, a method for determining the priority (or order) between the transformation kernels in the hybrid secondary transformation set will be described in detail.
[0396] When applying a secondary transform, the decoder may determine (or select) a set of secondary transforms to be applied to the current processing block according to predefined conditions. The set of secondary transforms may include a hybrid secondary transform set according to embodiments of the present disclosure. Additionally, the decoder may use a secondary transform index signaled from the encoder to determine a transform kernel of the secondary transform to be applied to the current processing block within the determined set of secondary transforms. The secondary transform index may indicate a transform kernel of the secondary transform to be applied to the current processing block within the determined set of secondary transforms. In one embodiment, the secondary transform index may be referred to as an NSST index, an LFNST index, etc.
[0397] As described above, since the secondary transform index is signaled from the encoder to the decoder, it is reasonable in terms of compression efficiency to assign lower indices to relatively more frequently occurring transform kernels for encoding / decoding with fewer bits. Accordingly, various embodiments of configuring a set of secondary transforms considering priorities will be described below.
[0398] In an embodiment, priorities among transform kernels in the set of secondary transforms may be determined differently according to the size of a region (or sub-block) to which the secondary transform is applied. For example, when the size of the current processing block is equal to or greater than a predetermined size, an 8×8 transform kernel may be used more frequently. As a result, the encoder / decoder may assign relatively fewer transform indices to the 8×8 transform kernel than to the 4×4 transform kernel. As an example, a hybrid secondary transform set may be configured as in Table 3 below.
[0399] [Table 3]
[0400] NSST index 4×4 NSST set 8×8 NSST set Hybrid NSST set 1 4×4 first kernel 8×8 first kernel 8×8 first kernel 2 4×4 second kernel 8×8 second kernel 8×8 second kernel 3 4×4 third kernel 8×8 third kernel 4×4 first kernel ... ... ... ...
[0401] In Table 3, an embodiment is described by assuming the use of NSST as the secondary transform, but the present disclosure is not limited to such a name. As described above, the secondary transform according to embodiments of the present disclosure may be referred to as an inseparable secondary transform (NSST), a low-frequency inseparable transform (LFNST), etc. Referring to Table 3, the set of secondary transforms of conventional video compression techniques consists only of transform kernels of the size of the region to which the secondary transform is applied (i.e., the 4×4 NSST set and the 8×8 NSST set in Table 3). The hybrid secondary transform set according to embodiments of the present disclosure (i.e., the hybrid NSST set) may include an 8×8 transform kernel and a 4×4 transform kernel.
[0402] In Table 3, it is assumed that the size of the current processing block is equal to or greater than a predetermined size. That is, when the minimum value of the width or height of the current block is equal to or greater than a predefined value (e.g., 8), the secondary transform set illustrated in Table 3 may be used. The hybrid secondary transform set may include an 8×8 transform kernel and a 4×4 transform kernel, and the 8×8 transform kernel is likely to be used. As a result, a low index value may be assigned to the 8×8 transform kernel.
[0403] In addition, in an embodiment, the priority among the transform kernels in the secondary transform set may be determined based on the order of the secondary transform kernels (first, second, and third). For example, the first 4×4 secondary transform kernel may have a higher priority than the second 4×4 secondary transform kernel, and a lower index value may be assigned to the first 4×4 secondary transform kernel. As an example, the hybrid secondary transform set may be configured as shown in Table 4 below.
[0404] [Table 4]
[0405]
[0406] In Table 4, the embodiment is described by assuming the use of NSST as the secondary transform, but the present disclosure is not limited to such a name. As described above, the secondary transform according to the embodiment of the present disclosure may be referred to as NSST, LFNST, etc.
[0407] Referring to Table 4, the hybrid secondary transform set (i.e., hybrid NSST set types 1, 2, and 3) according to the embodiment of the present disclosure may include 8×8 transform kernels and / or 4×4 transform kernels. In other words, like hybrid NSST set types 2 and 3, the hybrid secondary transform set may include 8×8 transform kernels and / or 4×4 transform kernels, and the priority among the 8×8 transform kernels and the priority among the 4×4 transform kernels may be configured based on the order of the corresponding secondary transform kernels.
[0408] Embodiment 7: Method for configuring a hybrid second-order transform set
[0409] In an embodiment of the present disclosure, a method for determining a secondary transform set by considering various conditions is proposed. Specifically, a method for determining a secondary transform set based on the size of a block and / or an intra prediction mode is proposed.
[0410] The encoder / decoder may configure a transform set suitable for the secondary transform of the current block based on the intra prediction mode. As an embodiment, the proposed method may be applied together with the above-described Embodiment 6. For example, the encoder / decoder may configure a secondary transform set based on the intra prediction mode and perform a secondary transform using transform kernels of various sizes included in each secondary transform set.
[0411] In an embodiment, the secondary transform set may be determined based on Table 5 below according to the intra prediction mode.
[0412] [Table 5]
[0413]
[0414] Referring to Table 5, the encoder / decoder can determine whether to apply (or configure) the hybrid quadratic transform set described in Embodiment 6 above based on the intra prediction mode. When the hybrid quadratic transform set is not applied, the encoder / decoder can apply (or configure) the quadratic transform set described above Figures 12 to 14 in the description.
[0415] Specifically, in the case of the intra prediction mode where the hybrid type value is defined as 1, the hybrid quadratic transform set can be configured according to the method described in Embodiment 6 above. Additionally, in the case of the intra prediction mode where the hybrid type value is defined as 0, the quadratic transform set can be configured according to a conventional method (i.e., the method described above Figures 12 to 14 in the description).
[0416] Table 5 shows the case of using the method for configuring two types of transform sets, but the present disclosure is not limited thereto. That is, two or more types of hybrid types indicating the method for configuring the transform set including the hybrid quadratic transform set can be configured (or defined). In this case, the hybrid type information can be configured as various values of N types (N > 2).
[0417] Furthermore, in the embodiment, the encoder / decoder can determine whether the transform set suitable for the current block is configured as a normal type or a hybrid type by considering the size of the current coding block (or transform block) and the intra prediction mode. Here, the normal type indicates the quadratic transform set configured according to a conventional method (i.e., the method described above Figures 12 to 14 in the description). For example, when the hybrid type (or mode type) value is 0, the encoder / decoder can configure the quadratic transform set by applying the method described Figures 12 to 14 in the description, and when the hybrid type value is 1, the encoder / decoder can configure the hybrid quadratic transform set including transform kernels of various sizes according to the size of the current block.
[0418] Figure 22 is a diagram illustrating a method for determining the transform type applied to the quadratic transform according to an embodiment of applying the present disclosure.
[0419] Referring to Figure 22 , the decoder is mainly described for convenience of description, but the present disclosure is not limited thereto, and even the method of performing the quadratic transform by determining the transform type can be applied to the encoder substantially equivalently.
[0420] The decoder performs inverse quantization (S2201) on the quantized transform block (or transform coefficients).
[0421] The decoder determines (or selects) a set of secondary transforms (or a set of transform types) for the current block for secondary transform based on the size of the current block and / or the intra prediction mode (S2202). In this case, various sets of secondary transforms can be predefined, and the method of configuring the set of secondary transforms described in this embodiment and / or Embodiment 6 can be applied.
[0422] As an embodiment, the decoder can determine whether to use a hybrid type (or a hybrid set of secondary transforms) to configure the set of secondary transforms based on the intra prediction mode of the current block. When using the hybrid type, a hybrid set of secondary transforms according to the method described in Embodiment 6 above can be used.
[0423] The decoder performs a secondary transform on the inverse quantized transform block (or transform coefficients) by using the determined transform kernel (S2203). In this case, the decoder can parse a secondary transform index indicating the transform kernel for the secondary transform of the current block within the set of secondary transforms determined in step S2202. In this case, the step of parsing the secondary transform index can be included in step S2203.
[0424] Embodiment 8: Method for encoding the second-order transform index
[0425] In an embodiment, a method for efficiently encoding / decoding a secondary transform index signaled from an encoder in the case of configuring a set of secondary transforms based on the size of an encoded block (or transform block) and / or the intra prediction mode is proposed.
[0426] As described above, when applying a secondary transform, the encoder / decoder can determine (or select) a set of secondary transforms applied to the current encoded block according to preconfigured conditions. In addition, the decoder can use the secondary transform index signaled from the encoder to derive the transform kernel applied to the current encoded block within the determined set of secondary transforms. Here, the secondary transform index represents a syntax element indicating the transform kernel of the secondary transform applied to the current block within the set of secondary transforms. In the present disclosure, the secondary transform index can be referred to as an NSST index, an LFNST index, etc.
[0427] As described in Embodiment 6 above, the encoder / decoder can configure a set of secondary transforms by using transform kernels of various sizes, and the number of transform kernels included in the set of secondary transforms is not fixed but can be variably determined.
[0428] Therefore, since the number of available transform kernels can be different for each set of quadratic transforms, the encoder / decoder can binarize the quadratic transform indices by using the truncated unary method for efficient binarization. As an implementation, the encoder / decoder can perform truncated unary binarization on the quadratic transform indices according to the maximum available value of the quadratic transform indices for each set of quadratic transforms by using Table 6 below.
[0429] [Table 6]
[0430]
[0431]
[0432] In Table 6, the implementation is described by assuming the use of NSST as the quadratic transform, but the present disclosure is not limited to such a name. As described above, the quadratic transform according to the implementation of the present disclosure can be referred to as an inseparable quadratic transform (NSST), a low-frequency inseparable transform (LFNST), etc.
[0433] Referring to Table 6, the NSST indices can be binarized by the truncated unary binarization method. In this case, the binarization length can be determined according to the maximum index value within the set of quadratic transforms.
[0434] In an implementation, the quadratic transform indices in Table 6 are encoded / decoded based on context, and in this case, the following variables can be considered for context modeling.
[0435] - The size of the coding block (or transform block)
[0436] - The intra prediction mode
[0437] - The hybrid type value
[0438] - The quadratic transform index value of the set of quadratic transforms
[0439] Embodiment 9: Simplified transform
[0440] In the implementation of the present disclosure, various implementations of the simplified transform proposed in Figures 15 to 22 are described to improve the complexity problem of the transform. As described above, the simplified transform proposed in the present disclosure can be applied regardless of whether it is a primary transform (e.g., DCT, DST) or a quadratic transform (e.g., NSST, low-frequency inseparable transform (LFNST)).
[0441] Figure 23 is a diagram illustrating a simplified transform structure based on a simplification factor as an application of the implementation of the present disclosure.
[0442] Referring to Figure 23, the decoder is mainly described for convenience of description, but the simplified transform proposed in the embodiments can be equivalently applied to the encoder.
[0443] The decoder can apply an inverse simplified transform to the inverse-quantized transform coefficients. In this case, the decoder can use a predetermined (or predefined) simplification factor (e.g., R or R / N) and / or a transform kernel to perform the simplified transform.
[0444] In one embodiment, a transform kernel can be selected based on available information such as the size (e.g., width / height) of the current block (encoded block or transform block), the intra / inter-frame prediction mode, CIdx, etc. When the current encoded block is a luminance block, the value of CIdx can be 0. Otherwise (i.e., if the current encoded block is a Cb or Cr block), CIdx can have a non-zero value such as 1.
[0445] Figure 24 is a diagram illustrating a method of performing decoding by adaptively applying a simplified transform, which is an embodiment to which the present disclosure can be applied.
[0446] Referring to Figure 24 , the decoder is mainly described for convenience of description, but the method of performing a transform using a simplified transform proposed in the embodiments can be equivalently applied to the encoder.
[0447] The decoder performs inverse quantization on the current block (S2401).
[0448] The decoder checks whether a transform is applied to (or used for) the current block (S2402). If no transform is applied to the current block, the decoder terminates the transform process.
[0449] When a transform is applied to the current block, the decoder parses a transform index indicating the transform kernel applied to the current block from the video signal (S2403).
[0450] The decoder checks whether the simplified inverse transform condition is satisfied (S2404). If the simplified inverse transform condition is not satisfied, the decoder performs a normal inverse transform on the current block (S2405). If the simplified inverse transform condition is satisfied, the decoder performs a simplified inverse transform on the current block (S2407). In this case, the decoder can select the transform kernel applied to the current block based on the transform index parsed in step S2403 (S2406). As an embodiment, a transform kernel can be selected based on available information such as the size (e.g., width / height) of the current block (encoded block or transform block), the intra / inter-frame prediction mode, CIdx, etc. In addition, when a simplified inverse transform is applied to the current block, step S2406 can include selecting a simplification factor.
[0451] In an embodiment, a simplified inverse transform condition can be applied to the above condition 6) (e.g., Tables 3 and 4). In other words, it can be determined whether to apply the simplified inverse transform based on the size of the current block (encoding block or transform block) and the transform type (or transform kernel).
[0452] As an embodiment, when the following specific conditions are met, a simplified transform can be used. In other words, a simplified transform can be applied to a block having a predetermined size or larger size (or greater than a predetermined size) that satisfies the following specific conditions.
[0453] - Width > TH && Height > TH (where TH is a predefined value representing a specific threshold (e.g., 4))
[0454] Or
[0455] - Width × Height > K && MIN(Width, Height) > TH (where K or TH is a predefined value and represents a specific threshold)
[0456] As another embodiment, when the following specific conditions are met, a simplified transform can be used. In other words, a simplified transform can be applied to a block having a predetermined size or smaller size (or less than a predetermined size) that satisfies the following specific conditions.
[0457] - Width <= TH && Height <= TH (where TH is a predefined value representing a specific threshold (e.g., 8))
[0458] Or
[0459] - Width × Height <= K && MIN(Width, Height) <= TH (where K or TH is a predefined value and represents a specific threshold)
[0460] As another example, a simplified transform can be applied only to a predetermined block group.
[0461] - Width == TH && Height == TH
[0462] Or
[0463] - Width == Height
[0464] As an embodiment, if the usage conditions for the simplified transform are not met, a normal transform can be applied. Specifically, the normal transform can be predefined and available for the encoder / decoder. Examples of the normal transform are shown below.
[0465] - DCT2, DCT4, DCT5, DCT7, DCT8
[0466] Or
[0467] - DST1, DST4, DST7
[0468] or
[0469] - Inseparable transform
[0470] or
[0471] - NSST(HyGT)
[0472] or
[0473] - LFNST(Low - frequency inseparable transform)
[0474] The above conditions can be interpreted based on logical operators, as shown in Table 7 below.
[0475] [Table 7]
[0476]
[0477] In addition, as Figure 24 illustrated in, the simplified transform condition can depend on the transform index Transform_idx indicating the transform applied to the current block. As an example, Transform_idx can be sent from the encoder to the decoder twice. One can be the horizontal transform index Transform_idx_h, and the other can be the vertical transform index Transform_idx_v.
[0478] Figure 25 is a diagram illustrating a method of performing decoding by adaptively applying a simplified transform, which can be an embodiment of the present disclosure.
[0479] Referring to Figure 25 , the decoder is mainly described for convenience of description, but the method of performing a transform using a simplified transform proposed in the embodiment can be equivalently applied to the encoder.
[0480] In an embodiment of the present disclosure, the above - mentioned simplified transform can be used for a secondary transform. In this regard, the description repeated with the method described in Figure 24 will be omitted.
[0481] 1) The decoder performs inverse quantization on the current block and then checks whether NSST is activated in the current block. The decoder can use a predefined condition to determine whether it is necessary to parse the NSST index.
[0482] 2) If NSST is activated, the decoder parses the NSST index and checks whether to apply a simplified secondary inverse transform.
[0483] 3) The decoder checks whether the simplified inverse transform condition is satisfied.
[0484] 4) If the simplified inverse transform condition is not satisfied, the decoder performs a normal inverse transform on the current block.
[0485] 5) If the simplified quadratic inverse transform condition is satisfied, the decoder performs the simplified quadratic inverse transform on the current block.
[0486] 6) In this case, the decoder can select the transform kernel to be applied to the current block based on the NSST index. As an implementation, the transform kernel can be selected based on available information such as the size (e.g., width / height) of the current block (encoded block or transform block), intra / inter prediction mode, CIdx, etc. Additionally, when applying the simplified quadratic inverse transform to the current block, the decoder can select the simplification factor.
[0487] In an implementation, it can be determined whether to apply the simplified inverse transform based on the size of the current block (encoded block or transform block) and the transform type (or transform kernel).
[0488] Embodiment 10: Simplified transform as a second-order transform with different block sizes
[0489] In an implementation of the present disclosure, a simplified transform considering various block sizes for quadratic transform / quadratic inverse transform is proposed. As an example, simplified transforms for different block sizes of 4×4, 8×8, and 16×16 for quadratic transform / quadratic inverse transform can be defined.
[0490] Figure 26 and Figure 27 is a diagram illustrating examples of the forward simplified quadratic transform and the inverse simplified quadratic transform and the pseudo-code for deriving them.
[0491] Referring to Figure 27 and 28 , the simplified quadratic transform and the simplified quadratic inverse transform are shown when the block to which the quadratic transform is applied is an 8×8 block and the simplification coefficient R is 16. The simplified quadratic transform and the simplified quadratic inverse transform illustrated in Figure 27 can be derived by using the pseudo-code illustrated in Figure 26 .
[0492] Embodiment 11: Simplified transform as a second-order transform with a non-rectangular shape
[0493] As described above, due to the complexity problem of applying the quadratic transform of the non-separable transform, in the image compression technology of the related art, the quadratic transform is applied to the upper left 4×4 or 8×8 region of the encoded block (or transform block).
[0494] The implementation of the present disclosure proposes a method of applying the simplified quadratic transform to various non-square graphs in addition to the 4×4 or 8×8 square regions.
[0495] Figure 28 is a diagram illustrating a method of applying the simplified quadratic transform to a non-square region as an implementation of the present disclosure.
[0496] Refer to Figure 28 In an embodiment, a simplified quadratic transform can be applied to only a part of a block, as exemplified in Figure 29 .
[0497] In Figure 28 , each square represents a 4×4 region. Thus, the encoder / decoder can apply the simplified quadratic transform to a 10×4 pixel, i.e., 160 pixel region. In this case, the simplification factor R = 16 and the entire RST matrix corresponds to a 16×160 matrix, thereby reducing the computational complexity of applying the quadratic transform.
[0498] Embodiment 12: Simplification factor
[0499] Figure 29 is a diagram illustrating a simplified transform controlled by a simplification factor as an embodiment of applying the present disclosure.
[0500] Refer to Figure 29 , as described above, the simplified transform according to an embodiment of the present disclosure can be controlled by a simplification factor, as exemplified in Figure 29 .
[0501] Specifically, modifying the simplification factor can modify the memory complexity and the number of multiplication operations. As previously mentioned as the simplification factor R / N in Figure 15 and Equation 6, memory and multiplication can be reduced by modifying the simplification factor. For example, for an 8×8 NSST with R = 16, memory and multiplication can be reduced by 1 / 4.
[0502] Embodiment 13: High-level syntax
[0503] Embodiments of the present disclosure propose a high-level syntax structure for controlling a simplified transform at a high level.
[0504] In an embodiment, as shown in the example of Table 8 below, information on whether to accept the simplified transform and on the size and / or simplification factor can be sent through a sequence parameter set (SPS). However, the present disclosure is not limited thereto, and the syntax can be signaled through a picture parameter set (PPS), a slice header, etc.
[0505] [Table 8]
[0506]
[0507]
[0508] Referring to Table 8, if Reduced_transform_enabled_flag is 1, the reduced transform may be available and applied. In the case where Reduced_transform_enabled_flag is 0, this case may indicate that the reduced transform is not available. If Reduced_transform_enabled_flag does not exist, the value may be estimated to be equal to zero.
[0509] Reduced_transform_factor represents a syntax element that specifies the number of reduced dimensions used for the reduced transform.
[0510] min_reduced_transform_size represents a syntax element that specifies the minimum transform size to which the reduced transform will be applied. If min_reduced_transform_size does not exist, the value may be estimated to be equal to zero.
[0511] max_reduced_transform_size represents a syntax element that specifies the maximum transform size to which the reduced transform will be applied. If max_reduced_transform_size does not exist, the value may be estimated to be equal to zero.
[0512] Reduced_transform_factor represents a syntax element that specifies the number of reduced dimensions used for the reduced transform. If Reduced_transform_factor does not exist, the value may be estimated to be equal to zero.
[0513] Embodiment 14: Second-order transform kernel
[0514] Embodiments of the present disclosure propose various secondary transform kernels.
[0515] In an embodiment, the 4×4 NSST kernel for the DC mode may be defined as shown in Table 9 below.
[0516] [Table 9]
[0517]
[0518]
[0519] Additionally, in an embodiment, the 4×4 NSST kernel for the planar mode may be defined as shown in Table 10 below.
[0520] [Table 10]
[0521]
[0522]
[0523] In addition, in the embodiments, an 8×4 NSST kernel for the DC mode can be defined as shown in Table 11 below.
[0524] [Table 11]
[0525]
[0526]
[0527]
[0528]
[0529]
[0530]
[0531] In addition, in the embodiments, an 8×8 NSST kernel for the planar mode can be defined as shown in Table 12 below.
[0532] [Table 12]
[0533]
[0534]
[0535]
[0536]
[0537]
[0538]
[0539] The transform kernels of Tables 9 to 12 above can be defined as smaller transform kernels for simplified transforms.
[0540] For example, for an 8×8 NSST with the DC mode and R = 8, the memory and multiplications can be reduced by 1 / 2. Thus, by keeping only the coefficients in the upper half of the transform kernel of Table 7 above (8×16 matrix), a simplified transform kernel can be defined in a smaller size as shown in Table 13.
[0541] [Table 13]
[0542]
[0543] The transform kernels of Tables 9 to 12 above can be defined as smaller transform kernels for simplified transforms.
[0544] For example, for the 4×4 NSST in DC mode with R = 8, the memory and multiplication can be reduced by 1 / 2. Thus, by keeping only the coefficients in the upper half of the transform kernel in Table 7 above (8×16 matrix), a simplified transform kernel can be defined with a smaller size as shown in Table 13.
[0545] [Table 14]
[0546]
[0547]
[0548]
[0549]
[0550]
[0551]
[0552] In the above example, each transform coefficient is represented by 9 bits (i.e., 1 bit: sign; 8 bits: absolute value from 0 to 255). In embodiments of the present disclosure, various precisions can be used to represent transform coefficients. For example, 8 bits can be used instead of 9 bits to represent each coefficient. In this case, the sign bit remains unchanged, but the range of the absolute value can change.
[0553] For ease of description, the above embodiments of the present disclosure have been described separately, but the present disclosure is not limited thereto. That is, the embodiments described in the above Embodiments 1 to 14 can be executed independently, and one or more various embodiments can be combined and executed.
[0554] Figure 30 is a diagram illustrating an inverse transform method according to an embodiment of applying the present disclosure.
[0555] Referring to Figure 30 , the decoder is mainly described for ease of description, but the method of performing transform / inverse transform based on the quadratic transform set according to an embodiment of the present disclosure can be applied to the encoder even substantially equivalently.
[0556] The decoder generates a dequantized transform block by performing dequantization on the current block (S3001).
[0557] The decoder obtains the intra prediction mode of the current block (S3002).
[0558] The decoder determines the quadratic transform set applied to the current block among multiple quadratic transform sets based on the intra prediction mode (S3003).
[0559] As described above, multiple sets of secondary transforms may include at least one hybrid set of secondary transforms.
[0560] Furthermore, as described above, a hybrid set of secondary transforms may include at least one 8×8 transform kernel applied to a region of 8×8 size and at least one 4×4 transform kernel applied to a region of 4×4 size.
[0561] In addition, as described above, when multiple sets of secondary transforms include multiple hybrid sets of secondary transforms, the multiple hybrid sets of secondary transforms may respectively include different numbers of transform kernels.
[0562] Moreover, as described above, step S3003 may include: determining whether to use a hybrid set of secondary transforms based on an intra prediction mode; and when it is determined to use a hybrid set of secondary transforms, determining, based on the size of the current block, the set of secondary transforms applied to the current block among the multiple sets of secondary transforms including at least one hybrid set of secondary transforms; and when it is determined not to use a hybrid set of secondary transforms, determining, based on the size of the current block, the set of secondary transforms applied to the current block among the remaining sets of secondary transforms other than the at least one hybrid set of secondary transforms.
[0563] The decoder derives the transform kernel applied to the current block in the determined set of secondary transforms (S3004).
[0564] As described above, step S3004 may further include obtaining a secondary transform index indicating the transform kernel applied to the current block in the determined set of secondary transforms. As an implementation, the secondary transform index may be binarized by a truncated unary scheme based on the maximum number of available transform kernels in the determined set of secondary transforms.
[0565] The decoder performs a secondary transform on a specific upper left region of the inverse quantized transform block by using the derived transform kernel (S3005).
[0566] Figure 31 FIG. is an illustration showing an inverse quantization unit and an inverse transform unit according to an embodiment of applying the present disclosure.
[0567] In Figure 31 for ease of description, the inverse transform unit 3100 is illustrated as a single block, but the inter prediction unit may be implemented as a component included in an encoder and / or a decoder.
[0568] Referring to Figure 31 , the inverse quantization unit 3101 and the inverse transform unit 3100 implement the above Figures 4 to 30The functions, processes, and / or methods proposed in
[0569] The dequantization unit 3101 generates a dequantized transform block by performing dequantization on the current block.
[0570] The intra prediction mode acquisition unit 3102 acquires the intra prediction mode of the current block.
[0571] The secondary transform set determination unit determines the secondary transform set to be applied to the current block among multiple secondary transform sets based on the intra prediction mode.
[0572] As described above, the multiple secondary transform sets may include at least one hybrid secondary transform set.
[0573] In addition, as described above, the hybrid secondary transform set may include at least one 8×8 transform kernel applied to an 8×8-sized region and at least one 4×4 transform kernel applied to a 4×4-sized region.
[0574] Furthermore, as described above, when the multiple secondary transform sets include multiple hybrid secondary transform sets, the multiple hybrid secondary transform sets may respectively include different numbers of transform kernels.
[0575] Moreover, as described above, the secondary transform set determination unit 3103 may determine whether to use the hybrid secondary transform set based on the intra prediction mode; and when it is determined to use the hybrid secondary transform set, based on the size of the current block, determine the secondary transform set to be applied to the current block among the multiple secondary transform sets including at least one hybrid secondary transform set; and when it is determined not to use the hybrid secondary transform set, based on the size of the current block, determine the secondary transform set to be applied to the current block among the remaining secondary transform sets other than the at least one hybrid secondary transform set.
[0576] The transform kernel derivation unit 3104 derives the transform kernel to be applied to the current block in the determined secondary transform set.
[0577] As described above, the transform kernel derivation unit 3104 may acquire a secondary transform index indicating the transform kernel to be applied to the current block in the determined secondary transform set. As an implementation, the secondary transform index may be binarized by a truncated unary scheme based on the maximum number of available transform kernels in the determined secondary transform set.
[0578] The secondary inverse transform unit 3105 performs a secondary inverse transform on a specific upper-left region of the dequantized transform block by using the derived transform kernel.
[0579] The multiple secondary transformation sets include at least one hybrid secondary transformation set.
[0580] Figure 32 An example of applying the video coding system of the present disclosure is illustrated.
[0581] The video coding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.
[0582] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0583] The video source may obtain video / images through processes such as capture, synthesis, or generation of video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like, and in this case, the video / image capture process may be replaced by a process of generating relevant data.
[0584] The encoding device may encode the input video / images. The encoding device may perform a series of processes including prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0585] The transmitter may transmit the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and may include elements for transmitting via a broadcast / communication network. The receiver may extract the bitstream and transmit the extracted bitstream to the decoding device.
[0586] The decoding device performs a series of processes including inverse quantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding device to decode the video / images.
[0587] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0588] Figure 33 It is a structural diagram of a content streaming system to which the embodiments of the present disclosure are applied.
[0589] Referring to Figure 33 , the content streaming system to which the present disclosure is applied may mainly include an encoding server, a streaming server, a network server, a media memory, a user device, and a multimedia input device.
[0590] The encoding server compresses the content input from a multimedia input device including a smart phone, a camera, a video camera, etc. into digital data for generating a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device including a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server may be omitted.
[0591] A bitstream can be generated by applying the encoding method or bitstream generation method of the present disclosure, and the streaming server can temporarily store the bitstream in the process of sending or receiving the bitstream.
[0592] The streaming server sends multimedia data to the user device based on a request from the user through the network server, and the network server serves as an intermediary for informing the user of what services are available. When the user requests a desired service from the network server, the network server transmits the requested service to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server is used to control commands / responses between various devices in the content streaming system.
[0593] The streaming server can receive content from the media memory and / or the encoding server. For example, when the streaming server receives content from the encoding server, the streaming server can receive the content in real time. In this case, the streaming server can store the bitstream for a predetermined time to provide a smooth streaming service.
[0594] Examples of user devices may include cellular phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, touchscreen PCs, tablet PCs, ultrabooks, wearable devices such as smart watches, smart glasses, or head-mounted displays (HMDs), etc.
[0595] Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be distributedly processed.
[0596] As described above, the embodiments described in the present disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0597] In addition, the decoders and encoders applying the present disclosure can be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video-on-demand (VoD) service providing device, a (set-top) over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, a vehicle terminal (e.g., a vehicle terminal, an aircraft terminal, a ship terminal, etc.), a medical video device, etc., and can be used to process video signals or data signals. For example, a set-top (OTT) video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0598] In addition, the processing method applying the present disclosure can be generated in the form of a program executable by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a flexible disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, a bit stream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted through a wired / wireless communication network.
[0599] In addition, the embodiments of the present disclosure can be implemented as a computer program product by program code, and the program code can be executed on a computer by the embodiments of the present disclosure. The program code can be stored on a computer-readable carrier.
[0600] In the above-described embodiments, the components and features of the present disclosure are combined in a predetermined form. Unless otherwise explicitly stated, each component or feature should be considered as an option. Each component or feature can be implemented without being associated with other components or features. Additionally, embodiments of the present disclosure can be configured by associating some components and / or features. The order of operations described in the embodiments of the present disclosure can be changed. Some components or features of any embodiment can be included in another embodiment or replaced with components and features corresponding to another embodiment. Obviously, after filing, it is possible to modify the claims not explicitly recited in the combination claims to form an embodiment or include them in new claims.
[0601] Embodiments of the present disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of implementation by hardware, according to the hardware implementation manner, the exemplary embodiments described herein can be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and the like.
[0602] In the case of implementation by firmware or software, embodiments of the present disclosure can be implemented in the form of modules, processes, functions, etc. to perform the above-described functions or operations. The software code can be stored in a memory and run by a processor. The memory can be internal or external to the processor, and can send data to / receive data from the processor by various known means.
[0603] It is obvious to those skilled in the art that the present disclosure can be implemented in other specific forms without departing from the essential characteristics of the present disclosure. Therefore, the above-mentioned detailed description should not be construed as restrictive in any way and should be considered exemplary. The scope of the present disclosure should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present disclosure are included within the scope of the present disclosure.
[0604] Industrial Applicability
[0605] In the foregoing, for illustrative purposes, preferred embodiments of the present disclosure have been disclosed, and hereinafter, those skilled in the art will make various modifications, changes, substitutions, or additions to other embodiments within the technical spirit and scope of the present disclosure disclosed in the appended claims.
Claims
1. A method for decoding a video signal, the method comprising the following steps: Generating an inverse-quantized transform block for a current block; Determining an intra prediction mode of the current block; Determining, based on the intra prediction mode, a secondary transform set among a plurality of secondary transform sets to be applied to the current block, wherein each of a group of intra prediction mode groups is assigned one of the plurality of secondary transform sets; Deriving, based on a secondary transform index, a transform kernel in the determined secondary transform set to be applied to the current block; and Performing an inverse secondary transform on coefficients of the inverse-quantized transform block based on the derived transform kernel, and wherein the number of coefficients on which the inverse secondary transform is performed is less than the number of coefficients obtained by performing the inverse secondary transform, and wherein the intra prediction mode is determined among 67 intra prediction modes and wide-angle intra prediction modes, A first secondary transform set is assigned to intra prediction modes 0 and 1, A second secondary transform set is assigned to intra prediction modes less than 0, greater than or equal to 2 and less than or equal to 12, and greater than or equal to 56, A third secondary transform set is assigned to intra prediction modes greater than or equal to 13 and less than or equal to 23 and greater than or equal to 45 and less than or equal to 55, and A fourth secondary transform set is assigned to intra prediction modes greater than or equal to 24 and less than or equal to 44.
2. The method according to claim 1, wherein, Binarizing the secondary transform index in a truncated unary scheme based on the number of available transform kernels in the determined secondary transform set.
3. The method according to claim 1, wherein The plurality of secondary transform sets includes at least one hybrid secondary transform set, the hybrid secondary transform set including at least one 8×8 transform kernel applied to an 8×8 size region and at least one 4×4 transform kernel applied to a 4×4 size region.
4. The method according to claim 3, wherein When the plurality of secondary transform sets includes a plurality of hybrid secondary transform sets, the plurality of hybrid secondary transform sets respectively include different numbers of transform kernels.
5. The method according to claim 4, Among them, Binarizing the secondary transform index in a truncated unary scheme based on the maximum number of available transform kernels in the determined secondary transform set.
6. The method according to claim 3, wherein, The step of determining the secondary transform set includes: Determining whether to use the hybrid secondary transform set based on the intra prediction mode, and When it is determined to use the hybrid secondary transform set, determining, based on the size of the current block, the secondary transform set among the plurality of secondary transform sets including the at least one hybrid secondary transform set to be applied to the current block, and when it is determined not to use the hybrid secondary transform set, determining, based on the size of the current block, the secondary transform set among the remaining secondary transform sets other than the at least one hybrid secondary transform set to be applied to the current block.
7. A method for encoding a video signal, the method comprising the following steps: Obtaining residual data of a current block; Performing a primary transform on the residual data to obtain a first block; Performing a secondary transform on a first number of coefficients of the first block to obtain a second number of transform coefficients; And Performing quantization on the transform coefficients, wherein the step of performing the secondary transform includes: Determine a secondary transform set among a plurality of secondary transform sets to be applied to the first block based on the intra prediction mode applied to the current block; Derive a transform kernel to be applied to the first block in the secondary transform set; Generate a secondary transform index including information about the transform kernel; and Apply the transform kernel to the first block, and wherein the first number is greater than the second number, and wherein the intra prediction mode is determined among 67 intra prediction modes and wide-angle intra prediction modes, The first secondary transform set is assigned to intra prediction modes 0 and 1, The second secondary transform set is assigned to intra prediction modes less than 0, greater than or equal to 2 and less than or equal to 12, and greater than or equal to 56, The third secondary transform set is assigned to intra prediction modes greater than or equal to 13 and less than or equal to 23 and greater than or equal to 45 and less than or equal to 55, and The fourth secondary transform set is assigned to intra prediction modes greater than or equal to 24 and less than or equal to 44.
8. The method according to claim 7, the method further comprising the steps of: Binarize the secondary transform index in a truncated unary scheme based on the number of available transform kernels in the secondary transform set.
9. The method according to claim 7, wherein, The plurality of secondary transform sets includes at least one hybrid secondary transform set, the hybrid secondary transform set containing at least one 8×8 transform kernel applied to an 8×8 size region and at least one 4×4 transform kernel applied to a 4×4 size region.
10. A non-transitory computer-readable storage medium for storing picture information generated by performing the following steps: Obtain residual data of a current block; Perform a primary transform on the residual data to obtain a first block; Perform a quadratic transform on a first number of coefficients of the first block to obtain a second number of transform coefficients; And Perform quantization on the transform coefficients, wherein, The step of performing the secondary transform includes the following steps: Determine a secondary transform set among a plurality of secondary transform sets to be applied to the first block based on the intra prediction mode applied to the current block; Derive a transform kernel to be applied to the first block in the secondary transform set; Generate a secondary transform index including information about the transform kernel; and Apply the transform kernel to the first block, and wherein the first number is greater than the second number, and wherein the intra prediction mode is determined among 67 intra prediction modes and wide-angle intra prediction modes, The first secondary transform set is assigned to intra prediction modes 0 and 1, The second secondary transform set is assigned to intra prediction modes less than 0, greater than or equal to 2 and less than or equal to 12 and greater than or equal to 56, The third secondary transform set is assigned to intra prediction modes greater than or equal to 13 and less than or equal to 23 and greater than or equal to 45 and less than or equal to 55, and The fourth secondary transform set is assigned to intra prediction modes greater than or equal to 24 and less than or equal to 44.
Citation Information
Patent Citations
Non-separable secondary transform for video coding
US20170094313A1
Binarizing secondary transform index
US20170324643A1