Devices for encoding and decoding video signals and devices for transmitting image data
By configuring the secondary transformation set and intra prediction mode to optimize the encoding and decoding process of video signals, the challenges of storage and processing capabilities in high-resolution video content processing are solved, and encoding efficiency and decoding performance are improved.
Patent Information
- Application Number
- CN202310313317.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-02
- Filing Date
- 2019-07-02
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-07-02
AI Technical Summary
The prior art is difficult to efficiently handle the problem of the sharp increase in memory storage, memory access rate and processing capabilities brought about by the high spatial resolution, high frame rate and high-dimensional scene presentation of next-generation video content, especially in terms of transformation encoding efficiency and complexity.
By configuring the quadratic transformation set, including a mixed quadratic transformation set and intra prediction mode, using inverse quantization, transform kernel derivation and quadratic inverse transformation, the encoding and decoding process of video signals is optimized, especially transform index encoding for different block sizes and prediction modes.
It improves the encoding efficiency and decoding efficiency of video signals, reduces storage and processing requirements, and improves the performance of video signal processing.
Smart Images

Figure CN116347076B_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application for invention with the original application number 201980045151.X (International Application Number: PCT / KR2019 / 008060, filing date: July 2, 2019, invention title: Method and apparatus for processing video signals based on secondary transformation). Technical Field
[0002] The present disclosure relates to a method and apparatus for processing video signals, and more particularly, to a method of designing and applying a secondary transformation. Background Art
[0003] Next-generation video content will have characteristics such as high spatial resolution, high frame rate, and high-dimensional scene rendering. To process the content, there will be a sharp increase in memory storage, memory access rate, and processing power.
[0004] Therefore, there is a need to design new coding tools for more efficiently processing next-generation video content. Specifically, when applying a transformation, there is a need to design a more efficient transformation in terms of coding efficiency and complexity. Summary of the Invention
[0005] Technical Problem
[0006] Embodiments of the present disclosure provide a method of configuring a secondary transformation set by considering various conditions for applying a secondary transformation.
[0007] In addition, embodiments of the present disclosure also provide a method of efficiently encoding / decoding a secondary transformation index signaled from an encoder when configuring a second transformation set based on the size and / or intra prediction mode of a coding block (or transformation block).
[0008] The technical objects of the present disclosure are not limited to the above-mentioned technical objects, and for those of ordinary skill in the art, other technical objects not mentioned above will become apparent from the following description.
[0009] Technical Solution
[0010] In one aspect of the present disclosure, there is provided a method of decoding a video signal, the method including the steps of: generating an inverse quantized transformation block by performing inverse quantization on a current block; obtaining an intra prediction mode of the current block; determining a secondary transformation set applied to the current block among a plurality of secondary transformation sets based on the intra prediction mode; deriving a transformation kernel applied to the current block in the determined secondary transformation set; and performing a secondary transformation on a specific upper left region of the inverse quantized transformation block by using the derived transformation kernel, wherein the plurality of secondary transformation sets may include at least one hybrid secondary transformation set.
[0011] Preferably, the hybrid secondary transform set may include at least one 8×8 transform kernel applied to a region of 8×8 size and at least one 4×4 transform kernel applied to a region of 4×4 size.
[0012] Preferably, when the plurality of secondary transform sets include a plurality of hybrid secondary transform sets, the plurality of hybrid secondary transform sets may respectively include different numbers of transform kernels.
[0013] Preferably, deriving the transform kernel may further include obtaining a secondary transform index indicating the transform kernel applied to the current block in the determined secondary transform set, and binarizing the secondary transform index by a truncated unary scheme based on the maximum number of available transform kernels in the determined secondary transform set.
[0014] Preferably, determining the secondary transform set may include: determining whether to use a hybrid secondary transform set based on the intra prediction mode, and when it is determined to use a hybrid secondary transform set, determining the secondary transform set applied to the current block among the plurality of secondary transform sets including the at least one hybrid secondary transform set based on the size of the current block, and when it is determined not to use a hybrid secondary transform set, determining the secondary transform set applied to the current block among the remaining secondary transform sets other than the at least one hybrid secondary transform set based on the size of the current block.
[0015] In another aspect of the present disclosure, there is provided an apparatus for decoding a video signal, the apparatus including: an inverse quantization unit that generates an inverse quantized transform block by performing inverse quantization on a current block; a prediction mode acquisition unit that acquires the intra prediction mode of the current block; a secondary transform set determination unit that determines the secondary transform set applied to the current block among a plurality of secondary transform sets based on the intra prediction mode; a transform kernel derivation unit that derives the transform kernel applied to the current block in the determined secondary transform set; and a secondary inverse transform unit that performs a secondary inverse transform on a specific upper left region of the inverse quantized transform block by using the derived transform kernel, wherein the plurality of secondary transform sets may include at least one hybrid secondary transform set.
[0016] Preferably, the hybrid secondary transform set may include at least one 8×8 transform kernel applied to a region of 8×8 size and at least one 4×4 transform kernel applied to a region of 4×4 size.
[0017] Preferably, when the plurality of secondary transform sets include a plurality of hybrid secondary transform sets, the plurality of hybrid secondary transform sets may respectively include different numbers of transform kernels.
[0018] Preferably, the transform kernel derivation unit may obtain a secondary transform index indicating the transform kernel in the determined secondary transform set applied to the current block, and may binarize the secondary transform index by a truncated unary scheme based on the maximum number of available transform kernels in the determined secondary transform set.
[0019] Preferably, the secondary transform set determination unit may determine whether to use a hybrid secondary transform set based on the intra prediction mode, and when it is determined to use the hybrid secondary transform set, determine the secondary transform set applied to the current block among a plurality of secondary transform sets including the at least one hybrid secondary transform set based on the size of the current block, and when it is determined not to use the hybrid secondary transform set, determine the secondary transform set applied to the current block among the remaining secondary transform sets other than the at least one hybrid secondary transform set based on the size of the current block.
[0020] Advantageous Effects
[0021] According to an embodiment of the present disclosure, a secondary transform set is configured to apply a secondary transform by considering various conditions, so as to efficiently select a transform for the secondary transform.
[0022] The effects obtainable in the present disclosure are not limited to the effects mentioned above, and those skilled in the art will clearly understand other unmentioned effects from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a schematic block diagram of an encoder that performs video signal encoding by applying an embodiment of the present disclosure.
[0024] Figure 2 is a schematic block diagram of a decoder that performs video signal decoding by applying an embodiment of the present disclosure.
[0025] Figures 3a to 3d illustrates an embodiment to which the present disclosure can be applied, Figure 3a is a diagram describing a quadtree (QT) (hereinafter, referred to as "QT") block partitioning structure, Figure 3b is a diagram describing a binary tree (BT) (hereinafter, referred to as "BT") block partitioning structure, Figure 3c is a diagram describing a ternary tree (TT) (hereinafter, referred to as "TT") block partitioning structure, and Figure 3d is a diagram describing an asymmetric tree (AT) (hereinafter, referred to as "AT") block partitioning structure.
[0026] Figure 4It is a schematic block diagram of a transform unit 120, a quantization unit 130, an inverse quantization unit 140, and an inverse transform unit 150 in an encoder to which an embodiment of the present disclosure is applied.
[0027] Figure 5 It is a schematic block diagram of an inverse quantization unit 220 and an inverse transform unit 230 in a decoder to which an embodiment of the present disclosure is applied.
[0028] Figure 6 It is a table showing a transform configuration group to which multiple transform selection (MTS) according to an embodiment of the present disclosure is applied.
[0029] Figure 7 It is a flowchart showing an encoding process for performing multiple transform selection (MTS) according to an embodiment of the present disclosure.
[0030] Figure 8 It is a flowchart showing a decoding process for performing multiple transform selection (MTS) according to an embodiment of the present disclosure.
[0031] Figure 9 It is a flowchart for describing an encoding process of an MTS flag and an MTS index according to an embodiment of the present disclosure.
[0032] Figure 10 It is a flowchart for describing a decoding process of applying a horizontal transform or a vertical transform to a row or a column based on an MTS flag and an MTS index according to an embodiment of the present disclosure.
[0033] Figure 11 It is a flowchart for performing an inverse transform based on transform-related parameters according to an embodiment of the present disclosure.
[0034] Figure 12 It is a table showing a transform set assigned for each intra prediction mode in NSST according to an embodiment of the present disclosure.
[0035] Figure 13 It is a calculation flowchart of Givens rotation according to an embodiment of the present disclosure.
[0036] Figure 14 It illustrates a round configuration in a 4×4 NSST composed of a Givens rotation layer and a permutation according to an embodiment of the present disclosure.
[0037] Figure 15 It is a block diagram for describing operations of a forward simplified transform and a reverse simplified transform according to an embodiment of the present disclosure.
[0038] Figure 16Is a diagram illustrating the process of performing backward scanning from the 64th to the 17th in the backward scanning order as an embodiment of applying the present disclosure.
[0039] Figure 17 Illustrates three forward scanning orders of a transform coefficient block (transform block) as an embodiment of applying the present disclosure.
[0040] Figure 18 Illustrates the positions and forward scanning orders of valid transform coefficients of each 4×4 block when applying diagonal scanning and 4×4 RST in the upper left 4×8 block as an embodiment of applying the present disclosure.
[0041] Figure 19 Illustrates the case where the valid transform coefficients of two 4×4 blocks are combined into one 4×4 block when applying diagonal scanning and 4×4 RST in the upper left 4×8 block as an embodiment of applying the present disclosure.
[0042] Figure 20 Is a flowchart for encoding a video signal based on a simplified quadratic transform as an embodiment of applying the present disclosure.
[0043] Figure 21 Is a flowchart for decoding a video signal based on a simplified quadratic transform as an embodiment of applying the present disclosure.
[0044] Figure 22 Is a diagram illustrating a method for determining a transform type applied to a quadratic transform according to an embodiment of applying the present disclosure.
[0045] Figure 23 Is a diagram illustrating a simplified transform structure based on a simplification factor that can be applied as an embodiment of the present disclosure.
[0046] Figure 24 Is a diagram illustrating a method for performing decoding by adaptively applying a simplified transform as an embodiment that can be applied to the present disclosure.
[0047] Figure 25 Is a diagram illustrating a method for performing decoding by adaptively applying a simplified transform as an embodiment that can be applied to the present disclosure.
[0048] Figure 26 and Figure 27 Is a diagram illustrating an example of a forward simplified quadratic transform and a backward simplified quadratic transform applied as an embodiment of the present disclosure and the pseudocode for deriving them.
[0049] Figure 28 Is a diagram illustrating a method for applying a simplified quadratic transform to a non-square region as an embodiment of applying the present disclosure.
[0050] Figure 29 is a diagram illustrating a simplified transform controlled by a simplification factor as an embodiment of applying the present disclosure.
[0051] Figure 30 is a diagram illustrating an inverse transform method according to an embodiment of applying the present disclosure.
[0052] Figure 31 is a diagram illustrating an inverse transform unit according to an embodiment of applying the present disclosure.
[0053] Figure 32 Illustrates a video coding system applying the present disclosure.
[0054] Figure 33 is a structural diagram of a content streaming system as an embodiment of applying the present disclosure. Detailed Embodiments
[0055] Hereinafter, the configuration and operation of the embodiments of the present disclosure will be described with reference to the drawings. The configuration and operation of the present disclosure described through the drawings will be described as one embodiment, and thus the technical spirit, core configuration, and operation of the present disclosure are not limited.
[0056] In addition, the terms used in the present disclosure are selected as general terms that are as widely used as possible at present. In specific cases, terms arbitrarily selected by the applicant will be used for description. In this case, since its meaning is clearly described in the detailed description of this part, it should not be simply interpreted only by the name of the term used in the description of the present disclosure. It should be understood that the meaning of the term should be interpreted.
[0057] In addition, when there are general terms selected for describing the present invention or other terms with similar meanings, the terms used in the present disclosure can be replaced with more appropriate interpretations. For example, signals, data, samples, pictures, frames, blocks, etc. can be appropriately replaced and interpreted in each coding process. In addition, separation, decomposition, segmentation, and division can be appropriately replaced and interpreted in each coding process.
[0058] In this document, multiple transform selection (MTS) may refer to a method for performing a transform using at least two transform types. This may also be expressed as adaptive multiple transform (AMT) or explicit multiple transform (EMT). Similarly, mts_idx may also be expressed as AMT_idx, EMT_idx, tu_mts_idx, AMT_TU_idx, EMT_TU_idx, transform index, or transform combination index, and the present disclosure is not limited to these expressions.
[0059] Figure 1It is a schematic block diagram of an encoder that executes video signal encoding as an embodiment of the present disclosure.
[0060] Referring to Figure 1 , the encoder 100 may be configured to include an image partitioning unit 110, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, a filtering unit 160, a decoded picture buffer (DPB) 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190.
[0061] The image partitioning unit 110 may partition an input image (or picture or frame) input to the encoder 100 into one or more processing units. For example, the processing unit may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transform unit (TU).
[0062] However, these terms are only for the convenience of describing the present disclosure, and the present disclosure is not limited to the definitions of these terms. Additionally, in the present disclosure, for convenience of description, the term "coding unit" is used as a unit used when encoding or decoding a video signal, but the present disclosure is not limited thereto, and can be appropriately interpreted according to the present disclosure.
[0063] The encoder 100 subtracts a prediction signal (or prediction block) output from the inter prediction unit 180 or the intra prediction unit 185 from the input image signal to generate a residual signal (or residual block), and the generated residual signal is sent to the transform unit 120.
[0064] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. The transform process may be applied to a quadtree-structured square block and blocks (square or rectangular) divided according to a binary tree structure, a ternary tree structure, or an asymmetric tree structure.
[0065] The transform unit 120 may perform a transform based on multiple transforms (or transform combinations), and the transform scheme may be referred to as multiple transform selection (MTS). MTS may also be referred to as adaptive multiple transform (AMT) or enhanced multiple transform (EMT).
[0066] MTS (or AMT or EMT) may refer to a transform scheme that is performed based on a transform (or transform combination) adaptively selected from multiple transforms (or transform combinations).
[0067] The multiple transforms (or transform combinations) may include the transforms (or transform combinations) described in the Figure 6 of the present disclosure. In the present disclosure, a transform or transform type may be represented, for example, as DCT type 2, DCT-II, DCT2, or DCT-2.
[0068] The transform unit 120 may perform the following embodiments.
[0069] The present disclosure provides methods for designing RSTs applicable to 4×4 blocks.
[0070] The present disclosure provides methods for configuring regions to which 4×4 RSTs are applied, arranging transform coefficients generated after applying 4×4 RSTs, scanning orders of the arranged transform coefficients, methods for sorting and combining transform coefficients generated for each block, etc.
[0071] The present disclosure provides methods for encoding transform indices that specify 4×4 RSTs.
[0072] The present disclosure provides methods for conditionally encoding corresponding transform indices by checking for the presence of non-zero transform coefficients in unacceptable regions when applying 4×4 RSTs.
[0073] The present disclosure provides methods for conditionally encoding corresponding transform indices after encoding the position of the last non-zero transform coefficient and then omitting the relevant residual encoding for unacceptable positions.
[0074] The present disclosure provides methods for applying different transform index encodings and residual encodings to luminance blocks and chrominance blocks when applying 4×4 RSTs.
[0075] Specific embodiments thereof will be described in more detail in the present disclosure.
[0076] The quantization unit 130 may quantize the transform coefficients and send the quantized transform coefficients to the entropy encoding unit 190, and the entropy encoding unit 190 may perform entropy encoding on the quantized signal and output the entropy-encoded quantized signal as a bitstream.
[0077] Although the transform unit 120 and the quantization unit 130 are described as separate functional units, the present disclosure is not limited thereto, and they may be combined into one functional unit. The inverse quantization unit 140 and the inverse transform unit 150 may also be similarly combined into one functional unit.
[0078] The quantized signal output from the quantization unit 130 may be used to generate a prediction signal. For example, the inverse quantization and inverse transform are cyclically applied to the quantized signal by the inverse quantization unit 140 and the inverse transform unit 150 to reconstruct the residual signal. The reconstructed residual signal is added to the prediction signal output from the inter-frame prediction unit 180 or the intra-frame prediction unit 185 to generate a reconstructed signal.
[0079] In addition, due to quantization errors that occur during such compression processing, deterioration showing block boundaries may occur. This phenomenon is called blocking artifact, which is one of the key elements for evaluating image quality. To reduce the deterioration, filtering processing can be performed. Through the filtering processing, the blocking artifact is eliminated and the errors in the current picture are reduced to enhance the image quality.
[0080] The filtering unit 160 applies filtering to the reconstructed signal and outputs the applied reconstructed signal to the reproduction device, or sends the output reconstructed signal to the decoded picture buffer 170. The inter-frame prediction unit 180 can use the filtered signal sent to the decoded picture buffer 170 as a reference picture. Thus, in the inter-picture prediction mode, the filtered picture is used as a reference picture to enhance the image quality and coding efficiency.
[0081] The decoded picture buffer 170 can store the filtered picture so as to use the filtered picture as a reference picture in the inter-frame prediction unit 180.
[0082] The inter-frame prediction unit 180 performs temporal prediction and / or spatial prediction in order to remove temporal redundancy and / or spatial redundancy by referring to the reconstructed pictures. Here, since the reference pictures used for prediction are transform signals quantized and dequantized in units of blocks during encoding / decoding at a previous time, there may be blocking artifacts or ringing effects.
[0083] Therefore, the inter-frame prediction unit 180 can interpolate signals between pixels in units of sub-pixels by applying a low-pass filter in order to solve the performance degradation caused by the discontinuity or quantization of this signal. Here, sub-pixels mean virtual pixels generated by applying an interpolation filter, and integer pixels mean actual pixels existing in the reconstructed picture. As an interpolation method, linear interpolation, bilinear interpolation, Wiener filter, etc. can be adopted.
[0084] An interpolation filter is applied to the reconstructed picture to enhance the prediction accuracy. For example, the inter-frame prediction unit 180 applies an interpolation filter to integer pixels to generate interpolated pixels, and can perform prediction by using the interpolated block composed of the interpolated pixels as a prediction block.
[0085] In addition, the intra prediction unit 185 may predict the current block by referring to samples near the block to be currently encoded. The intra prediction unit 185 may perform the following processes to perform intra prediction. First, reference samples, which are required for generating a prediction signal, may be prepared. Additionally, a prediction signal may be generated by using the prepared reference samples. Thereafter, the prediction mode is encoded. In this case, the reference samples may be prepared by zero-padding the reference samples and / or filtering the reference samples. Since the reference samples have undergone prediction and reconstruction processes, quantization errors may exist. Therefore, a reference sample filtering process may be performed for each prediction mode used for intra prediction to reduce this error.
[0086] The prediction signal generated by the inter prediction unit 180 or the intra prediction unit 185 may be used to generate a reconstructed signal or may be used to generate a residual signal.
[0087] Figure 2 is a schematic block diagram of a decoder that executes video signal decoding as an embodiment of applying the present disclosure.
[0088] Referring to Figure 2 , the decoder 200 may be configured to include a parsing unit (not illustrated), an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, a filtering unit 240, a decoded picture buffer (DPB) unit 250, an inter prediction unit 260, and an intra prediction unit 265.
[0089] Additionally, the reconstructed video signal output by the decoder may be reproduced by a reproducing device.
[0090] The decoder 200 may receive a signal output from the Figure 1 encoder 100 and may perform entropy decoding on the received signal by the entropy decoding unit 210.
[0091] The inverse quantization unit 220 obtains transform coefficients from the entropy-decoded signal by using quantization step information.
[0092] The inverse transform unit 230 performs an inverse transform on the transform coefficients to obtain a residual signal.
[0093] Here, the present disclosure provides a method of configuring a transform combination for each group of transform configurations divided by at least one of a prediction mode, a block size, or a block shape, and the inverse transform unit 230 may perform an inverse transform based on the transform combination configured by the present disclosure. Additionally, the embodiments described in the present disclosure may be applied.
[0094] The inverse transform unit 230 may perform the following embodiments.
[0095] The present disclosure provides a method of reconstructing a video signal based on a simplified quadratic transform.
[0096] The inverse transform unit 230 can derive a secondary transform corresponding to a secondary transform index, perform an inverse secondary transform on a transform coefficient block by using the secondary transform, and perform an inverse primary transform on the block on which the inverse secondary transform is performed. Here, the secondary transform refers to a simplified secondary transform, and the simplified secondary transform represents a transform that inputs N residual data (N×1 residual vector) to output L (L < N) transform coefficient data (L×1 transform coefficient vector).
[0097] The present disclosure is characterized in that a simplified secondary transform is applied to a specific region of a current block, and the specific region is the upper left M×M (M≤N) in the current block.
[0098] The present disclosure is characterized in that when performing the inverse secondary transform, a 4×4 simplified secondary transform is applied to each of the 4×4 blocks divided in the current block.
[0099] The present disclosure is characterized in that it is determined whether a secondary transform index is obtained based on the position of the last non-zero transform coefficient in the transform coefficient block.
[0100] The present disclosure is characterized in that when the last non-zero transform coefficient is not in a specific region, a secondary transform index is obtained, and the specific region indicates the remaining region except for the positions where non-zero transform coefficients may exist when arranging the transform coefficients according to the scan order in the case of applying the simplified secondary transform.
[0101] The inverse transform unit 230 can derive a transform combination corresponding to a primary transform index, and perform an inverse primary transform by using the transform combination. Here, the primary transform index corresponds to any one of a plurality of transform combinations constituted by a combination of DST7 and / or DCT8, and the transform combination includes a horizontal transform and a vertical transform. In this case, the horizontal transform and the vertical transform correspond to DST7 or DCT8.
[0102] Although the inverse quantization unit 220 and the inverse transform unit 230 are described as separate functional units, the present disclosure is not limited thereto, and they may be combined into one functional unit.
[0103] The obtained residual signal is added to the prediction signal output from the inter-frame prediction unit 260 or the intra-frame prediction unit 265 to generate a reconstructed signal.
[0104] The filtering unit 240 applies filtering to the reconstructed signal, and outputs the applied reconstructed signal to the generation device, or sends the output reconstructed signal to the decoded picture buffer unit 250. The inter-frame prediction unit 260 can use the filtered signal sent to the decoded picture buffer unit 250 as a reference picture.
[0105] In the present disclosure, the embodiments described in the respective functional units of the transform unit 120 and the encoder 100 can be equivalently applied to the corresponding functional units of the inverse transform unit 230 and the decoder, respectively.
[0106] Figures 3a to 3d Embodiments to which the present disclosure can be applied are illustrated. Figure 3a is a diagram for describing a quadtree (QT) (hereinafter referred to as "QT") block partitioning structure. Figure 3b is a diagram for describing a binary tree (BT) (hereinafter referred to as "BT") block partitioning structure. Figure 3c is a diagram for describing a ternary tree (TT) (hereinafter referred to as "TT") block partitioning structure, and Figure 3d is a diagram for describing an asymmetric tree (AT) (hereinafter referred to as "AT") block partitioning structure.
[0107] During video encoding, a block can be partitioned based on a quadtree (QT). Additionally, a sub-block partitioned by QT can be further recursively partitioned using QT. A leaf block that is no longer partitioned by QT can be partitioned by at least one of a binary tree (BT), a ternary tree (TT), and an asymmetric tree (AT). BT can have two partitioning types: horizontal BT (2N×N, 2N×N) and vertical BT (N×2N, N×2N). TT can have two partitioning types: horizontal TT (2N×1 / 2N, 2N×N, 2N×1 / 2N) and vertical TT (1 / 2N×2N, N×2N, 1 / 2N×2N). AT can have four partitioning types: horizontal-up AT (2N×1 / 2N, 2N×3 / 2N), horizontal-down AT (2N×3 / 2N, 2N×1 / 2N), vertical-left AT (1 / 2N×2N, 3 / 2N×2N), and vertical-right AT (3 / 2N×2N, 1 / 2N×2N). Each of BT, TT, and AT can be further recursively partitioned by using BT, TT, and AT.
[0108] Figure 3a An example of QT partitioning is illustrated. Block A can be partitioned into four sub-blocks A0, A1, A2, and A3 by QT. Sub-block A1 can be further partitioned into four sub-blocks B0, B1, B2, and B3 by QT again.
[0109] Figure 3b An example of BT partitioning is illustrated. Block B3 that is no longer partitioned by QT can be partitioned into vertical BT (C0, C1) or horizontal BT (D0, D1). Each sub-block, like block C0, can be further recursively partitioned in the form of horizontal BT (E0, E1) or vertical BT (F0, F1).
[0110] Figure 3cAn example of a TT partition is illustrated. The block B3 that is no longer partitioned by QT can be partitioned into a vertical TT (C0, C1, C2) or a horizontal TT (D0, D1, D2). Similar to block C1, each sub-block can be further recursively partitioned in the form of a horizontal TT (E0, E1, E2) or a vertical TT (F0, F1, F2).
[0111] Figure 3d An example of an AT partition is illustrated. The block B3 that is no longer partitioned by QT can be partitioned into a vertical AT (C0, C1) or a horizontal AT (D0, D1). Similar to block C1, each sub-block can be further recursively partitioned in the form of a horizontal AT (E0, E1) or a vertical TT (F0, F1).
[0112] In addition, BT, TT, and AT partitions can be used and performed together. For example, a sub-block partitioned by BT can be partitioned by TT or AT. Additionally, a sub-block partitioned by TT can be partitioned by BT or AT. Additionally, a sub-block partitioned by AT can be partitioned by BT or TT. For example, after a horizontal BT partition, each sub-block can be partitioned into a vertical BT, or after a vertical BT partition, each sub-block can be partitioned into a horizontal BT. The two types of partitioning methods are different from each other in terms of the partitioning order, but the same in terms of the final partition shape.
[0113] In addition, when a block is partitioned, the order of searching the block can be defined in various ways. Generally, the search can be performed from left to right and from top to bottom, and the search of the block can mean the order of deciding whether to further partition each partitioned sub-block, the encoding order of each sub-block when each sub-block is no longer partitioned, or the search order when referring to the information of another neighboring block in the sub-block.
[0114] Figure 4 and Figure 5 illustrates an embodiment applying the present disclosure Figure 4 is a schematic block diagram of a transform unit 120, a quantization unit 130, an inverse quantization unit 140, and an inverse transform unit 150 in an encoder, and Figure 5 is a schematic block diagram of an inverse quantization unit 220 and an inverse transform unit 230 in a decoder.
[0115] Referring to Figure 4 the transform unit 120 and the quantization unit 130 may include a primary transform unit 121, a secondary transform unit 122, and a quantization unit 130. The inverse quantization unit 140 and the inverse transform unit 150 may include an inverse quantization unit 140, an inverse secondary transform unit 151, and an inverse primary transform unit 152.
[0116] Referring to Figure 5, the dequantization unit 220 and the inverse transformation unit 230 may include a dequantization unit 220, an inverse secondary transformation unit 231, and an inverse primary transformation unit 232.
[0117] In the present disclosure, when performing a transformation, the transformation may be performed in multiple steps. For example, two steps of a primary transformation and a secondary transformation may be applied as exemplified in Figure 4 , or more transformation steps may be used according to an algorithm. Here, the primary transformation may also be referred to as a core transformation.
[0118] The primary transformation unit 121 may apply a primary transformation to the residual signal, and here, the primary transformation may be defined in a table of an encoder and / or a decoder.
[0119] The primary transformation may employ a discrete cosine transform type 2 (hereinafter, referred to as "DCT2").
[0120] Alternatively, only in specific cases, a discrete sine transform type 7 (hereinafter, referred to as "DST7") may be employed. For example, DST7 may be applied to a 4×4 block in an intra prediction mode.
[0121] In addition, the primary transformation may employ a combination of various transforms of multiple transform selection (MTS) DST 7, DCT 8, DST 1, and DCT 5. For example, Figure 6 .
[0122] The secondary transformation unit 122 may apply a secondary transformation to the signal after the primary transformation, and here, the secondary transformation may be defined in a table of an encoder and / or a decoder.
[0123] As an implementation, the secondary transformation may conditionally employ an inseparable secondary transformation (hereinafter, referred to as "NSST"). For example, NSST may be applied only to intra prediction blocks and may have a set of transforms suitable for each prediction mode group.
[0124] Here, the prediction mode groups may be configured based on symmetry with respect to the prediction direction. For example, since prediction mode 52 and prediction mode 16 are symmetric based on prediction mode 34 (diagonal direction), the same set of transforms may be applied by forming a group. In this case, when applying the transform for prediction mode 52, since prediction mode 52 has the same set of transforms as prediction mode 16, the input data is transposed and then applied.
[0125] In addition, since there is no symmetry in the direction in the case of the planar mode and the DC mode, each mode has a different set of transforms, and the corresponding set of transforms may be composed of two transforms. With respect to the remaining direction modes, each set of transforms may be composed of three transforms.
[0126] As another embodiment, the secondary transform may employ a combination of various transforms DST 7, DCT 8, DST 1, and DCT 5 of multiple transform selection (MTS). For example, it may employ Figure 6 .
[0127] As another embodiment, DST 7 may be applied to the secondary transform.
[0128] As another embodiment, the secondary transform may be applied not to the entire primary transform block but only to a specific upper-left region. For example, when the block size is 8×8 or larger, 8×8 NSST is applied, and when the block size is less than 8×8, a 4×4 secondary transform may be applied. In this case, the block may be divided into 4×4 blocks, and then the 4×4 secondary transform may be applied to each divided block.
[0129] As another embodiment, even in the case of 4×N / N×4 (N >= 16), a 4×4 secondary transform may be applied.
[0130] The secondary transform (e.g., NSST), 4×4 secondary transform, and 8×8 secondary transform will be described in more detail with reference to Figures 12 to 15 and other embodiments in the specification.
[0131] The quantization unit 130 may perform quantization on the signal after the secondary transform.
[0132] The dequantization unit 140 and the inverse transform unit 150 perform the above processes conversely, and redundant descriptions thereof will be omitted.
[0133] Figure 5 is a schematic block diagram of the dequantization unit 220 and the inverse transform unit 230 in the decoder.
[0134] Referring to Figure 5 , the dequantization unit 220 and the inverse transform unit 230 may include a dequantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.
[0135] The dequantization unit 220 obtains transform coefficients from the entropy-decoded signal by using quantization step information.
[0136] The inverse secondary transform unit 231 performs an inverse secondary transform on the transform coefficients. Here, the inverse secondary transform represents the inverse transform of the secondary transform described in Figure 4 .
[0137] As another embodiment, the secondary transform may employ a combination of various transforms DST 7, DCT 8, DST 1, and DCT 5 of multiple transform selection (MTS). For example, it may employ Figure 6 .
[0138] The inverse primary transform unit 232 performs an inverse primary transform on the signal (or block) after the inverse secondary transform, and obtains a residual signal. Here, the inverse primary transform represents the inverse transform of the primary transform described in Figure 4 .
[0139] As an implementation, the primary transform may employ a combination of various transforms DST 7, DCT 8, DST 1, and DCT 5 of the multiple transform selection (MTS). For example, Figure 6 may be adopted.
[0140] As an implementation of the present disclosure, DST 7 may be applied to the primary transform.
[0141] As an implementation of the present disclosure, DST 8 may be applied to the primary transform.
[0142] The present disclosure provides a method for configuring a transform combination for each transform configuration group divided by at least one of a prediction mode, a block size, or a block shape, and the inverse primary transform unit 232 may perform an inverse transform based on the transform combination configured by the present disclosure. Additionally, the implementations described in the present disclosure may be applied.
[0143] Figure 6 is a table showing a transform configuration group that applies the implementation of the present disclosure and uses the multiple transform selection (MTS).
[0144] Transformation configuration group applying multiple transform selection (MTS)
[0145] In the present disclosure, the j-th transform combination candidate for the transform configuration group G i is represented by the pair shown in Equation 1 below.
[0146] [Equation 1]
[0147] (H(G i , j), V(G i , j))
[0148] Here, H(Gi,j) indicates the horizontal transform of the j-th candidate, and V(Gi,j) indicates the vertical transform of the j-th candidate. For example, in Figure 6 , it may be represented that H(G3,2) = DST7 and V(G3,2) = DCT8. According to the context, the value assigned to H(G i , j) or V(G i , j) may be a nominal value as in the above example to distinguish the transform, or may be an index value indicating the transform, or may be a two-dimensional (2D) matrix for the transform.
[0149] Additionally, in the present disclosure, the 2D matrix values of the DCT and DST may be represented as shown in Equations 2 and 3 below.
[0150] [Formula 2]
[0151] DCT Type 2: DCT Type 8:
[0152] [Formula 3]
[0153] DST Type 7: DST Type 4:
[0154] Here, whether DST or DCT is represented by S or C, the type number is represented in Roman numerals in the superscript, and N in the subscript represents an N×N transform. Additionally, 2D matrices such as and are assumed to have column vectors forming the transform basis.
[0155] Referring to Figure 6 , the transform configuration groups can be determined based on the prediction mode, and the number of groups can be a total of six groups G0 to G5. Additionally, G0 to G4 correspond to the cases where intra-frame prediction is applied, and G5 represents the transform combination (or transform set and transform combination set) applied to the residual block generated by inter-frame prediction.
[0156] A transform combination can be composed of a horizontal transform (or row transform) applied to the rows of the corresponding 2D block and a vertical transform (or column transform) applied to the columns.
[0157] Here, each of all the transform configuration groups can have four transform combination candidates. The four transform combinations can be selected or determined by the transform combination indices 0 to 3, and they are sent from the encoder to the decoder by encoding the transform combination indices.
[0158] As an implementation, according to the intra-frame prediction mode, the residual data (or residual signal) obtained by intra-frame prediction can have different statistical characteristics. Therefore, as Figure 6 illustrated in, transforms other than the general cosine transform can be applied to each intra-frame prediction mode.
[0159] Referring to Figure 6 , the cases of using 35 intra-frame prediction modes and the cases of using 67 intra-frame prediction modes are illustrated. Multiple transform combinations can be applied to each transform configuration group divided in each intra-frame prediction mode column. For example, these multiple transform combinations can be composed of four (row-direction transform and column-direction transform) combinations. As a specific example, DST-7 and DST-5 can be applied in the row (horizontal) direction and column (vertical) direction in group 0, and as a result, a total of four combinations are available.
[0160] Since a total of four transform kernel combinations can be applied to each intra prediction mode, a transform combination index for selecting one of the transform kernel combinations can be sent for each transform unit. In the present disclosure, the transform combination index may be referred to as an MTS index and denoted as mts_idx.
[0161] In addition, in addition to the transform kernels shown above Figure 6 It is also possible that due to the characteristics of the residual signal, DCT2 is optimal for both the row direction and the column direction. Therefore, an MTS flag is defined for each coding unit to perform the transform adaptively. Here, when the MTS flag is 0, DCT 2 can be applied to both the row direction and the column direction, and when the MTS flag is 1, one of the four combinations can be selected or determined by the MTS index.
[0162] As an implementation, when the MTS flag is 1, if the number of non-zero transform coefficients for a transform unit is not greater than a threshold, DST-7 can be applied to both the row direction and the column direction without applying Figure 6 the transform kernel. For example, the threshold can be set to 2, and this threshold can be set differently based on the block size or the size of the transform unit. This also applies to other implementations in the specification.
[0163] As an implementation, if the number of non-zero transform coefficients is not greater than the threshold, by first parsing the transform coefficient values, the amount of additional information transmission can be reduced by applying DST-7 without parsing the MTS index.
[0164] As an implementation, when the MTS flag is 1, if the number of non-zero transform coefficients is greater than the threshold of a transform unit, the MTS index can be parsed, and the horizontal transform and the vertical transform can be determined based on the MTS index.
[0165] As an implementation, MTS can be applied only when both the width and height of the transform unit are equal to or less than 32.
[0166] As an implementation, it can be pre-configured through offline training Figure 6 .
[0167] As an implementation, the MTS index can be defined as an index that can simultaneously indicate the horizontal transform and the vertical transform. Alternatively, the MTS index can be separately defined as a horizontal transform index and a vertical transform index.
[0168] As an implementation manner, the MTS flag or MTS index may be defined at at least one level of a sequence, a picture, a slice, a block, a coding unit, a transform unit, or a prediction unit. For example, the MTS flag or MTS index may be defined at at least one level of a sequence parameter set (SPS), a coding unit, or a transform unit. Additionally, as an example, a syntax flag for enabling / disabling MTS may be defined at at least one level of a sequence parameter set (SPS), a picture parameter set (PPS), or a slice header.
[0169] As another implementation manner, a transform combination (horizontal transform or vertical transform) corresponding to a transform index may be configured without depending on an MTS flag, a prediction mode, and / or a block shape. For example, the transform combination may be constituted by at least one of DCT2, DST7, and / or DCT8. As a specific example, when the transform index is 0, 1, 2, 3, or 4, each transform combination may be (DCT2, DCT2), (DST7, DST7), (DCT8, DST7), (DST7, DCT8), or (DCT8, DCT8).
[0170] Figure 7 is a flowchart showing an encoding process of performing multiple transform selection (MTS) as an implementation manner of applying the present disclosure.
[0171] In the present disclosure, an implementation manner in which a transform is applied to the horizontal direction and the vertical direction respectively is basically described, but the transform combination may be configured as an inseparable transform.
[0172] Alternatively, the transform combination may be constituted by a mixture of a separable transform and an inseparable transform. In this case, when an inseparable transform is used, row / column transform selection or horizontal / vertical direction selection may not be required, and only when a separable transform is selected, the Figure 6 transform combination can be used.
[0173] Additionally, the scheme proposed by the present disclosure may be applied regardless of whether it is a primary transform or a secondary transform. That is, there is no limitation that the scheme should be applied only to either the primary transform or the secondary transform, and the scheme may be applied to both the primary transform and the secondary transform. Here, the primary transform may mean a transform for first transforming a residual block, and the secondary transform may mean a transform for applying a transform to a block generated as a result of the primary transform.
[0174] First, the encoder may determine a transform configuration group corresponding to the current block. Here, the transform configuration group may mean Figure 6 the transform configuration group of, the present disclosure is not limited thereto, and the transform configuration group may be constituted by other transform combinations.
[0175] The encoder may perform a transform on candidate transform combinations available in a transform configuration group (S720).
[0176] As a result of performing the transform, the encoder may determine or select a transform combination having a minimum rate distortion (RD) cost (S730).
[0177] The encoder may encode a transform combination index corresponding to the selected transform combination (S740).
[0178] Figure 8 FIG. is a flowchart showing a decoding process of performing multiple transform selection (MTS) as an embodiment of applying the present disclosure.
[0179] First, the decoder may determine a transform configuration group of a current block (S810).
[0180] The decoder may parse (or obtain) a transform combination index from a video signal, and here, the transform combination index may correspond to any one of multiple transform combinations in the transform configuration group (S820). For example, the transform configuration group may include discrete sine transform type (DST) 7 and discrete cosine transform type (DCT) 8. The transform combination index may be referred to as an MTS index.
[0181] As an embodiment, the transform configuration group may be configured based on at least one of a prediction mode, a block size, or a block shape of the current block.
[0182] The decoder may derive a transform combination corresponding to the transform combination index (S830). Here, the transform combination may be composed of a horizontal transform and a vertical transform, and may include at least one of DST-7 or DCT-8.
[0183] In addition, the transform combination may mean the transform combination described in Table 6, but the present disclosure is not limited thereto. That is, depending on other embodiments in the present disclosure, the transform combination may be composed of other transform combinations.
[0184] The decoder may perform an inverse transform on the current block based on the transform combination (S840). When the transform combination is composed of a row (horizontal) transform and a column (vertical) transform, the column (vertical) transform may be applied after first applying the row (horizontal) transform. However, the present disclosure is not limited thereto, and the transform order may be reversed, or when the transform combination is composed of an inseparable transform, the inseparable transform may be applied immediately.
[0185] As an embodiment, when the vertical transform or the horizontal transform is DST-7 or DCT-8, the inverse transform of DST-7 or DCT-8 may be applied to each column and then to each column.
[0186] As an implementation manner, for vertical transformation or horizontal transformation, different transformations can be applied to each row and / or each column.
[0187] As an implementation manner, a transformation combination index can be obtained based on an MTS flag indicating whether to perform MTS. That is, when MTS is performed according to the MTS flag, a transformation combination index can be obtained.
[0188] As an implementation manner, the decoder can check whether the number of non-zero transformation coefficients is greater than a threshold. In this case, a transformation can be obtained when the number of non-zero transformation coefficients is greater than the threshold.
[0189] As an implementation manner, the MTS flag or MTS index can be defined at at least one level of sequence, picture, slice, block, coding unit, transformation unit, or prediction unit.
[0190] As an implementation manner, an inverse transformation can be applied only when both the width and height of the transformation unit are equal to or less than 32.
[0191] On the other hand, as another implementation manner, the process of determining a transformation configuration group and the process of parsing a transformation combination index can be performed simultaneously. Alternatively, step S810 can be pre-configured and omitted in the encoder and / or decoder.
[0192] Figure 9 is a flowchart for describing the encoding process of the MTS flag and MTS index as an implementation manner of applying the present disclosure.
[0193] The encoder can determine whether to apply multiple transformation selection (MTS) to the current block (S910).
[0194] When applying multiple transformation selection (MTS), the encoder can encode MTS flag = 1 (S920).
[0195] In addition, the encoder can determine an MTS index based on at least one of the prediction mode, horizontal transformation, and vertical transformation of the current block (S930). Here, the MTS index can mean an index indicating any one of multiple transformation combinations for each intra prediction mode, and the MTS index can be sent for each transformation unit.
[0196] When the MTS index is determined, the encoder can encode the MTS index (S940).
[0197] On the other hand, when not applying multiple transformation selection (MTS), the encoder can encode MTS flag = 0 (S920).
[0198] Figure 10It is a flowchart for describing a decoding process of applying a horizontal transformation or a vertical transformation to a row or a column based on an MTS flag and an MTS index as an embodiment of the present disclosure.
[0199] The decoder can parse the MTS flag from the bitstream (S1010). Here, the MTS flag can indicate whether to apply multiple transform selection (MTS) to the current block.
[0200] The decoder can determine whether to apply multiple transform selection (MTS) to the current block based on the MTS flag (S1020). For example, it can be checked whether the MTS flag is 1.
[0201] When the MTS flag is 1, the decoder can check whether the number of non-zero transform coefficients is greater than (or equal to or greater than) a threshold (S1030). For example, the threshold can be set to 2, and this threshold can be set differently based on the block size or the size of the transform unit.
[0202] When the number of non-zero transform coefficients is greater than the threshold, the decoder can parse the MTS index (S1040). Here, the MTS index can mean any one of multiple transform combinations for each intra prediction mode, and the MTS index can be sent for each transform unit. Alternatively, the MTS index can mean an index indicating any one of the transform combinations defined in a pre-configured transform combination table, and here, the pre-configured transform combination table can mean Figure 6 , but the present disclosure is not limited thereto.
[0203] The decoder can derive or determine the horizontal transformation and the vertical transformation based on at least one of the MTS index and the prediction mode (S1050).
[0204] Alternatively, the decoder can derive the transform combination corresponding to the MTS index. For example, the decoder can derive or determine the horizontal transformation and the vertical transformation corresponding to the MTS index.
[0205] When the number of non-zero transform coefficients is not greater than the threshold, the decoder can apply a pre-configured vertical inverse transform (S1060). For example, the vertical inverse transform can be the inverse transform of DST7.
[0206] In addition, the decoder can apply a pre-configured horizontal inverse transform to each row (S1070). For example, the horizontal inverse transform can be the inverse transform of DST7. That is, when the number of non-zero transform coefficients is not greater than the threshold, a transform kernel pre-configured by the encoder or the decoder can be used. For example, a transform kernel that is not defined in the Figure 6 illustrated transform combination table but is widely used (such as DCT-2, DST-7, or DCT-8) can be used.
[0207] In addition, when the MTS flag is 0, the decoder can apply a pre-configured inverse vertical transform to each column (S1080). For example, the inverse vertical transform can be the inverse transform of DCT2.
[0208] In addition, the decoder can apply a pre-configured inverse horizontal transform to each row (S1090). For example, the inverse horizontal transform can be the inverse transform of DCT2. That is, when the MTS flag is 0, a transform kernel pre-configured by the encoder or the decoder can be used. For example, a transform kernel that is not defined in the transform combination table illustrated in Figure 6 but is widely used can be used.
[0209] Figure 11 is a flowchart for performing an inverse transform based on transform-related parameters in an embodiment applying the present disclosure.
[0210] The decoder applying the present disclosure can obtain sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag (S1110). Here, sps_mts_intra_enabled_flag indicates whether there is a tu_mts_flag in the residual coding syntax of an intra-coded unit. For example, when sps_mts_intra_enabled_flag = 0, there is no tu_mts_flag in the residual coding syntax of the intra-coded unit, and when sps_mts_intra_enabled_flag = 0, there is a tu_mts_flag in the residual coding syntax of the intra-coded unit. In addition, sps_mts_inter_enabled_flag indicates whether there is a tu_mts_flag in the residual coding syntax of an inter-coded unit. For example, when sps_mts_inter_enabled_flag = 0, there is no tu_mts_flag in the residual coding syntax of the intra-coded unit, and when sps_mts_inter_enabled_flag = 0, there is a tu_mts_flag in the residual coding syntax of the inter-coded unit.
[0211] The decoder can obtain tu_mts_flag based on sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag (S1120). For example, when sps_mts_intra_enabled_flag = 1 or sps_mts_inter_enabled_flag = 1, the decoder can obtain tu_mts_flag. Here, tu_mts_flag indicates whether multiple transform selection (hereinafter referred to as "MTS") is applied to the residual samples of the luma transform block. For example, when tu_mts_flag = 0, MTS is not applied to the residual samples of the luma transform block, and when tu_mts_flag = 1, MTS is applied to the residual samples of the luma transform block.
[0212] As another example, at least one of the embodiments in this document can be applied to tu_mts_flag.
[0213] The decoder can obtain mts_idx based on tu_mts_flag (S1130). For example, when tu_mts_flag = 1, the decoder can obtain mts_idx. Here, mts_idx indicates which transform kernel is applied to the luma residual samples along the horizontal and / or vertical directions of the current transform block.
[0214] For example, at least one of the embodiments in this document can be applied to mts_idx. As a specific example, the embodiments that can be applied Figure 6 are at least one of the embodiments.
[0215] The decoder can derive the transform kernel corresponding to mts_idx (S1140). For example, it can be defined by dividing the transform kernel corresponding to mts_idx into horizontal transform and vertical transform.
[0216] As another example, different transform kernels can be applied to the horizontal transform and the vertical transform. However, the present disclosure is not limited thereto, and the same transform kernel can be applied to the horizontal transform and the vertical transform.
[0217] As an embodiment, mts_idx can be defined as shown in Table 1 below.
[0218] [Table 1]
[0219] mts_idx[x0][y0] trTypeHor trTypeVer 0 0 0 1 1 1 2 2 1 3 1 2 4 2 2
[0220] In addition, the decoder can perform inverse transform based on the transform kernel (S1150).
[0221] Above Figure 11Among them, the following implementation manners are mainly described: obtaining tu_mts_flag to determine whether to apply MTS, and obtaining mts_idx according to the subsequently obtained value of tu_mts_flag to determine the transform kernel, but the present disclosure is not limited thereto. As an example, the decoder directly parses mts_idx without parsing tu_mts_flag to determine the transform kernel. In this case, Table 1 above can be used. That is, when the mts_idx value indicates 0, DCT-2 can be applied in the horizontal / vertical direction, and when the mts_idx value indicates a value other than 0, DST-7 and / or DCT-8 can be applied according to the mts_idx value.
[0222] As another implementation manner of the present disclosure, a decoding process for performing transform processing is described.
[0223] The decoder may check the transform size nTbS (S10). Here, the transform size nTbS may be a variable representing the horizontal sample size of the scaled transform coefficients.
[0224] The decoder may check the transform kernel type trType (S20). Here, the transform kernel type trType may be a variable representing the transform kernel type, and various implementation manners of this document may be applied. The transform kernel type trType may include a horizontal transform kernel type trTypeHor and a vertical transform kernel type trTypeVer.
[0225] Referring to Table 1, when the transform kernel type trType is 0, the transform kernel type may represent DCT2, when the transform kernel type trType is 1, the transform kernel type may represent DST7, and when the transform kernel type trType is 2, the transform kernel type may represent DCT8.
[0226] The decoder may perform transform matrix multiplication (S30) based on at least one of the transform size nTbS or the transform kernel type.
[0227] As another example, when the transform kernel type is 1 and the transform size is 4, a predetermined transform matrix 1 may be applied when performing transform matrix multiplication.
[0228] As another example, when the transform kernel type is 1 and the transform size is 8, a predetermined transform matrix 2 may be applied when performing transform matrix multiplication.
[0229] As another example, when the transform kernel type is 1 and the transform size is 16, a predetermined transform matrix 3 may be applied when performing transform matrix multiplication.
[0230] As another example, when the transform kernel type is 1 and the transform size is 32, a predefined transform matrix 4 can be applied when performing the transform matrix multiplication.
[0231] Similarly, when the transform kernel type is 2 and the transform sizes are 4, 8, 16, or 32, predefined transform matrices 5, 6, 7, and 8 can be applied, respectively.
[0232] Here, each of the predefined transform matrices 1 to 8 can correspond to any one of various types of transform matrices. As an example, a transform matrix of the type exemplified in Figure 6 can be applied.
[0233] The decoder can derive the transformed samples (or transform coefficients) based on the transform matrix multiplication (S40).
[0234] Each of the above embodiments can be used, but the present disclosure is not limited thereto, and it can be used in combination with the above embodiments and other embodiments of the present disclosure.
[0235] Figure 12 is a table showing the assignment of transform sets for each intra prediction mode in NSST as an embodiment of applying the present disclosure.
[0236] Non-separable second-order transform (NSST)
[0237] The secondary transform unit can apply a secondary transform to the signal after the primary transform, and here, the secondary transform can be defined in the tables of the encoder and / or decoder.
[0238] As an embodiment, the secondary transform can conditionally employ an inseparable secondary transform (hereinafter, referred to as "NSST"). For example, NSST can be applied only to intra prediction blocks and can have transform sets suitable for each prediction mode group.
[0239] Here, the prediction mode groups can be configured based on the symmetry with respect to the prediction direction. For example, since prediction mode 52 and prediction mode 16 are symmetric based on prediction mode 34 (diagonal direction), the same transform set can be applied by forming a group. In such a case, when applying the transform for prediction mode 52, since prediction mode 52 has the same transform set as prediction mode 16, the input data is transposed and then applied.
[0240] In addition, since there is no symmetry in the direction in the case of the planar mode and the DC mode, each mode has a different transform set, and the corresponding transform set can be composed of two transforms. With respect to the remaining direction modes, each transform set can be composed of three transforms. However, the present disclosure is not limited thereto, and each transform set can be composed of multiple transforms.
[0241] In an embodiment, a transform set table other than the transform set table exemplified in Figure 12 may be defined. For example, as shown in Table 2 below, a transform set may be determined from a predefined table according to an intra prediction mode (or a group of intra prediction modes). Syntax indicating a specific transform in the transform set determined according to the intra prediction mode may be signaled from an encoder to a decoder.
[0242] [Table 2]
[0243] IntraPredMode Transform set index IntraPredMode<0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= IntraPredMode 1
[0244] Referring to Table 2, predefined transform sets may be assigned to groups of intra prediction modes (or groups of intra prediction modes). Here, the IntraPredMode value may be a mode value transformed in consideration of wide-angle intra prediction (WAIP).
[0245] Figure 13 is a flowchart of the calculation of the Givens rotation as an embodiment of applying the present disclosure.
[0246] As another embodiment, the secondary transform may be applied not to the entire primary transform block but only to the upper left 8×8 region. For example, when the block size is 8×8 or larger, 8×8 NSST is applied, and when the block size is less than 8×8, 4×4 NSST is applied, and in this case, the block is divided into 4×4 blocks, and then 4×4 NSST is applied to each of the divided blocks.
[0247] As another embodiment, 4×4 NSST may be applied even in the case of 4×N / N×4 (N >= 16).
[0248] Since both 8×8 NSST and 4×4 NSST follow the transform combination configuration described in this document and are non-separable transforms, 8×8 NSST receives 64 data and outputs 64 data, and 4×4 NSST has 16 inputs and 16 outputs.
[0249] Both 8×8 NSST and 4×4 NSST are configured by a hierarchical combination of Givens rotations. The matrix corresponding to one Givens rotation is shown in Equation 4, and the matrix product is shown in Equation 5 below.
[0250] [Equation 4]
[0251]
[0252] [Equation 5]
[0253] t m = x m cosθ - x n sinθ
[0254] t n = x m sinθ - x n cosθ
[0255] As Figure 13 illustrated in, since one Givens rotation rotates two data, in order to process 64 data (for 8×8 NSST) or 16 data (for 4×4 NSST), a total of 32 or 8 Givens rotations are required.
[0256] Therefore, a bundle of 32 or 8 is used to form the Givens rotation layer. The output data of one Givens rotation layer is transmitted as the input data of the next Givens rotation layer through the determined permutation.
[0257] Figure 14 Illustrated is one-round configuration in the 4×4 NSST composed of a Givens rotation layer and a permutation as an embodiment applying the present disclosure.
[0258] Referring to Figure 14 , illustrated is processing four Givens rotation layers in sequence in the case of 4×4 NSST. As Figure 14 illustrated in, the output data of one Givens rotation layer is transmitted as the input data of the next Givens rotation layer through the determined permutation (or shuffling).
[0259] As Figure 14 illustrated in, the pattern to be permuted is determined regularly, and in the case of 4×4 NSST, four Givens rotation layers and their corresponding permutations are combined to form one round.
[0260] In the case of 8×8 NSST, six Givens rotation layers and corresponding permutations form one round. The 4×4 NSST goes through two rounds, and the 8×8 NSST goes through four rounds. Different rounds use the same permutation pattern, but the Givens rotation angles applied are different. Therefore, it is necessary to store the angle data of all Givens rotations constituting each transformation.
[0261] As the last step, finally, a permutation is further performed on the data output through the Givens rotation layer, and the corresponding permutation information is stored separately for each transformation. In the forward NSST, the corresponding permutation is performed finally, while on the contrary, in the reverse NSST, the corresponding reverse permutation is applied first.
[0262] In the case of the reverse NSST, the Givens rotation layer and the permutation applied to the forward NSST are performed in the reverse order, and the angle of each Givens rotation is even rotated by taking the negative value.
[0263] Figure 15It is a block diagram for describing operations of forward and inverse simplification transforms which are embodiments of applying the present disclosure.
[0264] Reduced second-order transform (RST)
[0265] When assuming that the orthogonal matrix representing a transform has an N×N form, among N transform basis vectors (R < N), the simplification transform (hereinafter referred to as "RT") only leaves R transform basis vectors. The matrix of the forward RT for generating transform coefficients is given by Equation 6 below.
[0266] [Equation 6]
[0267]
[0268] Since the matrix of the inverse RT becomes the transpose matrix of the forward RT matrix, the applications of the forward RT and the inverse RT are exemplified as in Figure 15 . Here, the simplification factor is defined as R / N (R < N).
[0269] The number of elements of the simplification transform is R×N, which is smaller than the size of the entire matrix (N×N). In other words, the required matrix is R / N of the entire matrix. Additionally, the required number of multiplications is R×N, which is lower by R / N than the original N×N. When applying the simplification transform, R coefficients are provided, and as a result, only R coefficient values can be sent instead of N coefficients.
[0270] Assume a case where RT is applied to the upper left 8×8 block of a transform block that has undergone a primary transform. This RT can be referred to as an 8×8 reduced secondary transform (8×8RST).
[0271] When the R value in Equation 6 above is 16, the forward 8×8RST has a 16×64 matrix form, and the inverse 8×8RST has a 64×16 matrix form.
[0272] In addition, a transform set configuration identical to the one exemplified in Figure 12 can even be applied to 8×8RST. That is, the corresponding 8×8RST can be applied according to the transform set in Figure 12 .
[0273] As an embodiment, when two or three transforms form a transform set according to the intra prediction mode in Figure 12 , it can be configured to select one of up to 4 transforms including the case where no secondary transform is applied. Here, one transform can be regarded as an identity matrix.
[0274] When indices 0, 1, 2, and 3 are assigned to four transforms respectively, a syntax element called the NSST index can be signaled for each transform block, thereby specifying the corresponding transform. That is, in the case of NSST, an 8×8 NSST can be specified for the 8×8 upper-left block through the NSST index, and an 8×8 RST can be specified in the RST configuration. Additionally, in this case, index 0 can be assigned to the case where the identity matrix, i.e., the quadratic transform, is not applied.
[0275] When the forward 8×8 RST shown in Equation 6 is applied, 16 valid transform coefficients are generated. As a result, it can be considered that the 64 input data constituting the 8×8 region are reduced to 16 output data. From the perspective of a two-dimensional region, only one-fourth of the region is filled with valid transform coefficients. Therefore, Figure 16 the 4×4 upper-left region in
[0276] Figure 16 is a diagram illustrating the process of performing a backward scan from the 64th to the 17th in the backward scan order as an embodiment applying the present disclosure.
[0277] Figure 16 illustrates the scan from the 17th coefficient to the 64th coefficient in the forward scan order (when the forward scan order starts from 1). However, Figure 16 illustrates the backward scan, and it illustrates performing the backward scan from the 64th coefficient to the 17th coefficient.
[0278] Referring to Figure 16 , the upper-left 4×4 region is the region of interest (ROI), valid transform coefficients are assigned to the ROI, and the remaining regions are empty. That is, by default, the value 0 can be assigned to the remaining regions.
[0279] If there are valid transform coefficients other than 0 in the regions of Figure 16 except for the ROI region, this means that the 8x8 RST is not applied. As a result, in this case, the corresponding NSST index coding can be omitted.
[0280] Conversely, if there are no non-zero transform coefficients outside the ROI region of Figure 16 (when applying the 8×8 RST and assigning 0 to the regions other than the ROI), then it is possible that the 8×8 RST is applied. As a result, the NSST index can be encoded.
[0281] Thus, since the existence of non-zero transform coefficients must be checked, conditional NSST index coding can be performed after the residual coding process.
[0282] The present disclosure provides an RST design method and an associated optimization method that can be applied to 4×4 blocks from an RST structure. In addition to 4×4 RSTs, the embodiments disclosed in the present disclosure can also be applied to 8×8 RSTs or another type of transform.
[0283] Figure 17 Illustrated are three forward scan orders of a transform coefficient block (transform block) that is an embodiment applying the present disclosure.
[0284] Embodiment 1: RST applicable to 4×4 blocks
[0285] An inseparable transform that can be applied to one 4×4 block is a 16×16 transform. That is, when data elements constituting a 4×4 block are arranged in a row-major or column-major order, a 16×1 vector is used to apply the inseparable transform.
[0286] The forward 16×16 transform is composed of 16 row-direction transform basis vectors, and when an inner product is applied to the 16×1 vector and each transform basis vector, transform coefficients for the transform basis vectors are obtained. The process of obtaining transform coefficients corresponding to all 16 transform basis vectors is equivalent to multiplying the input 16×1 vector by a 16×16 inseparable transform matrix.
[0287] The transform coefficients obtained through matrix multiplication have the form of a 16×1 vector, and for each transform coefficient, the statistical characteristics may be different. For example, when a 16×1 transform coefficient vector is composed of the 0th element to the 15th element, the variance of the 0th element may be greater than the variance of the 15th element. In other words, as the element position is earlier, the corresponding variance value of the element is larger, such that the element may have a larger energy value.
[0288] When applying an inverse 16×16 inseparable transform to the 16×1 transform coefficients, the original 4×4 block signal can be restored. When the forward 16×16 inseparable transform is an orthogonal transform, the corresponding inverse 16×16 transform can be obtained through the transpose matrix for the forward 16×16 transform.
[0289] When the 16×1 transform coefficient vector is multiplied by the inverse 16×16 inseparable transform matrix, data in the form of a 16×1 vector can be obtained, and when the obtained data is arranged in the row-major or column-major order applied first, the 4×4 block signal can be restored.
[0290] As described above, the elements constituting the 16×1 transform coefficient vector can have different statistical characteristics.
[0291] If the transform coefficients arranged on the front side (close to the 0th element) have large energy, a signal very close to the original signal can be restored even by applying the inverse transform to some of the first-occurring transform coefficients without using all the transform coefficients. For example, when the inverse 16×16 non-separable transform consists of 16 column basis vectors, only L column basis vectors are left to form a 16×L matrix. In addition, when the 16×L matrix and L×1 are multiplied by each other after only L important transform coefficients are left among the transform coefficients (L×1 vector), a 16×1 vector with a small error from the 16×1 vector data of the original input can be restored.
[0292] As a result, since only L coefficients are used for data restoration, when obtaining the transform coefficients, an L×1 transform coefficient vector is obtained instead of a 16×1 transform coefficient vector. That is, when configuring an L×16 transform by selecting L corresponding row direction vectors in the forward 16×16 non-separable transform matrix and multiplying the configured L×16 transform by a 16×1 input vector, L important transform coefficients can be obtained.
[0293] The range of the L value is 1 ≤ L < 16, and generally, L vectors can be selected from among the 16 transform basis vectors by any method, but from the perspectives of encoding and decoding, selecting the transform basis vectors with high importance in terms of signal energy can be advantageous in terms of encoding efficiency.
[0294] Embodiment 2: Configuration of the application area of 4×4 RST and arrangement of transform coefficients
[0295] 4×4RST can be applied as a secondary transform and can be additionally applied to the blocks to which a primary transform such as DCT type 2 is applied. When the size of the block to which the primary transform is applied is N×N, the size of the block to which the primary transform is applied is generally larger than 4×4. Therefore, when applying 4×4RST to an N×N block, there are the following two methods.
[0296] Embodiment 2-1) 4×4RST is not applied to all N×N regions, but can be applied only to certain regions. For example, 4×4RST can be applied only to the upper-left M×M region (M ≤ N).
[0297] Embodiment 2-2) The region to which the secondary transform is to be applied can be divided into 4×4 blocks, and then 4×4RST can be applied to each divided block.
[0298] As an embodiment, Embodiment 2-1) and 2-2) can be mixed and applied. For example, only the upper-left M×M region can be divided into 4×4 blocks, and then 4×4RST can be applied.
[0299] As an implementation, the secondary transformation can be applied only to the upper-left 8×8 region. When the N×N block is equal to or larger than 8×8, 8×8 RST can be applied. When the N×N block is smaller than 8×8 (4×4, 8×4, and 4×8), the N×N block can be divided into 4×4 blocks, and then 4×4 RST can be applied to each of the 4×4 blocks, as in Embodiment 2-2. Additionally, even in the case of 4×N / N×4 (N >= 16), 4×4 NSST can be applied.
[0300] When L (1 ≤ L < 16) transform coefficients are generated after applying 4×4 RST, degrees of freedom for arranging the L transform coefficients are generated. However, since there will be a predetermined order when processing the transform coefficients in the residual coding step, the coding performance may vary depending on the arrangement of the L transform coefficients in the 2D block.
[0301] For example, in the case of residual coding in HEVC, the coding starts from the position farthest from the DC position. This is to enhance the coding performance by taking advantage of the fact that the quantization coefficient values become zero or approach zero as they move away from the DC position.
[0302] Therefore, in terms of coding performance, it may be advantageous to arrange the more important coefficients with high energy for the L transform coefficients so that the L transform coefficients are subsequently coded in the order of residual coding.
[0303] Figure 17 Illustrated are three forward scan orders in units of 4×4 transform blocks (coefficient groups (CG)) applied in HEVC. Residual coding follows Figure 17 the reverse order of the scan order (i.e., coding is performed in the order from 16 to 1).
[0304] Since the three scan orders presented in Figure 17 are selected according to the intra prediction mode, the present disclosure can be configured to similarly determine the scan order for even the L transform coefficients according to the intra prediction mode.
[0305] Figure 18 Illustrated are the positions and forward scan order of the valid transform coefficients in each 4×4 block when diagonal scan is applied and 4×4 RST is applied to the upper-left 4×8 block as an implementation of applying the present disclosure.
[0306] When following the Figure 17 diagonal scan order in Figure 18 and dividing the upper-left 4×8 block into 4×4 blocks and applying 4×4 RST to each 4×4 block, if the L value is 8 (i.e., if only 8 out of 16 transform coefficients are left), the transform coefficients can be set as in
[0307] Only half of each 4×4 block can have transform coefficients, and by default, a value of 0 can be applied to the positions marked with X.
[0308] Therefore, residual coding is performed by arranging L transform coefficients for each 4×4 block according to the Figure 17 scanning order illustrated in and assuming that the remaining (16–L) positions of each 4×4 block are filled with zeros.
[0309] Figure 19 Illustrated is the case where, when applying diagonal scanning and 4×4 RST in the upper left 4×8 block as an embodiment of applying the present disclosure, the valid transform coefficients of two 4×4 blocks are combined into one 4×4 block.
[0310] Referring to Figure 19 , the L transform coefficients arranged in two 4×4 blocks can be combined into one. Specifically, when the value of L is 8, since the transform coefficients of the two 4×4 blocks are combined and completely fill one 4×4 block, no transform coefficients are left in the other 4×4 block.
[0311] Therefore, since most residual coding is not required for empty 4×4 blocks, the corresponding coded_sub_block_flag can be coded as 0.
[0312] In addition, as an embodiment of the present disclosure, various schemes can even be applied to how to mix the transform coefficients of two 4×4 blocks. The transform coefficients can be combined according to a random order, but the present disclosure can provide the following methods.
[0313] 1) The transform coefficients of two 4×4 blocks are alternately mixed in the scanning order. That is, in Figure 18 , when the transform coefficients of the top block are and and the transform coefficients of the bottom block are and , the transform coefficients can be alternately mixed one by one like and . Alternatively, the order of and can be changed. In other words, the order can be configured such that appears first.
[0314] 2) The transform coefficients for the first 4×4 block can be arranged first, and then the transform coefficients for the second 4×4 block can be arranged. In other words, the transform coefficients can be concatenated and arranged like … and . Alternatively, the order can be like … and Changed in the same way.
[0315] Embodiment 3: Method for encoding NSST indices for 4×4 RST
[0316] When applying 4×4 RST as exemplified in Figure 18 it is possible to fill positions 16 to L+1 with 0 values according to the transform coefficient scan order of each 4×4 block.
[0317] Therefore, when non-zero values are generated at positions 16 to L+1 even in one of two 4×4 blocks, it can be known that this is a case where 4×4 RST is not applied.
[0318] When 4×4 RST also has a structure in which one of the transform sets prepared for NSST is selected and applied, it is possible to signal the transform index to which the transform is to be applied (in this embodiment, it may be referred to as the NSST index).
[0319] It is assumed that any decoder can know the NSST index by bitstream parsing, and the parsing is performed after residual decoding.
[0320] When residual decoding is performed and it is confirmed that there is at least one non-zero transform coefficient between positions L+1 and 16, 4×4 RST is not applied, so the NSST index can be configured not to be parsed.
[0321] Therefore, the NSST index is selectively parsed only when it is necessary to reduce signaling costs.
[0322] If, as Figure 18 exemplified in, 4×4 RST is applied to multiple 4×4 blocks in a specific region (for example, the same 4×4 RST can be applied to all of multiple 4×4 blocks, or different 4×4 RSTs can be applied), it is possible to specify the 4×4 RST applied to all 4×4 blocks by one NSST index. In this case, the same 4×4 RST can be specified, or the 4×4 RST applied to each of all 4×4 blocks can be specified.
[0323] Since it is determined whether 4×4 RST is applied to all 4×4 blocks by one NSST index, it is possible to check whether there are non-zero transform coefficients at positions 16 to L+1 of all 4×4 blocks during the residual decoding process. As a result of the check, when non-zero transform coefficients exist at positions (positions 16 to L+1) that are not acceptable even in one 4×4 block, the NSST index can be configured not to be encoded.
[0324] The NSST index can be signaled separately for the luminance block and the chrominance block. In the case of the chrominance block, separate NSST indexes can be signaled for Cb and Cr, and one NSST index can be shared.
[0325] When Cb and Cr share one NSST index, the 4×4 RST specified by the same NSST index can be applied. In this case, the 4×4 RST for Cb and Cr can be the same, or the NSST indexes can be the same, but separate 4×4 RSTs can be provided.
[0326] To apply conditional signaling to the shared NSST index, check whether there are non-zero transform coefficients at positions from the (L + 1)-th position to the 16-th position in all 4×4 blocks of Cb and Cr. When there are non-zero transform coefficients, the NSST index can be configured not to be signaled.
[0327] As Figure 19 illustrated, even for the case where the transform coefficients of two 4×4 blocks are combined, when applying the 4×4 RST, check whether there are non-zero transform coefficients at positions where there are no valid transform coefficients, and then it can be determined whether to signal the NSST.
[0328] For example, as Figure 19 illustrated in (b) of, when the L value is 8 and there are no valid transform coefficients in a 4×4 block (the block marked with X) when applying the 4×4 RST, the coded_sub_block_flag of the block where there are no valid transform coefficients can be checked. In this case, when the coded_sub_block_flag is 1, the NSST index can be configured not to be signaled.
[0329] Embodiment 4: Optimization method for the case where encoding of NSST indices is performed before residual encoding
[0330] When encoding the NSST index before residual coding, whether to apply the 4×4 RST is predetermined. As a result, for positions where 0 is assigned to the transform coefficients, residual coding can be omitted.
[0331] Here, whether to apply the 4×4 RST can be configured to be known through the NSST index. For example, when the NSST index is 0, the 4×4 RST is not applied.
[0332] Alternatively, the NSST index can be signaled through a separate syntax element (e.g., the NSST flag). For example, if the separate syntax element is called the NSST flag, first parse the NSST flag to determine whether to apply the 4×4 RST. If the NSST flag value is 1, residual coding can be omitted for positions where there cannot be valid transform coefficients.
[0333] As an implementation manner, when performing residual coding, first, the position of the last non-zero transform coefficient on the TU is encoded. When encoding the NSST index is performed after encoding the position of the last non-zero transform coefficient and assuming that 4×4RST is applied to the position of the last non-zero transform coefficient, if the position of the last non-zero transform coefficient is determined to be a position where no non-zero transform coefficient can be generated, 4×4RST can be configured not to be applied to the position of the last non-zero transform coefficient without encoding the NSST index.
[0334] For example, since in the case of the positions marked with X in Figure 18 , no valid transform coefficients are set when 4×4RST is applied (for example, these positions can be filled with zero values), so when the last non-zero transform coefficient is in the area marked with X, encoding of the NSST index can be omitted. When the last non-zero transform coefficient is not in the area marked with X, encoding of the NSST index can be performed.
[0335] As an implementation manner, when checking whether to apply 4×4RST by conditionally encoding the NSST index after encoding the position of the last non-zero transform coefficient, the remaining residual coding part can be processed by the following two schemes.
[0336] 1) When 4×4RST is not applied, the general residual coding is kept as it is. That is, encoding is performed assuming that non-zero transform coefficients may exist at any position from the non-zero transform coefficient position to the DC.
[0337] 2) When 4×4RST is applied, since there are no transform coefficients at specific positions or specific 4×4 blocks (for example, the X positions in Figure 18 , which can be filled with 0 by default), the residual of the corresponding position or block can not be performed.
[0338] For example, in the case of reaching the positions marked with X in Figure 18 , encoding of the sig_coeff_flag can be omitted. Here, the sig_coeff_flag means a flag indicating whether there is a non-zero transform coefficient at the corresponding position.
[0339] When, as exemplified in Figure 19 , the transform coefficients of two blocks are combined, encoding of the coded_sub_block_flag can be omitted for the 4×4 blocks assigned 0, and the corresponding value can be deduced as 0, and all corresponding 4×4 blocks can be deduced as 0 values without separate encoding.
[0340] In the case of encoding the NSST index after encoding the positions of non-zero transform coefficients, when the x-position P x and y-position P y of the last non-zero transform coefficient are respectively less than T x and T y , the NSST index encoding can be omitted and the 4×4RST may not be applied.
[0341] For example, the case where T x = 1 and T y = 1 means that for the case where there is a non-zero transform coefficient at the DC position, the NSST index encoding is omitted.
[0342] The scheme of determining whether to encode the NSST index by comparing with a threshold can be differently applied to luminance and chrominance. For example, different T x and T y can be applied to luminance and chrominance, and the threshold can be applied to luminance but not to chrominance. Or vice versa.
[0343] The above two methods can be applied simultaneously, that is, the first method of omitting the NSST index encoding when the non-zero transform coefficient is in a region where there is no valid transform coefficient, and the second method of omitting the NSST index encoding when each of the X coordinate and Y coordinate of the non-zero transform coefficient is less than a predetermined threshold.
[0344] For example, the threshold of the position coordinates of the last non-zero transform coefficient can be checked first, and then it can be checked whether the last non-zero transform coefficient is in a region where there is no valid transform coefficient. Alternatively, the order can be changed.
[0345] The method proposed in Embodiment 4 can even be applied to 8×8RST. That is, when the last non-zero transform coefficient is in a region of the upper left 8×8 area other than the upper left 4×4, the encoding of the NSST index can be omitted, otherwise, the NSST index encoding can be performed.
[0346] In addition, when both the X and Y coordinate values for the non-zero transform coefficient are less than the threshold, the encoding of the NSST index can be omitted. Alternatively, the two methods can be applied together.
[0347] Embodiment 5: Applying different NSST index encoding and residual encoding schemes to luminance and chrominance when applying RST
[0348] The schemes described in Embodiments 3 and 4 can be differently applied to luminance and chrominance respectively. That is, the NSST index encoding and residual encoding schemes for luminance and chrominance can be applied differently.
[0349] For example, the luminance may adopt the solution of Embodiment 4 and the chrominance may adopt the solution of Embodiment 3. Alternatively, the luminance may adopt the conditional NSST index coding proposed in Embodiment 3 or 4, and the chrominance may not adopt the conditional NSST index coding. Or vice versa.
[0350] Figure 20 is a flowchart for encoding a video signal based on a simplified quadratic transform as an embodiment of applying the present disclosure.
[0351] The encoder may determine (or select) a forward quadratic transform (S2010) based on at least one of the prediction mode, block shape, and / or block size of the current block. In this case, candidates for the forward quadratic transform may include Figure 6 and / or Figure 12 at least one of the embodiments of.
[0352] The encoder may determine the optimal forward quadratic transform through rate distortion optimization. The optimal forward quadratic transform may correspond to one of a plurality of transform combinations, and these plurality of transform combinations may be defined by a transform index. For example, for RD optimization, the results of performing all of the forward quadratic transform, quantization, residual coding, etc. may be compared for each candidate. In this case, formulas such as cost = rate + λ·distortion or cost = distortion + λ·rate may be used, but the present disclosure is not limited thereto.
[0353] The encoder may signal the quadratic transform index corresponding to the optimal forward quadratic transform (S2020). Here, the quadratic transform index may adopt other embodiments described in the present disclosure.
[0354] For example, the quadratic transform index may adopt Figure 12 the transform set configuration of. Since, according to the intra prediction mode, one transform set consists of two or three transforms, in addition to the case where no quadratic transform is applied, one of up to four transforms can be configured to be selected. When indices 0, 1, 2, and 3 are assigned to the four transforms respectively, the applied transform can be specified by signaling the quadratic transform index for each transform coefficient block. In this case, index 0 can be assigned to the case where no unit matrix, i.e., quadratic transform, is applied.
[0355] As another embodiment, signaling of the quadratic transform index may be performed in any one of the following steps: 1) before residual coding; 2) in the middle of residual coding (after coding the positions of non-zero transform coefficients); or 3) after residual coding. Hereinafter, the embodiments will be described in detail.
[0356] 1) Method of Signaling the Quadratic Transform Index before Residual Coding
[0357] The encoder can determine the forward secondary transform.
[0358] The encoder can signal the secondary transform index corresponding to the forward secondary transform.
[0359] The encoder can encode the position of the last non-zero transform coefficient.
[0360] The encoder can perform residual coding on syntax elements other than the position of the last non-zero transform coefficient.
[0361] 2) Method for Signaling the Secondary Transform Index in the Middle of Residual Coding
[0362] The encoder can determine the forward secondary transform.
[0363] The encoder can encode the position of the last non-zero transform coefficient.
[0364] When the non-zero transform coefficients are not in a specific region, the encoder can encode the secondary transform index corresponding to the forward secondary transform. Here, in the case of applying the simplified secondary transform, when arranging the transform coefficients according to the scan order, the specific region represents the remaining region except for the positions where non-zero transform coefficients may exist. However, the present disclosure is not limited thereto.
[0365] The encoder can perform residual coding on syntax elements other than the position of the last non-zero transform coefficient.
[0366] 3) Method for Signaling the Secondary Transform Index Before Residual Coding
[0367] The encoder can determine the forward secondary transform.
[0368] The encoder can encode the position of the last non-zero transform coefficient.
[0369] When the non-zero transform coefficients are not in a specific region, the encoder can perform residual coding on syntax elements other than the position of the last non-zero transform coefficient. Here, in the case of applying the simplified secondary transform, when arranging the transform coefficients according to the scan order, the specific region represents the remaining region except for the positions where non-zero transform coefficients may exist. However, the present disclosure is not limited thereto.
[0370] The encoder can encode the secondary transform index corresponding to the forward secondary transform.
[0371] In addition, the encoder can perform a forward primary transform (S2030) on the current block (residual block). Here, step S2010 and / or step S2020 can be similarly applied to the forward primary transform.
[0372] The encoder may perform a forward quadratic transform on the current block by using an optimal forward quadratic transform (S2040). For example, the optimal forward quadratic transform may be a simplified quadratic transform. The simplified quadratic transform refers to a transform that inputs N residual data (N×1 residual vector) and outputs L (L < N) transform coefficient data (L×1 transform coefficient vector).
[0373] As an implementation, the simplified quadratic transform may be applied to a specific region of the current block. For example, when the current block is N×N, the specific region may refer to the upper left N / 2×N / 2 region. However, the present disclosure is not limited thereto, and different configurations may be made according to at least one of the prediction mode, block shape, or block size. For example, when the current block is N×N, the specific region may refer to the upper left M×M region (M ≤ N).
[0374] In addition, the encoder performs quantization on the current block to generate a transform coefficient block (S2050).
[0375] The encoder performs entropy coding on the transform coefficient block to generate a bitstream.
[0376] Figure 21 is a flowchart for decoding a video signal based on a simplified quadratic transform as an implementation of applying the present disclosure.
[0377] The decoder may obtain a quadratic transform index from the bitstream (S2110). Here, the quadratic transform index may adopt other implementations described in the present disclosure. For example, the quadratic transform index may include Figure 6 and / or Figure 12 at least one of the implementations of.
[0378] As another implementation, the obtaining of the quadratic transform index may be performed in any one of the following steps: 1) before residual coding; 2) in the middle of residual coding (after decoding the positions of non-zero transform coefficients); or 3) after residual coding.
[0379] The decoder may derive a quadratic transform corresponding to the quadratic transform index (S2120). In this case, candidates for the forward quadratic transform may include Figure 6 and / or Figure 12 at least one of the implementations of.
[0380] However, steps S2110 and S2120 are implementations, and the present disclosure is not limited thereto. For example, the decoder may not obtain the quadratic transform index, but instead derive the quadratic transform based on at least one of the prediction mode, block shape, and / or block size of the current block.
[0381] In addition, the decoder may obtain a transform coefficient block by performing entropy decoding on the bitstream, and may perform inverse quantization on the transform coefficient block (S2130).
[0382] The decoder can perform an inverse quadratic transform (S2140) on the inverse quantized transform coefficient block. For example, the inverse quadratic transform can be a simplified quadratic transform. The simplified quadratic transform refers to a transform that takes N residual data (N×1 residual vector) as input and outputs L (L < N) transform coefficient data (L×1 transform coefficient vector).
[0383] As an implementation, the simplified quadratic transform can be applied to a specific region of the current block. For example, when the current block is N×N, the specific region can refer to the upper left N / 2×N / 2 region. However, the present disclosure is not limited thereto, and different configurations can be made according to at least one of the prediction mode, block shape, or block size. For example, when the current block is N×N, the specific region can refer to the upper left M×M (M≤N) or M×L (M≤N, L≤N) region.
[0384] In addition, the decoder can perform an inverse primary transform on the inverse quadratic transform result (S2150).
[0385] The decoder generates a residual block through step S2150, and the residual block and the prediction block are added together to generate a reconstructed block.
[0386] Embodiment 6: Method for configuring a hybrid second-order transform set
[0387] In an implementation, a method for configuring a set of quadratic transforms by considering various conditions when applying the quadratic transform is proposed.
[0388] In the present disclosure, the quadratic transform indicates the transform performed on all or some of the primary transform coefficients after applying the primary transform on the encoder side as described above, and the quadratic transform can be referred to as an inseparable quadratic transform (NSST), a low-frequency inseparable transform (LFNST), etc. After applying the quadratic transform to all or some of the inverse quantized transform coefficients, the decoder can apply the primary transform to all or some of the inverse quantized transform coefficients.
[0389] In addition, in the present disclosure, the set of hybrid quadratic transforms represents the set of transforms applicable to the quadratic transform, and the present disclosure is not limited to such a name. For example, the set of hybrid quadratic transforms (or transform kernels or transform types) can be referred to as a quadratic transform group, a quadratic transform table, a quadratic transform candidate, a quadratic transform candidate list, a hybrid quadratic transform group, a hybrid quadratic transform table, a hybrid quadratic transform candidate, a hybrid quadratic transform candidate list, etc., and the set of hybrid quadratic transforms can include multiple transform kernels (or transform types).
[0390] A secondary transformation can be applied to a sub-block with a specific size in the upper left side in the current block according to predefined conditions. In conventional video compression techniques, a 4×4 secondary transformation set or an 8×8 secondary transformation set is used according to the size of the selected sub-block. In this case, the 4×4 secondary transformation set only includes transformation kernels (hereinafter referred to as 4×4 transformation kernels) applied to a 4×4 sized area (or block), and the 8×8 secondary transformation set only includes transformation kernels (hereinafter referred to as 8×8 transformation kernels) applied to an 8×8 sized area (or block). In other words, in conventional video compression techniques, the secondary transformation set is composed of transformation kernels with a finite size only according to the size of the area to which the secondary transformation is applied.
[0391] Therefore, the present disclosure proposes a hybrid secondary transformation set including transformation kernels that can be applied to areas with various sizes.
[0392] As an implementation, the size of the transformation kernels included in the hybrid secondary transformation set (i.e., the size of the area to which the corresponding transformation kernel is applied) is not fixed, but can be variably determined (or set). For example, the hybrid secondary transformation set can include 4×4 transformation kernels and 8×8 transformation kernels.
[0393] In addition, in an implementation, the number of transformation kernels included in the hybrid secondary transformation set is not fixed, but can be variably determined (or set). In other words, the hybrid secondary transformation set can include multiple transformation sets, and each transformation set can include a different number of transformation kernels. For example, the first transformation set can include three transformation kernels, and the second transformation set can include four transformation kernels.
[0394] In addition, in an implementation, the order (or priority) among the transformation kernels included in the hybrid secondary transformation set is not fixed, but can be variably determined (or set). In other words, the hybrid secondary transformation set can include multiple transformation sets, and the order among the transformation kernels in each transformation set can be independently defined. In addition, different indexes can be mapped (or assigned) among the transformation kernels in each transformation set. For example, assuming that the first transformation set and the second transformation set each include a first transformation kernel, a second transformation kernel, and a third transformation kernel, the index values 1, 2, and 3 can be mapped to the first transformation kernel, the second transformation kernel, and the third transformation kernel in the first transformation set respectively, and the index values 3, 2, and 1 can be mapped to the first transformation kernel, the second transformation kernel, and the third transformation kernel in the second transformation set respectively.
[0395] Hereinafter, a method for determining the priority (or order) among the transformation kernels in the hybrid secondary transformation set will be described in detail.
[0396] When applying a secondary transform, the decoder may determine (or select) a set of secondary transforms to be applied to the current processing block according to predefined conditions. The set of secondary transforms may include a hybrid secondary transform set according to embodiments of the present disclosure. Additionally, the decoder may use a secondary transform index signaled from the encoder to determine a transform kernel of the secondary transform applied to the current processing block within the determined set of secondary transforms. The secondary transform index may indicate the transform kernel of the secondary transform applied to the current processing block within the determined set of secondary transforms. In one embodiment, the secondary transform index may be referred to as an NSST index, an LFNST index, etc.
[0397] As described above, since the secondary transform index is signaled from the encoder to the decoder, it is reasonable in terms of compression efficiency to assign lower indices to relatively more frequently occurring transform kernels for encoding / decoding with fewer bits. Accordingly, various embodiments for configuring a set of secondary transforms considering priorities will be described below.
[0398] In an embodiment, priorities among transform kernels in the set of secondary transforms may be determined differently according to the size of the region (or sub-block) to which the secondary transform is applied. For example, when the size of the current processing block is equal to or greater than a predetermined size, an 8×8 transform kernel may be used more frequently. As a result, the encoder / decoder may assign relatively fewer transform indices to the 8×8 transform kernel than to the 4×4 transform kernel. As an example, a hybrid secondary transform set may be configured as in Table 3 below.
[0399] [Table 3]
[0400] NSST index 4×4 NSST set 8×8 NSST set Hybrid NSST set 1 4×4 first kernel 8×8 first kernel 8×8 first kernel 2 4×4 second kernel 8×8 second kernel 8×8 second kernel 3 4×4 third kernel 8×8 third kernel 4×4 first kernel ... ... ... ...
[0401] In Table 3, embodiments are described by assuming the use of NSST as the secondary transform, but the present disclosure is not limited to such a name. As described above, the secondary transform according to embodiments of the present disclosure may be referred to as non-separable secondary transform (NSST), low-frequency non-separable transform (LFNST), etc. Referring to Table 3, the set of secondary transforms of conventional video compression techniques consists only of transform kernels of the size of the region to which the secondary transform is applied (i.e., the 4×4 NSST set and the 8×8 NSST set in Table 3). The hybrid secondary transform set according to embodiments of the present disclosure (i.e., the hybrid NSST set) may include an 8×8 transform kernel and a 4×4 transform kernel.
[0402] In Table 3, it is assumed that the size of the current processing block is equal to or greater than a predetermined size. That is, when the minimum value of the width or height of the current block is equal to or greater than a predefined value (e.g., 8), the secondary transform set illustrated in Table 3 may be used. The hybrid secondary transform set may include an 8×8 transform kernel and a 4×4 transform kernel, and it is likely that the 8×8 transform kernel is used. As a result, a low index value may be assigned to the 8×8 transform kernel.
[0403] In addition, in an embodiment, the priority between transform kernels in a secondary transform set can be determined based on the order of the secondary transform kernels (first, second, and third). For example, the first 4×4 secondary transform kernel can have a higher priority than the second 4×4 secondary transform kernel, and a lower index value can be assigned to the first 4×4 secondary transform kernel. As an example, the hybrid secondary transform set can be configured as shown in Table 4 below.
[0404] [Table 4]
[0405]
[0406] In Table 4, the embodiment is described by assuming that NSST is used as the secondary transform, but the present disclosure is not limited to such a name. As described above, the secondary transform according to the embodiment of the present disclosure can be referred to as NSST, LFNST, etc.
[0407] Referring to Table 4, the hybrid secondary transform set (i.e., hybrid NSST set types 1, 2, and 3) according to the embodiment of the present disclosure can include 8×8 transform kernels and / or 4×4 transform kernels. In other words, like hybrid NSST set types 2 and 3, the hybrid secondary transform set can include 8×8 transform kernels and / or 4×4 transform kernels, and the priority between 8×8 transform kernels and the priority between 4×4 transform kernels can be configured based on the order of the corresponding secondary transform kernels.
[0408] Embodiment 7: Method for configuring a hybrid second-order transform set
[0409] In an embodiment of the present disclosure, a method for determining a secondary transform set by considering various conditions is proposed. Specifically, a method for determining a secondary transform set based on the size of a block and / or an intra prediction mode is proposed.
[0410] The encoder / decoder can configure a transform set suitable for the secondary transform of the current block based on the intra prediction mode. As an embodiment, the proposed method can be applied together with the above-described Embodiment 6. For example, the encoder / decoder can configure a secondary transform set based on the intra prediction mode and perform a secondary transform using transform kernels of various sizes included in each secondary transform set.
[0411] In an embodiment, the secondary transform set can be determined based on Table 5 below according to the intra prediction mode.
[0412] [Table 5]
[0413] Intra mode 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 Hybrid type 1 1 1 1 1 0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 0 0 0 0 0 0 0 0 0 0 0 1 1 Intra mode 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 Hybrid type 1 1 1 0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0
[0414] Referring to Table 5, the encoder / decoder can determine whether to apply (or configure) the hybrid quadratic transform set described in Embodiment 6 above based on the intra prediction mode. When the hybrid quadratic transform set is not applied, the encoder / decoder can apply (or configure) the quadratic transform set described above Figures 12 to 14 in the description.
[0415] Specifically, in the case of the intra prediction mode where the hybrid type value is defined as 1, the hybrid quadratic transform set can be configured according to the method described in Embodiment 6 above. Additionally, in the case of the intra prediction mode where the hybrid type value is defined as 0, the quadratic transform set can be configured according to a conventional method (i.e., the method described above Figures 12 to 14 in the description).
[0416] Table 5 shows a case of using the method for configuring two types of transform sets, but the present disclosure is not limited thereto. That is, two or more types of hybrid types indicating methods for configuring transform sets including the hybrid quadratic transform set can be configured (or defined). In this case, the hybrid type information can be configured as various, i.e., N types (N>2) of values.
[0417] Furthermore, in the embodiment, the encoder / decoder can determine whether the transform set suitable for the current block is configured as a normal type or a hybrid type by considering the size of the current coding block (or transform block) and the intra prediction mode. Here, the normal type indicates the quadratic transform set configured according to a conventional method (i.e., the method described above Figures 12 to 14 in the description). For example, when the hybrid type (or mode type) value is 0, the encoder / decoder can configure the quadratic transform set by applying the method described Figures 12 to 14 in the description, and when the hybrid type value is 1, the encoder / decoder can configure a hybrid quadratic transform set including transform kernels of various sizes according to the size of the current block.
[0418] Figure 22 is a diagram illustrating a method for determining a transform type to be applied to a quadratic transform according to an embodiment of applying the present disclosure.
[0419] Referring to Figure 22 , the decoder is mainly described for convenience of description, but the present disclosure is not limited thereto, and a method for performing a quadratic transform by determining a transform type can be applied to the encoder substantially equivalently.
[0420] The decoder performs inverse quantization (S2201) on the quantized transform block (or transform coefficients).
[0421] The decoder determines (or selects) a secondary transform set (or set of transform types) for the secondary transform of the current block based on the size of the current block and / or the intra prediction mode (S2202). In this case, various secondary transform sets can be predefined, and the method of configuring the secondary transform set described in this embodiment and / or Embodiment 6 can be applied.
[0422] As an embodiment, the decoder can determine whether to use a hybrid type (or hybrid secondary transform set) to configure the secondary transform set based on the intra prediction mode of the current block. When using the hybrid type, a hybrid secondary transform set according to the method described in Embodiment 6 above can be used.
[0423] The decoder performs a secondary transform on the inverse quantized transform block (or transform coefficients) by using the determined transform kernel (S2203). In this case, the decoder can parse a secondary transform index indicating the transform kernel for the secondary transform of the current block within the secondary transform set determined in step S2202. In this case, the step of parsing the secondary transform index can be included in step S2203.
[0424] Embodiment 8: Method for encoding second-order transform indices
[0425] In an embodiment, a method for efficiently encoding / decoding a secondary transform index signaled from an encoder is proposed in the case of configuring a secondary transform set based on the size and / or intra prediction mode of an encoded block (or transform block).
[0426] As described above, when applying a secondary transform, the encoder / decoder can determine (or select) a secondary transform set to be applied to the current encoded block according to preconfigured conditions. In addition, the decoder can use the secondary transform index signaled from the encoder to derive the transform kernel applied to the current encoded block within the determined secondary transform set. Here, the secondary transform index represents a syntax element indicating the transform kernel of the secondary transform applied to the current block within the secondary transform set. In the present disclosure, the secondary transform index can be referred to as an NSST index, an LFNST index, etc.
[0427] As described in Embodiment 6 above, the encoder / decoder can configure a secondary transform set by using transform kernels of various sizes, and the number of transform kernels included in the secondary transform set is not fixed but can be variably determined.
[0428] Therefore, since the number of available transform kernels can be different for each set of quadratic transforms, the encoder / decoder can binarize the quadratic transform index by using the truncated unary method for efficient binarization. As an implementation, the encoder / decoder can perform truncated unary binarization according to the maximum available value of the quadratic transform index for each set of quadratic transforms by using Table 6 below.
[0429] [Table 6]
[0430]
[0431]
[0432] In Table 6, the implementation is described by assuming the use of NSST as the quadratic transform, but the present disclosure is not limited to such a name. As described above, the quadratic transform according to the implementation of the present disclosure can be referred to as an inseparable quadratic transform (NSST), a low-frequency inseparable transform (LFNST), etc.
[0433] Referring to Table 6, the NSST index can be binarized by the truncated unary binarization method. In this case, the binarization length can be determined according to the maximum index value within the set of quadratic transforms.
[0434] In an implementation, the quadratic transform index in Table 6 is encoded / decoded based on context, and in this case, the following variables can be considered for context modeling.
[0435] - The size of the coding block (or transform block)
[0436] - The intra prediction mode
[0437] - The hybrid type value
[0438] - The quadratic transform index value of the set of quadratic transforms
[0439] Embodiment 9: Reduced transform
[0440] In the implementation of the present disclosure, various implementations of the simplified transform proposed in Figures 15 to 22 are described to improve the complexity problem of the transform. As described above, the simplified transform proposed in the present disclosure can be applied regardless of whether it is a primary transform (e.g., DCT, DST) or a quadratic transform (e.g., NSST, low-frequency inseparable transform (LFNST)).
[0441] Figure 23 is a diagram illustrating a simplified transform structure based on a simplification factor as an application of the implementation of the present disclosure.
[0442] Referring to Figure 23, the decoder is mainly described for convenience of description, but the simplified transformation proposed in the embodiments can be equivalently applied to the encoder.
[0443] The decoder can apply an inverse simplified transformation to the inverse-quantized transform coefficients. In this case, the decoder can use a predetermined (or predefined) simplification factor (e.g., R or R / N) and / or a transform kernel to perform the simplified transformation.
[0444] In one embodiment, a transform kernel can be selected based on available information such as the size (e.g., width / height) of the current block (encoded block or transform block), the intra / inter-frame prediction mode, CIdx, etc. When the current encoded block is a luminance block, the value of CIdx can be 0. Otherwise (i.e., if the current encoded block is a Cb or Cr block), CIdx can have a non-zero value such as 1.
[0445] Figure 24 is a diagram illustrating a method of performing decoding by adaptively applying a simplified transformation as an embodiment to which the present disclosure can be applied.
[0446] Refer to Figure 24 , the decoder is mainly described for convenience of description, but the method of performing a transformation using a simplified transformation proposed in the embodiments can be equivalently applied to the encoder.
[0447] The decoder performs inverse quantization on the current block (S2401).
[0448] The decoder checks whether a transformation is applied to (or used for) the current block (S2402). If no transformation is applied to the current block, the decoder terminates the transformation process.
[0449] When a transformation is applied to the current block, the decoder parses a transform index indicating the transform kernel applied to the current block from the video signal (S2403).
[0450] The decoder checks whether the simplified inverse transform condition is satisfied (S2404). If the simplified inverse transform condition is not satisfied, the decoder performs a normal inverse transform on the current block (S2405). If the simplified inverse transform condition is satisfied, the decoder performs a simplified inverse transform on the current block (S2407). In this case, the decoder can select the transform kernel applied to the current block based on the transform index parsed in step S2403 (S2406). As an embodiment, a transform kernel can be selected based on available information such as the size (e.g., width / height) of the current block (encoded block or transform block), the intra / inter-frame prediction mode, CIdx, etc. In addition, when a simplified inverse transform is applied to the current block, step S2406 can include selecting a simplification factor.
[0451] In an embodiment, the simplified inverse transform condition may be applied to the above condition 6) (e.g., Tables 3 and 4). In other words, it may be determined whether to apply the simplified inverse transform based on the size of the current block (encoding block or transform block) and the transform type (or transform kernel).
[0452] As an embodiment, when the following specific conditions are met, the simplified transform may be used. In other words, the simplified transform may be applied to blocks having a predetermined size or larger size (or greater than a predetermined size) that satisfy the following specific conditions.
[0453] - Width > TH && Height > TH (where TH is a predefined value representing a specific threshold (e.g., 4))
[0454] Or
[0455] - Width × Height > K && MIN(Width, Height) > TH (where K or TH is a predefined value and represents a specific threshold)
[0456] As another embodiment, when the following specific conditions are met, the simplified transform may be used. In other words, the simplified transform may be applied to blocks having a predetermined size or smaller size (or less than a predetermined size) that satisfy the following specific conditions.
[0457] - Width <= TH && Height <= TH (where TH is a predefined value representing a specific threshold (e.g., 8))
[0458] Or
[0459] - Width × Height <= K && MIN(Width, Height) <= TH (where K or TH is a predefined value and represents a specific threshold)
[0460] As another example, the simplified transform may be applied only to a predetermined block group.
[0461] - Width == TH && Height == TH
[0462] Or
[0463] - Width == Height
[0464] As an embodiment, if the usage conditions of the simplified transform are not met, the normal transform may be applied. Specifically, the normal transform may be predefined and available for the encoder / decoder. Examples of the normal transform are shown below.
[0465] - DCT2, DCT4, DCT5, DCT7, DCT8
[0466] Or
[0467] - DST1, DST4, DST7
[0468] or
[0469] - Inseparable transform
[0470] or
[0471] - NSST(HyGT)
[0472] or
[0473] - LFNST(Low - frequency inseparable transform)
[0474] The above conditions can be interpreted based on logical operators, as shown in Table 7 below.
[0475] [Table 7]
[0476]
[0477] In addition, as exemplified in Figure 24 , the simplified transform condition can depend on the transform index Transform_idx indicating the transform applied to the current block. As an example, Transform_idx can be sent from the encoder to the decoder twice. One can be the horizontal transform index Transform_idx_h, and the other can be the vertical transform index Transform_idx_v.
[0478] Figure 25 is a diagram illustrating a method of performing decoding by adaptively applying a simplified transform, which can be an embodiment of the present disclosure.
[0479] Referring to Figure 25 , the decoder is mainly described for convenience of description, but the method of performing a transform using a simplified transform proposed in the embodiment can be equivalently applied to the encoder.
[0480] In an embodiment of the present disclosure, the above - mentioned simplified transform can be used for the secondary transform. In this regard, the description repeated with the method described in Figure 24 will be omitted.
[0481] 1) The decoder performs inverse quantization on the current block and then checks whether NSST is activated in the current block. The decoder can use predefined conditions to determine whether it is necessary to parse the NSST index.
[0482] 2) If NSST is activated, the decoder parses the NSST index and checks whether to apply the simplified secondary inverse transform.
[0483] 3) The decoder checks whether the simplified inverse transform condition is satisfied.
[0484] 4) If the simplified inverse transform condition is not satisfied, the decoder performs a normal inverse transform on the current block.
[0485] 5) If the simplified quadratic inverse transform condition is satisfied, the decoder performs a simplified quadratic inverse transform on the current block.
[0486] 6) In this case, the decoder can select a transform kernel to be applied to the current block based on the NSST index. As an implementation, the transform kernel can be selected based on available information such as the size (e.g., width / height) of the current block (encoded block or transform block), intra / inter prediction mode, CIdx, etc. Additionally, when applying a simplified quadratic inverse transform to the current block, the decoder can select a simplification factor.
[0487] In an implementation, it can be determined whether to apply a simplified inverse transform based on the size of the current block (encoded block or transform block) and the transform type (or transform kernel).
[0488] Embodiment 10: Reduced transform as a second-order transform with different block sizes
[0489] In an implementation of the present disclosure, a simplified transform considering various block sizes for quadratic transform / quadratic inverse transform is proposed. As an example, simplified transforms for different block sizes of 4×4, 8×8, and 16×16 for quadratic transform / quadratic inverse transform can be defined.
[0490] Figure 26 and Figure 27 are diagrams illustrating examples of forward simplified quadratic transform and inverse simplified quadratic transform and the pseudo-code for deriving them.
[0491] Referring to Figure 27 and 28 , the simplified quadratic transform and the simplified quadratic inverse transform are shown when the block to which the quadratic transform is applied is an 8×8 block and the simplification coefficient R is 16. The simplified quadratic transform and the simplified quadratic inverse transform illustrated in Figure 27 can be derived by using the pseudo-code illustrated in Figure 26 .
[0492] Embodiment 11: Reduced transform as a second-order transform with a non-rectangular shape
[0493] As described above, due to the complexity problem of applying the quadratic transform of the inseparable transform, in the image compression technology of the related art, the quadratic transform is applied to the upper left 4×4 or 8×8 region of the encoded block (or transform block).
[0494] The implementation of the present disclosure proposes a method of applying a simplified quadratic transform to various non-square graphs in addition to the 4×4 or 8×8 square region.
[0495] Figure 28 is a diagram illustrating a method of applying a simplified quadratic transform to a non-square region as an implementation of the present disclosure.
[0496] Refer to Figure 28 , in an embodiment, a simplified quadratic transform can be applied to only a part of a block, as exemplified in Figure 29 .
[0497] In Figure 28 , each square represents a 4×4 area. Thus, the encoder / decoder can apply the simplified quadratic transform to a 10×4 pixel area, i.e., a 160 pixel area. In this case, the simplification factor R = 16 and the entire RST matrix corresponds to a 16×160 matrix, thereby reducing the computational complexity of applying the quadratic transform.
[0498] Embodiment 12: Reduction factor
[0499] Figure 29 is a diagram illustrating a simplified transform controlled by a simplification factor as an embodiment of applying the present disclosure.
[0500] Refer to Figure 29 , as described above, the simplified transform according to an embodiment of the present disclosure can be controlled by a simplification factor, as exemplified in Figure 29 .
[0501] Specifically, modifying the simplification factor can modify the memory complexity and the number of multiplication operations. As previously mentioned as the simplification factor R / N in Figure 15 and Equation 6, the memory and multiplication can be reduced by modifying the simplification factor. For example, for an 8×8 NSST with R = 16, the memory and multiplication can be reduced by 1 / 4.
[0502] Embodiment 13: High-level syntax
[0503] Embodiments of the present disclosure propose a high-level syntax structure for controlling a simplified transform at a high level.
[0504] In an embodiment, as shown in the example of Table 8 below, information on whether to accept the simplified transform and on the size and / or simplification factor can be sent through a sequence parameter set (SPS). However, the present disclosure is not limited thereto, and the syntax can be signaled through a picture parameter set (PPS), a slice header, etc.
[0505] [Table 8]
[0506]
[0507]
[0508] Referring to Table 8, if the Reduced_transform_enabled_flag is 1, the reduced transform may be available and applied. In the case where the Reduced_transform_enabled_flag is 0, this case may indicate that the reduced transform is not available. If the Reduced_transform_enabled_flag does not exist, the value may be estimated to be equal to zero.
[0509] Reduced_transform_factor represents a syntax element that specifies the number of reduced dimensions used for the reduced transform.
[0510] min_reduced_transform_size represents a syntax element that specifies the minimum transform size to which the reduced transform will be applied. If min_reduced_transform_size does not exist, the value may be estimated to be equal to zero.
[0511] max_reduced_transform_size represents a syntax element that specifies the maximum transform size to which the reduced transform will be applied. If max_reduced_transform_size does not exist, the value may be estimated to be equal to zero.
[0512] Reduced_transform_factor represents a syntax element that specifies the number of reduced dimensions used for the reduced transform. If Reduced_transform_factor does not exist, the value may be estimated to be equal to zero.
[0513] Embodiment 14: Second-order transform kernel
[0514] Embodiments of the present disclosure propose various secondary transform kernels.
[0515] In an embodiment, the 4×4 NSST kernel for the DC mode may be defined as shown in Table 9 below.
[0516] [Table 9]
[0517]
[0518]
[0519] Additionally, in an embodiment, the 4×4 NSST kernel for the planar mode may be defined as shown in Table 10 below.
[0520] [Table 10]
[0521]
[0522]
[0523] In addition, in the embodiments, an 8×4 NSST kernel for the DC mode can be defined as shown in Table 11 below.
[0524] [Table 11]
[0525]
[0526]
[0527]
[0528]
[0529]
[0530]
[0531] In addition, in the embodiments, an 8×8 NSST kernel for the planar mode can be defined as shown in Table 12 below.
[0532] [Table 12]
[0533]
[0534]
[0535]
[0536]
[0537]
[0538]
[0539] The transform kernels in Tables 9 to 12 above can be defined as smaller transform kernels for simplified transforms.
[0540] For example, for an 8×8 NSST with the DC mode and R = 8, the memory and multiplications can be reduced by 1 / 2. Therefore, by keeping only the coefficients in the upper half of the transform kernel in Table 7 above (8×16 matrix), a simplified transform kernel can be defined in a smaller size as shown in Table 13.
[0541] [Table 13]
[0542]
[0543] The transform kernels in Tables 9 to 12 above can be defined as smaller transform kernels for simplified transforms.
[0544] For example, for the 4×4 NSST in DC mode and R = 8, the memory and multiplication can be reduced by 1 / 2. Therefore, by keeping only the coefficients in the upper half of the transform kernel in Table 7 above (8×16 matrix), a simplified transform kernel can be defined with a smaller size as shown in Table 13.
[0545] [Table 14]
[0546]
[0547]
[0548]
[0549]
[0550]
[0551]
[0552] In the above example, each transform coefficient is represented by 9 bits (i.e., 1 bit: negative sign; 8 bits: absolute value from 0 to 255). In an embodiment of the present disclosure, various precisions can be used to represent the transform coefficients. For example, 8 bits can be used instead of 9 bits to represent each coefficient. In this case, the sign bit does not change, but the range of the absolute value can change.
[0553] For convenience of description, the above-described embodiments of the present disclosure have been described separately, but the present disclosure is not limited thereto. That is, the embodiments described in the above Embodiments 1 to 14 can be independently executed, and one or more various embodiments can be combined and executed.
[0554] Figure 30 is a diagram illustrating an inverse transform method according to an embodiment of applying the present disclosure.
[0555] Referring to Figure 30 , the decoder is mainly described for convenience of description, but the method of performing transform / inverse transform based on the quadratic transform set according to an embodiment of the present disclosure can be applied to the encoder even substantially equivalently.
[0556] The decoder generates a dequantized transform block by performing dequantization on the current block (S3001).
[0557] The decoder obtains the intra prediction mode of the current block (S3002).
[0558] The decoder determines the quadratic transform set applied to the current block among a plurality of quadratic transform sets based on the intra prediction mode (S3003).
[0559] As described above, a plurality of secondary transform sets may include at least one hybrid secondary transform set.
[0560] In addition, as described above, a hybrid secondary transform set may include at least one 8×8 transform kernel applied to a region of size 8×8 and at least one 4×4 transform kernel applied to a region of size 4×4.
[0561] Furthermore, as described above, when a plurality of secondary transform sets include a plurality of hybrid secondary transform sets, the plurality of hybrid secondary transform sets may respectively include different numbers of transform kernels.
[0562] In addition, as described above, step S3003 may include: determining whether to use a hybrid secondary transform set based on an intra prediction mode; and when it is determined to use a hybrid secondary transform set, determining, based on the size of the current block, a secondary transform set applied to the current block among the plurality of secondary transform sets including at least one hybrid secondary transform set; and when it is determined not to use a hybrid secondary transform set, determining, based on the size of the current block, a secondary transform set applied to the current block among the remaining secondary transform sets other than the at least one hybrid secondary transform set.
[0563] The decoder derives the transform kernel applied to the current block in the determined secondary transform set (S3004).
[0564] As described above, step S3004 may further include obtaining a secondary transform index indicating the transform kernel applied to the current block in the determined secondary transform set. As an implementation, the secondary transform index may be binarized by a truncated unary scheme based on the maximum number of available transform kernels in the determined secondary transform set.
[0565] The decoder performs a secondary transform on a specific upper left region of the inverse quantized transform block by using the derived transform kernel (S3005).
[0566] Figure 31 FIG. is an illustration showing an inverse quantization unit and an inverse transform unit according to an embodiment of the present disclosure.
[0567] In Figure 31 , for convenience of description, the inverse transform unit 3100 is illustrated as a block, but the inter prediction unit may be implemented as a component included in an encoder and / or a decoder.
[0568] Referring to Figure 31 , the inverse quantization unit 3101 and the inverse transform unit 3100 implement the above Figures 4 to 30The functions, processes, and / or methods proposed in [reference]. Specifically, the inverse transform unit 3100 may be configured to include an intra prediction mode acquisition unit 3102, a secondary transform set determination unit 3103, a transform kernel derivation unit 3104, and a secondary inverse transform unit 3105.
[0569] The inverse quantization unit 3101 generates a dequantized transform block by performing inverse quantization on the current block.
[0570] The intra prediction mode acquisition unit 3102 acquires the intra prediction mode of the current block.
[0571] The secondary transform set determination unit determines the secondary transform set applied to the current block among multiple secondary transform sets based on the intra prediction mode.
[0572] As described above, the multiple secondary transform sets may include at least one hybrid secondary transform set.
[0573] In addition, as described above, the hybrid secondary transform set may include at least one 8×8 transform kernel applied to an 8×8-sized region and at least one 4×4 transform kernel applied to a 4×4-sized region.
[0574] Furthermore, as described above, when the multiple secondary transform sets include multiple hybrid secondary transform sets, the multiple hybrid secondary transform sets may respectively include different numbers of transform kernels.
[0575] In addition, as described above, the secondary transform set determination unit 3103 may determine whether to use the hybrid secondary transform set based on the intra prediction mode; and when it is determined to use the hybrid secondary transform set, determine the secondary transform set applied to the current block among the multiple secondary transform sets including at least one hybrid secondary transform set based on the size of the current block; and when it is determined not to use the hybrid secondary transform set, determine the secondary transform set applied to the current block among the remaining secondary transform sets other than the at least one hybrid secondary transform set based on the size of the current block.
[0576] The transform kernel derivation unit 3104 derives the transform kernel applied to the current block in the determined secondary transform set.
[0577] As described above, the transform kernel derivation unit 3104 may acquire a secondary transform index indicating the transform kernel applied to the current block in the determined secondary transform set. As an implementation, the secondary transform index may be binarized by a truncated unary scheme based on the maximum number of available transform kernels in the determined secondary transform set.
[0578] The secondary inverse transform unit 3105 performs a secondary inverse transform on a specific upper-left region of the dequantized transform block by using the derived transform kernel.
[0579] The multiple secondary transformation sets include at least one hybrid secondary transformation set.
[0580] Figure 32 An example of applying the video coding system of the present disclosure is illustrated.
[0581] The video coding system may include a source device and a receiving device. The source device may transmit the encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.
[0582] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0583] The video source may obtain video / images through processes such as capture, synthesis, or generation of video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like, and in this case, the video / image capture process may be replaced by a process of generating relevant data.
[0584] The encoding device may encode the input video / image. The encoding device may perform a series of processes including prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0585] The transmitter may transmit the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include an element for generating a media file in a predetermined file format, and may include an element for transmitting via a broadcast / communication network. The receiver may extract the bitstream and send the extracted bitstream to the decoding device.
[0586] The decoding device performs a series of processes including inverse quantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding device to decode the video / image.
[0587] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0588] Figure 33 It is a structural diagram of a content streaming system to which embodiments of the present disclosure are applied.
[0589] Refer to Figure 33 , the content streaming system to which the present disclosure is applied can mainly include an encoding server, a streaming server, a network server, a media memory, a user device, and a multimedia input device.
[0590] The encoding server compresses the content input from a multimedia input device including a smart phone, a camera, a video camera, etc. into digital data for generating a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device including a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted.
[0591] A bitstream can be generated by applying the encoding method or bitstream generation method of the present disclosure, and the streaming server can temporarily store the bitstream in the process of sending or receiving the bitstream.
[0592] The streaming server sends multimedia data to the user device based on a request from the user through the network server, and the network server serves as an intermediary for informing the user of what services are available. When the user requests a desired service from the network server, the network server transmits the requested service to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system can include a separate control server, and in this case, the control server is used to control commands / responses between various devices in the content streaming system.
[0593] The streaming server can receive content from the media memory and / or the encoding server. For example, when the streaming server receives content from the encoding server, the streaming server can receive the content in real time. In this case, the streaming server can store the bitstream for a predetermined time in order to provide a smooth streaming service.
[0594] Examples of user devices can include cellular phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, touchscreen PCs, tablet PCs, ultrabooks, wearable devices such as smart watches, smart glasses, or head-mounted displays (HMDs), etc.
[0595] Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be distributedly processed.
[0596] As described above, the embodiments described in the present disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0597] In addition, the decoder and encoder applying the present disclosure can be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video-on-demand (VoD) service providing device, an (over-the-top) OTT video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, a vehicle terminal (e.g., a vehicle terminal, an aircraft terminal, a ship terminal, etc.), a medical video device, etc., and can be used to process video signals or data signals. For example, an over-the-top (OTT) video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0598] In addition, the processing method applying the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a flexible disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted through a wired / wireless communication network.
[0599] In addition, the embodiments of the present disclosure can be implemented as a computer program product by program code, and the program code can be executed on a computer by the embodiments of the present disclosure. The program code can be stored on a computer-readable carrier.
[0600] In the above-described embodiments, the components and features of the present disclosure are combined in a predetermined form. Unless otherwise explicitly stated, each component or feature should be regarded as optional. Each component or feature can be implemented without being associated with other components or features. Additionally, embodiments of the present disclosure can be configured by associating some components and / or features. The order of operations described in the embodiments of the present disclosure can be changed. Some components or features of any embodiment can be included in another embodiment or replaced with components and features corresponding to another embodiment. Obviously, after filing, claims not explicitly recited in the combination claims can be modified to form an embodiment or included in new claims.
[0601] Embodiments of the present disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of implementation by hardware, according to the hardware implementation manner, the exemplary embodiments described herein can be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and the like.
[0602] In the case of implementation by firmware or software, embodiments of the present disclosure can be implemented in the form of modules, processes, functions, etc. to perform the above-described functions or operations. The software code can be stored in a memory and run by a processor. The memory can be internal or external to the processor, and data can be sent to / received from the processor by various known means.
[0603] It is obvious to those skilled in the art that the present disclosure can be implemented in other specific forms without departing from the essential characteristics of the present disclosure. Therefore, the above-mentioned detailed description should not be construed as restrictive in any way and should be considered exemplary. The scope of the present disclosure should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present disclosure are included within the scope of the present disclosure.
[0604] Industrial Applicability
[0605] Above, for illustrative purposes, preferred embodiments of the present disclosure have been disclosed, and hereinafter, those skilled in the art will make various modifications, changes, substitutions, or additions to other embodiments within the technical spirit and scope of the present disclosure disclosed in the appended claims.
Claims
1. An apparatus for decoding a video signal, the apparatus comprising: a memory for storing the video signal; and a processor coupled to the memory, wherein the processor is configured to: generate an inverse-quantized transform block for a current block; obtain an intra prediction mode of the current block; determine a secondary transform set among a plurality of secondary transform sets to be applied to the current block based on the intra prediction mode; derive a transform kernel in the determined secondary transform set to be applied to the current block; and perform an inverse secondary transform on coefficients of the inverse-quantized transform block based on the derived transform kernel, wherein, to derive the transform kernel, the processor is configured to: obtain a secondary transform index including information about the transform kernel in the determined secondary transform set to be applied to the current block, and wherein the number of coefficients on which the inverse secondary transform is performed is less than the number of coefficients obtained by performing the inverse secondary transform.
2. The device according to claim 1, wherein, Binarize the secondary transform index in a truncated unary scheme based on the number of available transform kernels in the determined secondary transform set.
3. The device according to claim 1, wherein The plurality of secondary transform sets includes at least one hybrid secondary transform set including at least one 8×8 transform kernel applied to an 8×8 size region and at least one 4×4 transform kernel applied to a 4×4 size region.
4. The device according to claim 3, wherein, When the plurality of secondary transform sets includes a plurality of hybrid secondary transform sets, the plurality of hybrid secondary transform sets respectively include different numbers of transform kernels.
5. The apparatus according to claim 4, wherein the secondary transform index is binarized in a truncated unary scheme based on the maximum number of available transform kernels in the determined secondary transform set.
6. The device according to claim 3, wherein To determine the secondary transform set, the processor is configured to: determine whether to use a hybrid secondary transform set based on the intra prediction mode, and when it is determined to use a hybrid secondary transform set, determine the secondary transform set among the plurality of secondary transform sets including the at least one hybrid secondary transform set to be applied to the current block based on the size of the current block, and when it is determined not to use a hybrid secondary transform set, determine the secondary transform set among the remaining secondary transform sets other than the at least one hybrid secondary transform set to be applied to the current block based on the size of the current block.
7. An apparatus for encoding a video signal, the apparatus comprising: a memory for storing the video signal; and a processor coupled to the memory, wherein the processor is configured to: obtain residual data of a current block; perform a primary transform on the residual data to obtain a first block; perform a secondary transform on a first number of coefficients of the first block to obtain a second number of transform coefficients; and perform quantization on the transform coefficients, wherein, to perform the secondary transform, the processor is configured to: determine a secondary transform set among a plurality of secondary transform sets to be applied to the first block based on the intra prediction mode applied to the current block, derive a transform kernel in the secondary transform set to be applied to the first block; Generate a secondary transform index including information about the transform kernel; and Apply the transform kernel to the first block, and wherein the first number is greater than the second number.
8. The apparatus according to claim 7, wherein the processor is further configured to: Binarize the secondary transform index in a truncated unary scheme based on the number of available transform kernels in the secondary transform set.
9. The device according to claim 7, wherein The plurality of secondary transform sets include at least one hybrid secondary transform set including at least one 8×8 transform kernel applied to an 8×8 sized region and at least one 4×4 transform kernel applied to a 4×4 sized region.
10. An apparatus for transmitting data for an image, the apparatus comprising: A processor configured to obtain a bitstream for the image; And A transmitter configured to transmit the data including the bitstream, wherein the processor is configured to: Obtain the bitstream for the image; and Transmit the data including the bitstream, wherein, to obtain the bitstream, the processor is configured to: Obtain residual data of a current block; Perform a primary transform on the residual data to obtain a first block; Perform a secondary transform on a first number of coefficients of the first block to obtain a second number of transform coefficients; and Perform quantization on the transform coefficients, wherein performing the secondary transform includes: Determine a secondary transform set among the plurality of secondary transform sets to be applied to the first block based on an intra prediction mode applied to the current block, Derive a transform kernel in the secondary transform set to be applied to the first block; Generate a secondary transform index including information about the transform kernel; and Apply the transform kernel to the first block, and wherein the first number is greater than the second number.
Citation Information
Patent Citations
Binarizing secondary transform index
US20170324643A1
Image encoding method and apparatus, and image decoding method and apparatus
WO2017138791A1