Method and apparatus for processing a video signal using a reduced quadratic transform
The application of a Reduced Secondary Transform (RST) to 4x4 blocks with specific coding methods addresses the inefficiencies in processing high-resolution video content, reducing complexity and enhancing coding efficiency.
Patent Information
- Application Number
- JP2025016756
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-04-01
- Filing Date
- 2025-02-04
- Publication Date
- 2025-11-12
- Estimated Expiration
- 2039-04-01
AI Technical Summary
Next-generation video content with high spatial resolution, frame rate, and dimensionality requires more efficient coding tools to manage increased memory storage, memory access rate, and processing power, particularly in transform designs.
A Reduced Secondary Transform (RST) is applied to 4x4 blocks with specific transform index coding and scanning order, conditionally coding transform indices, and omitting residual coding for impermissible positions, differing between luma and chroma blocks.
This approach significantly reduces calculation complexity and improves coding efficiency for encoding and decoding still or moving images, especially when compared to non-separable secondary transforms.
Smart Images

Figure 0007769157000013 
Figure 0007769157000014 
Figure 0007769157000015
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and apparatus for processing a video signal, and more particularly to a design of a Reduced Secondary Transform (RST) that can be applied to a 4x4 block, an arrangement of transform coefficients generated after applying the 4x4 RST, a scanning order, and a transform index coding method for specifying the 4x4 RST to be applied. [Background technology]
[0002] Next-generation video content is characterized by high spatial resolution, high frame rate, and high dimensionality of scene representation. Processing such content will bring about significant increases in memory storage, memory access rate, and processing power.
[0003] Therefore, it is necessary to design coding tools to process next-generation video content more efficiently. In particular, when applying transforms, it is necessary to design transforms that are much more efficient in terms of coding efficiency and complexity. Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention proposes an encoder / decoder structure to reflect the new transform design.
[0005] The present invention proposes a design of an RST that can be applied to a 4x4 block, a transform index coding method and structure for specifying the arrangement and scanning order of transform coefficients generated after applying the 4x4 RST, and the 4x4 RST to be applied. [Means for solving the problem]
[0006] The present invention provides a method to reduce complexity and improve coding efficiency through a novel transform design.
[0007] The present invention provides a design method for an RST that can be applied to a 4x4 block.
[0008] The present invention provides the configuration of an area to which 4x4 RST is applied, a method for arranging transform coefficients generated after applying 4x4 RST, a scan order of the arranged transform coefficients, and a method for aligning and matching transform coefficients generated for each block.
[0009] The present invention provides a method for coding a transform index that specifies a 4x4 RST.
[0010] The present invention provides a method for conditionally coding the corresponding transform index by checking whether a non-zero transform coefficient exists in an area that is not allowed when 4x4 RST is applied.
[0011] The present invention provides a method for coding the position of the last non-zero transform coefficient, then conditionally coding its transform index, and then omitting the associated residual coding for impermissible positions.
[0012] The present invention provides a method for applying different transform index coding and residual coding to luma blocks and chroma blocks, respectively, when applying 4x4 RST. [Effects of the Invention]
[0013] The present invention applies 4x4 RST when encoding still images or moving images, and can significantly reduce the amount of calculation compared to when other Nsst (non-separable secondary transform) is applied.
[0014] In addition, taking into account the fact that valid transform coefficients do not exist in a specific region when 4x4 RST is applied, performance can be improved by conditionally coding a transform index specifying 4x4 RST and applying related residual coding optimization.
[0015] In this way, the new low-complexity algorithm can reduce the complexity of the calculations and improve the coding efficiency. [Brief explanation of the drawings]
[0016] [Figure 1] 1 shows a schematic block diagram of an encoder for encoding a video signal as an embodiment to which the present invention is applied. [Figure 2] 1 shows a schematic block diagram of a decoder in which a video signal is decoded as an embodiment to which the present invention is applied. [Figure 3A] 1 is a diagram for explaining a block division structure using QT (QuadTree, hereinafter referred to as "QT") as an embodiment to which the present invention can be applied. [Figure 3B] 1 is a diagram for explaining a block division structure using a Binary Tree (hereinafter referred to as "BT") as an embodiment to which the present invention can be applied. [Figure 3C] 1 is a diagram for explaining a block division structure using a TT (Ternary Tree, hereinafter referred to as "TT") as an embodiment to which the present invention can be applied. FIG. [Figure 3D] 1 is a diagram for explaining a block division structure using an AT (Asymmetric Tree, hereinafter referred to as "AT") as an embodiment to which the present invention can be applied. [Figure 4] As an embodiment to which the present invention is applied, a schematic block diagram of a transform and quantization unit 120 / 130 and an inverse quantization and inverse transform unit 140 / 150 in an encoder is shown. [Figure 5] As an embodiment to which the present invention is applied, a schematic block diagram of an inverse quantization and inverse transform unit 220 / 230 in a decoder is shown. [Figure 6] FIG. 10 is a diagram showing a table of transform configuration groups to which MTS (Multiple Transform Selection) is applied, as an embodiment to which the present invention is applied. [Figure 7] 1 is a flowchart illustrating an encoding process in which MTS (Multiple Transform Selection) is performed as an embodiment to which the present invention is applied. [Figure 8] 1 is a flowchart illustrating a decoding process in which MTS (Multiple Transform Selection) is performed as an embodiment to which the present invention is applied. [Figure 9] 10 is a flowchart illustrating a process of encoding an MTS flag and an MTS index as an embodiment to which the present invention is applied. [Figure 10] 10 is a flowchart illustrating a decoding process for applying a horizontal transform or a vertical transform to a row or a column based on an MTS flag and an MTS index, as an embodiment to which the present invention is applied. [Figure 11] As an embodiment to which the present invention is applied, a flowchart for performing an inverse transformation based on transformation-related parameters is shown. [Figure 12] 10 is a table showing allocation of a transform set for each intra prediction mode in Nsst as an embodiment to which the present invention is applied. [Figure 13] As an embodiment to which the present invention is applied, a calculation flow diagram for Givens rotation is shown. [Figure 14] As an embodiment to which the present invention is applied, a one-round configuration in 4x4 Nsst composed of a Gibbons rotation layer and permutation is shown. [Figure 15]1 is a block diagram illustrating the operations of a forward reduced transform and an inverse reduced transform as an embodiment to which the present invention is applied. [Figure 16] 10 is a diagram illustrating a process of performing a reverse scan from the 64th to the 17th blocks based on the reverse scan order, as an embodiment to which the present invention is applied. [Figure 17] As an embodiment to which the present invention is applied, three forward scan sequences of a transform coefficient block (transform block) are shown. [Figure 18] In an embodiment to which the present invention is applied, when a diagonal scan is applied to the upper left 4x8 block and a 4x4 RST is applied, the positions of valid transform coefficients and the order of forward scan for each 4x4 block are shown. [Figure 19] As an embodiment to which the present invention is applied, a case will be shown in which a diagonal scan is applied to the upper left 4x8 block and valid transform coefficients of two 4x4 blocks are combined into one 4x4 block when a 4x4 RST is applied. [Figure 20] 1 shows a flowchart for encoding a video signal based on a reduced quadratic transform as an embodiment to which the present invention is applied. [Figure 21] 1 shows a flowchart for decoding a video signal based on a reduced quadratic transform as an embodiment to which the present invention is applied. [Figure 22] As an embodiment to which the present invention is applied, a structural diagram of a content streaming system is shown.
[0017] [Best Mode for Carrying Out the Invention] The present invention provides a method for restoring a video signal based on a reduced secondary transform, comprising: obtaining a secondary transform index from the video signal; inducing a secondary transform corresponding to the secondary transform index, where the secondary transform means a reduced secondary transform, and the reduced secondary transform indicates a transform in which N residual data (Nx1 residual vector) are input and L transform coefficient data (Lx1 transform coefficient vector) are output (L < N); performing entropy decoding and inverse quantization on the current block (NxN) to obtain a transform coefficient block; performing an inverse secondary transform on the transform coefficient block using the reduced secondary transform; performing an inverse primary transform on the block on which the inverse secondary transform has been performed; and restoring the current block using the block on which the inverse primary transform has been performed.
[0018] In the present invention, the reduced secondary transform is applied to a specific region of the current block, and the specific region is the upper left MxM (M ≦ N) region within the current block.
[0019] In the present invention, when the inverse secondary transform is performed, for each of the divided 4x4 blocks within the current block, a 4x4 reduced secondary transform is applied.
[0020] In the present invention, whether to obtain the secondary transform index is determined based on the position of the last non-zero transform coefficient within the transform coefficient block.
[0021] In the present invention, when the last non-zero transform coefficient is not located in the specific region, the secondary transform index is obtained, and the specific region indicates the remaining region excluding the positions where non-zero transform coefficients can exist when the transform coefficients are arranged in the scan order when the reduced secondary transform is applied.
[0022] In the present invention, the method includes the step of obtaining a primary conversion index of the current block from the video signal, where the primary conversion index corresponds to any one of a plurality of combinations of conversions composed of a combination of DST7 and / or DCT8, and further includes the step of deriving a combination of conversions corresponding to the primary conversion index. The combination of conversions is composed of a horizontal conversion and a vertical conversion, the horizontal conversion and the vertical conversion correspond to any one of the DST7 or the DCT8, and the inverse primary conversion is performed using the combination of conversions.
[0023] The present invention provides an apparatus for restoring a video signal based on a reduced secondary conversion, including an analysis (parsing) unit that obtains a secondary conversion index from the video signal, a conversion unit that derives a secondary conversion corresponding to the secondary conversion index, where the secondary conversion means a reduced secondary conversion, and the reduced secondary conversion indicates a conversion in which N residual data (Nx1 residual vector) is input and L conversion coefficient data (Lx1 conversion coefficient vector) is output (L < N). The apparatus further includes an entropy decoding unit that performs entropy decoding on the current block (NxN), an inverse quantization unit that performs inverse quantization on the current block after the entropy decoding to obtain a conversion coefficient block, the conversion unit that performs an inverse secondary conversion on the conversion coefficient block using the reduced secondary conversion and performs an inverse primary conversion on the block after the inverse secondary conversion, and a restoration unit that restores the current block using the block after the inverse primary conversion.
Embodiments for Carrying out the Invention
[0024] Hereinafter, the configuration and operation of embodiments of the present invention will be described with reference to the accompanying drawings. The configuration and operation of the present invention described by the drawings are described as one embodiment, and thereby the technical idea, core configuration, and operation of the present invention are not limited.
[0025] In addition, the terms used in this invention have been selected to be as widely used as possible, but in specific cases, terms arbitrarily selected by the applicant will be used for explanation. In such cases, the meanings will be clearly stated in the detailed description of the relevant part, so it is made clear that the terms used in this invention should not be simply analyzed based on their names alone, but should also be analyzed by understanding the meanings of the relevant terms.
[0026] Furthermore, the terms used in the present invention are general terms selected to explain the invention, but if there are other terms with similar meanings, they can be substituted for more appropriate analysis. For example, in the case of signal, data, sample, picture, frame, block, etc., they can be appropriately substituted and analyzed in each coding process. In addition, in the case of partitioning, decomposition, splitting, division, etc., they can be appropriately substituted and analyzed in each coding process.
[0027] In this document, MTS (Multiple Transform Selection, hereinafter referred to as "MTS") may refer to a method of performing a transform using at least two or more transform types. This may also be expressed as AMT (Adaptive Multiple Transform) or EMT (Explicit Multiple Transform). Similarly, mts_idx may also be expressed as AMT_idx, EMT_idx, tu_mts_idx, AMT_TU_idx, EMT_TU_idx, a transform index, or a transform combination index, but the present invention is not limited to such expressions.
[0028] FIG. 1 is a schematic block diagram of an encoder for encoding a video signal, as an embodiment to which the present invention is applied.
[0029] As shown in FIG. 1, the encoder 100 includes an image division unit 110, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, a filtering unit 160, a decoded picture buffer (DPB) 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190.
[0030] The image division unit 110 divides an input image (or picture, frame) input to the encoder 100 into one or more processing units. For example, the processing units may be coding tree units (CTUs), coding units (CUs), prediction units (PUs), or transform units (TUs).
[0031] However, the above terms are used merely for the convenience of explanation of the present invention, and the present invention is not limited to the definitions of the terms. Also, for the convenience of explanation, the present specification uses the term "coding unit" to refer to a unit used in the process of encoding or decoding a video signal, but the present invention is not limited thereto, and can be appropriately analyzed depending on the content of the invention.
[0032] The encoder 100 subtracts a prediction signal output from the inter prediction unit 180 or the intra prediction unit 185 from the input image signal to generate a residual signal, and the generated residual signal is transmitted to the conversion unit 120.
[0033] The transform unit 120 applies a transform technique to the residual signal to generate transform coefficients. The transform process can be applied to square blocks of a quadtree structure, a binary tree structure, a ternary tree structure, or blocks (square or rectangular) divided by an asymmetric tree structure.
[0034] The transform unit 120 may perform a transform based on multiple transforms (or a combination of transforms), which may be referred to as Multiple Transform Selection (MTS), which may also be referred to as Adaptive Multiple Transform (AMT) or Enhanced Multiple Transform (EMT).
[0035] The above (MTS, AMT, EMT) may refer to a conversion method that is performed based on a conversion (or a combination of conversions) adaptively selected from a plurality of conversions (or a combination of conversions).
[0036] The plurality of transforms (or combinations of transforms) may include the transforms (or combinations of transforms) described in Figure 6 herein. In this specification, the transforms or transform types may be expressed as, for example, DCT-Type 2, DCT-II, DCT2, or DCT2.
[0037] The converter 120 may implement the following embodiments.
[0038] The present invention provides a design method for an RST that can be applied to a 4x4 block.
[0039] The present invention provides the configuration of an area to which 4x4 RST is applied, a method for arranging transform coefficients generated after applying 4x4 RST, a scan order of the arranged transform coefficients, and a method for aligning and matching transform coefficients generated for each block.
[0040] The present invention provides a method for coding a transform index that specifies a 4x4 RST.
[0041] The present invention provides a method for conditionally coding the corresponding transform index by checking whether a non-zero transform coefficient exists in an area that is not allowed when 4x4 RST is applied.
[0042] The present invention provides a method for coding the position of the last non-zero transform coefficient, then conditionally coding its transform index, and then omitting the associated residual coding for impermissible positions.
[0043] The present invention provides a method for applying different transform index coding and residual coding to luma blocks and chroma blocks, respectively, when applying 4x4 RST.
[0044] Specific embodiments for this are described in more detail herein.
[0045] The quantization unit 130 quantizes the transform coefficients and transmits the quantized coefficients to the entropy encoding unit 190, and the entropy encoding unit 190 entropy codes the quantized signal and outputs it as a bitstream.
[0046] Although the transform unit 120 and the quantization unit 130 are described as separate functional units, the present invention is not limited thereto and they may be combined into a single functional unit. The inverse quantization unit 140 and the inverse transform unit 150 may also be combined into a single functional unit.
[0047] The quantized signal output from the quantization unit 130 can be used to generate a prediction signal. For example, the quantized signal can be subjected to inverse quantization and inverse transformation via the inverse quantization unit 140 and the inverse transform unit 150 in a loop to reconstruct a residual signal. A reconstructed signal can be generated by adding the reconstructed residual signal to a prediction signal output from the inter prediction unit 180 or the intra prediction unit 185.
[0048] Meanwhile, quantization errors occurring during the compression process can cause visible block boundary degradation. This phenomenon is called blocking artifacts and is one of the important factors in evaluating picture quality. A filtering process can be performed to reduce this degradation. This filtering process removes blocking artifacts and reduces errors in the current picture, thereby improving picture quality.
[0049] The filtering unit 160 applies filtering to the reconstructed signal and outputs the signal to a playback device or transmits it to the decoded picture buffer 170. The filtered signal transmitted to the decoded picture buffer 170 can be used as a reference picture by the inter prediction unit 180. In this way, by using the filtered picture as a reference picture in the inter prediction mode, not only the image quality but also the coding efficiency can be improved.
[0050] The decoded picture buffer 170 can store the filtered pictures for use as reference pictures from the inter predictor 180.
[0051] The inter prediction unit 180 performs temporal prediction and / or spatial prediction to remove temporal redundancy and / or spatial redundancy by referring to a reconstructed picture. Here, the reference picture used for performing the prediction is a transformed signal that has undergone quantization and dequantization in block units during previous encoding / decoding, so blocking artifacts and ringing artifacts may exist.
[0052] Therefore, in order to solve the performance degradation caused by such signal discontinuities and quantization, the inter predictor 180 may apply a low-pass filter to interpolate signals between pixels in sub-pixel units. Here, a sub-pixel refers to a virtual pixel generated by applying an interpolation filter, and an integer pixel refers to an actual pixel present in a reconstructed picture. As an interpolation method, linear interpolation, bi-linear interpolation, a Wiener filter, etc. may be applied.
[0053] The interpolation filter may be applied to a reconstructed picture to improve prediction accuracy. For example, the inter prediction unit 180 may apply an interpolation filter to integer pixels to generate interpolated pixels, and perform prediction using an interpolated block formed of the interpolated pixels as a prediction block.
[0054] Meanwhile, the intra prediction unit 185 may predict a current block by referring to samples surrounding a block to be currently encoded. The intra prediction unit 185 may perform the following processes to perform intra prediction. First, reference samples necessary for generating a prediction signal may be prepared. Then, a prediction signal may be generated using the prepared samples. Then, a prediction mode is encoded. In this case, the reference samples may be prepared through reference sample padding and / or reference sample filtering. Since the reference samples have undergone prediction and reconstruction processes, quantization errors may exist in the reference samples. Therefore, to reduce such errors, a reference sample filtering process may be performed for each prediction mode used in intra prediction.
[0055] A prediction signal generated via the inter prediction unit 180 or the intra prediction unit 185 can be used to generate a reconstructed signal or can be used to generate a residual signal.
[0056] FIG. 2 shows a schematic block diagram of a decoder in which a video signal is decoded, according to an embodiment of the present invention.
[0057] Referring to FIG. 2, the decoder 200 may include an analysis unit (not shown), an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, a filtering unit 240, a decoded picture buffer (DPB) 250, an inter prediction unit 260, and an intra prediction unit 265.
[0058] The reconstructed video signal output through the decoder 200 can then be played back through a playback device.
[0059] The decoder 200 can receive the signal output from the encoder 100 of FIG. 1, and the received signal can be entropy decoded via an entropy decoding unit 210 .
[0060] The inverse quantization unit 220 obtains transform coefficients from the entropy decoded signal using the quantization step size information.
[0061] The inverse transform unit 230 performs an inverse transform on the transform coefficients to obtain a residual signal.
[0062] Here, the present invention provides a method for configuring a transform combination for each transform configuration group classified by at least one of a prediction mode, a block size, or a block shape, and the inverse transform unit 230 can perform an inverse transform based on the transform combination configured according to the present invention. Also, the embodiments described in this specification can be applied.
[0063] The inverse transform unit 230 can execute the following embodiments.
[0064] The present invention provides a method for restoring a video signal based on a reduced secondary transform.
[0065] The inverse transform unit 230 induces a secondary transform corresponding to a secondary transform index, uses the secondary transform to perform an inverse secondary transform on a transform coefficient block, and can perform an inverse primary transform on the block on which the inverse secondary transform has been performed. Here, the secondary transform means a reduced secondary transform, and the reduced secondary transform indicates a transform in which N residual data (Nx1 residual vector) is input and L transform coefficient data (Lx1 transform coefficient vector) is output.
[0066] In the present invention, the reduced secondary transform is applied to a specific region of the current block, and the specific region is the upper left MxM (M≦N) region within the current block.
[0067] In the present invention, when the inverse secondary transform is executed, a 4x4 reduced secondary transform is applied to each of the divided 4x4 blocks within the current block.
[0068] In the present invention, whether to obtain the secondary transform index is determined based on the position of the last non-zero transform coefficient in the transform coefficient block.
[0069] In the present invention, the secondary transform index is obtained when the last non-zero transform coefficient is not located in a specific area, and the specific area indicates the remaining area excluding positions where non-zero transform coefficients can exist when the transform coefficients are arranged according to scan order when the reduced secondary transform is applied.
[0070] The inverse transform unit 230 may derive a transform combination corresponding to a primary transform index and perform an inverse primary transform using the transform combination. Here, the primary transform index corresponds to one of a plurality of transform combinations configured by a combination of DST7 and / or DCT8, and the transform combination is configured by a horizontal transform and a vertical transform. At this time, the horizontal transform and the vertical transform correspond to one of the DST7 or the DCT8.
[0071] Although the inverse quantization unit 220 and the inverse transform unit 230 are described as separate functional units, the present invention is not limited thereto and they may be combined into one functional unit.
[0072] A reconstructed signal is generated by adding the obtained residual signal to a prediction signal output from the inter prediction unit 260 or the intra prediction unit 265.
[0073] The filtering unit 240 applies filtering to the reconstructed signal and outputs it to a playback device or transmits it to the decoded picture buffer unit 250. The filtered signal transmitted to the decoded picture buffer unit 250 can be used as a reference picture in the inter prediction unit 260.
[0074] In this specification, the embodiments described for the transform unit 120 and each functional unit of the encoder 100 can be similarly applied to the inverse transform unit 230 and corresponding functional units of the decoder, respectively.
[0075] Figure 3 is a diagram illustrating block division structures using QT (QuadTree, hereinafter referred to as "QT") as an embodiment to which the present invention can be applied, Figure 3A being a block division structure using QT (QuadTree, hereinafter referred to as "QT"), Figure 3B being a block division structure using BT (Binary Tree, hereinafter referred to as "BT"), Figure 3C being a block division structure using TT (Ternary Tree, hereinafter referred to as "TT"), and Figure 3D being a block division structure using AT (Asymmetric Tree, hereinafter referred to as "AT").
[0076] In video coding, a block can be partitioned based on a Quad Tree (QT). A subblock partitioned by QT can be further partitioned recursively using QT. A leaf block that is not further partitioned by QT can be partitioned using at least one of a Binary Tree (BT), a Ternary Tree (TT), or an Asymmetric Tree (AT). BT has two types of partitioning: horizontal BT (2N×N, 2N×N) and vertical BT (N×2N, N×2N). TT has two types of partitioning: horizontal TT (2N×½N, 2N×N, 2N×½N) and vertical TT (½N×2N, N×2N, ½N×2N). AT has four forms of partitioning: horizontal-up AT (2Nx1 / 2N, 2Nx3 / 2N), horizontal-down AT (2Nx3 / 2N, 2Nx1 / 2N), vertical-left AT (1 / 2Nx2N, 3 / 2Nx2N), and vertical-right AT (3 / 2Nx2N, 1 / 2Nx2N). Each BT, TT, and AT can be further partitioned recursively using BT, TT, and AT.
[0077] 3A shows an example of QT division. Block A is divided into four sub-blocks (A0, A1, A2, A3) by QT. Sub-block A1 is further divided into four sub-blocks (B0, B1, B2, B3) by QT.
[0078] 3B shows an example of BT division. Block B3, which is not further divided by QT, is divided by vertical BT (C0, C1) or horizontal BT (D0, D1). Like block C0, each sub-block can be further divided recursively into horizontal BT (E0, E1) or vertical BT (F0, F1).
[0079] 3C shows an example of TT division. Block B3, which is not further divided by QT, is divided by vertical TT (C0, C1, C2) or horizontal TT (D0, D1, D2). Like block C1, each sub-block can be further divided recursively into horizontal TT (E0, E1, E2) or vertical TT (F0, F1, F2) forms.
[0080] 3D shows an example of AT division. Block B3, which is not further divided by QT, is divided by vertical AT(C0, C1) or horizontal AT(D0, D1). Like block C1, each sub-block can be further divided recursively into horizontal AT(E0, E1) or vertical TT(F0, F1).
[0081] Meanwhile, BT, TT, and AT divisions can be used together for division. For example, sub-blocks divided by BT can be divided by TT or AT. Also, sub-blocks divided by TT can be divided by BT or AT. Sub-blocks divided by AT can be divided by BT or TT. For example, after horizontal BT division, each sub-block can be divided by vertical BT, or after vertical BT division, each sub-block can be divided by horizontal BT. The above two division methods have different division orders, but the final divided shapes are the same.
[0082] Also, when a block is divided, the order of searching the block can be defined in various ways. Generally, searching from left to right and from top to bottom means the order of determining whether to perform additional block division of each divided sub-block, the coding order of each sub-block if the block is not further divided, or the search order when referring to information of other neighboring blocks in a sub-block.
[0083] 4 and 5 show an embodiment to which the present invention is applied. FIG. 4 shows a schematic block diagram of a transform and quantization unit 120 / 130 and an inverse quantization and inverse transform unit 140 / 150 in an encoder, and FIG. 5 shows a schematic block diagram of an inverse quantization and inverse transform unit 220 / 230 in a decoder.
[0084] 4, the transform and quantization unit 120 / 130 includes a primary transform unit 121, a secondary transform unit 122, and a quantization unit 130. The inverse quantization and inverse transform unit 140 / 150 includes an inverse quantization unit 140, an inverse secondary transform unit 151, and an inverse primary transform unit 152.
[0085] As shown in FIG. 5, the inverse quantization and inverse transform unit 220 / 230 includes an inverse quantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.
[0086] In the present invention, transformation can be performed through multiple stages. For example, as shown in Figure 4, two stages of a primary transform and a secondary transform can be applied, or more stages can be used depending on the algorithm. Here, the primary transform can also be called a core transform.
[0087] The primary transform unit 121 applies a primary transform to the residual signal, where the primary transform may already be defined as a table in the encoder and / or decoder.
[0088] For the primary transform, a Discrete Cosine Transform type 2 (hereinafter referred to as "DCT2") may be applied. Alternatively, a Discrete Sine Transform type 7 (hereinafter referred to as "DST7") may be applied in specific cases. For example, DST7 may be applied to a 4x4 block in an intra prediction mode.
[0089] In addition, in the case of the primary transform, a combination of multiple transforms (DST 7, DCT 8, DST 1, DCT 5) of Multiple Transform Selection (MTS) can be applied. For example, FIG. 6 can be applied.
[0090] The secondary transform unit 122 can apply a secondary transform to the primary transformed signal, where the secondary transform can be predefined in a table from the encoder and / or decoder.
[0091] In one embodiment, the secondary transform may be a Non-Separable Secondary Transform (hereinafter referred to as "NSST") that may be conditionally applied. For example, the NSST may be applied only to intra-predicted blocks, and may have a set of applicable transforms for each prediction mode group.
[0092] Here, the prediction mode group may be set based on the symmetry of the prediction direction. For example, since prediction mode 52 and prediction mode 16 are symmetric based on prediction mode 34 (diagonal direction), they form one group and the same transform set may be applied. In this case, when applying the transform of prediction mode 52, input data is transposed before application, because the transform set is the same as that of prediction mode 16.
[0093] On the other hand, for the planar mode and DC mode, there is no directional symmetry, so each mode has a transformation set, which can consist of two transformations. For the remaining directional modes, each transformation set can consist of three transformations.
[0094] In another embodiment, in the case of the secondary transform, a combination of various transforms (DST7, DCT8, DST1, DCT5) of Multiple Transform Selection (MTS) may be applied. For example, Figure 6 can be applied. In another embodiment, DST7 can be applied as a secondary transform.
[0095] In another embodiment, the Nsst may be applied only to the top-left 8x8 region rather than to the entire primary transformed block. For example, if the block size is 8x8 or more, the 8x8 Nsst is applied, and if it is less than 8x8, the 4x4 Nsst is applied. In this case, the block is divided into 4x4 blocks and then the 4x4 Nsst is applied to each block.
[0096] In another embodiment, 4x4 Nsst can be applied even when 4xN / Nx4 (N>=16).
[0097] The Nsst, 4x4 Nsst and 8x8 Nsst will be described in more detail below with reference to FIGS. 12 to 15 and other embodiments of the specification.
[0098] The quantization unit 130 may perform quantization on the quadratic transformed signal. The inverse quantization and inverse transformation unit 140 / 150 performs the above-described process in reverse, and therefore, a duplicated description will be omitted.
[0099] FIG. 5 shows a schematic block diagram of the inverse quantization and inverse transform unit (220 / 230) within the decoder.
[0100] 5, the inverse quantization and inverse transform unit (220 / 230) may include an inverse quantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.
[0101] The inverse quantization unit 220 obtains transform coefficients from the entropy-decoded signal using quantization step size information.
[0102] The inverse secondary transform unit 231 performs an inverse secondary transform on the transform coefficients. Here, the inverse secondary transform refers to the inverse transform of the secondary transform described with reference to FIG.
[0103] In another embodiment, in the case of the secondary transformation, a combination of various transformations (DST7, DCT8, DST1, DCT5) of Multiple Transform Selection (MTS) can be applied. For example, FIG. 6 can be applied.
[0104] The inverse primary transform unit 232 performs an inverse primary transform on the inverse quadratic transformed signal (or block) to obtain a residual signal. Here, the inverse primary transform refers to the inverse transform of the primary transform described in FIG. 4.
[0105] In one embodiment, for the primary transform, a combination of various transforms (DST7, DCT8, DST1, DCT5) of Multiple Transform Selection (MTS) may be applied, for example, as shown in FIG.
[0106] In one embodiment of the present invention, DST7 can be applied as a primary transform.
[0107] In one embodiment of the present invention, a DCT8 may be applied as the primary transform.
[0108] The present invention provides a method for configuring a transform combination for each transform configuration group, which is divided by at least one of a prediction mode, a block size, and a block shape, and the inverse primary transform unit 232 can perform an inverse transform based on the transform combination configured according to the present invention. Also, the embodiments described herein can be applied.
[0109] FIG. 6 is a table showing a transform configuration group to which MTS (Multiple Transform Selection) is applied as an embodiment to which the present invention is applied.
[0110] Transformation settings group to which MTS (Multiple Transform Selection) is applied
[0111] In this specification, candidates for the j-th transformation combination of the transformation setting group Gi are expressed as pairs as shown in the following formula (1).
[0112] [Number 1] (H(Gi, j), V(Gi, j))
[0113] Here, H(Gi,j) refers to the horizontal transform of the jth candidate, and V(Gi,j) refers to the vertical transform of the jth candidate. For example, in FIG. 6, it can be written as H(G3,2) = DST7 and V(G3,2) = DCT8. Depending on the context, the value assigned to H(Gi,j) or V(Gi,j) may be a nominal value for distinguishing the transform, as in the above example, an index value pointing to the transform, or a 2-dimensional matrix of the transform.
[0114] In this specification, the 2D matrix values of the DCT and DST can be expressed as in the following Equations 2 and 3.
[0115] [Number 2] DCT type 2:, DCT type 8:
[0116] [Number 3] DST type 7:, DST type 4:
[0117] Here, DST or DCT is indicated by S or C, the type number is written in superscript Roman numeral form, and the subscript N indicates an NxN transform. Also, for 2D matrices such as the above, it is assumed that the column vector forms the transform basis.
[0118] 6, transform configuration groups are determined based on prediction modes, and there can be a total of six groups (G0 to G5). G0 to G4 correspond to cases where intra prediction is applied, and G5 indicates a combination (or transform set, or transform combination set) of transforms applied to a residual block generated by inter prediction.
[0119] A single transformation combination can consist of a horizontal transform (or row transform) applied to the rows of the 2D block and a vertical transform (or column transform) applied to the columns.
[0120] Here, each transform setting group may have four transform combination candidates, which may be selected or determined by transform combination indexes of 0 to 3, and the transform combination indexes may be encoded and transmitted from the encoder to the decoder.
[0121] In one embodiment, residual data (or residual signals) obtained through intra prediction may have different statistical characteristics depending on the intra prediction mode. Therefore, as shown in FIG. 6, a transform other than the general cosine transform may be applied to each intra prediction mode.
[0122] 6 shows cases where 35 intra prediction modes are used and cases where 67 intra prediction modes are used. A plurality of transform combinations can be applied to each transform setting group in each intra prediction mode column. For example, the plurality of transform combinations can be configured as four (row-wise transform, column-wise transform) combinations. For example, in group 0, DST-7 and DCT-5 can be applied to both the row (horizontal) and column (vertical) directions, resulting in a total of four possible combinations.
[0123] Since a total of four transform kernel combinations can be applied to each intra prediction mode, a transform combination index for selecting one of the four transform kernel combinations can be transmitted for each transform unit. In this specification, the transform combination index can be referred to as an MTS index and can be represented by mts_idx.
[0124] In addition to the transform kernels shown in Figure 6, there may be cases where DCT2 is optimal for all row and column directions due to the characteristics of the residual signal. Therefore, by defining an MTS flag for each coding unit, transforms can be performed adaptively. Here, when the MTS flag is 0, DCT2 is applied to all row and column directions, and when the MTS flag is 1, one of the four combinations can be selected or determined via the MTS index.
[0125] In one embodiment, when the MTS flag is 1, if the number of non-zero transform coefficients for one transform unit is not greater than a threshold, DST-7 can be applied to all rows and columns without applying the transform kernel of Figure 6. For example, the threshold can be set to 2, which can be set differently depending on the block size or the size of the transform unit. This is also applicable to other embodiments of the specification.
[0126] In one embodiment, the values of the transform coefficients are first analyzed, and if the number of non-zero transform coefficients is not greater than a threshold, the amount of additional information transmitted can be reduced by applying DST-7 without analyzing the MTS index.
[0127] In one embodiment, when the MTS flag is 1, if the number of non-zero transform coefficients for one transform unit is greater than a threshold, the MTS index can be analyzed, and the horizontal transform and vertical transform can be determined based on the MTS index.
[0128] In one embodiment, MTS can be applied only if the width and height of the transform unit are both 32 or less.
[0129] In one embodiment, FIG. 6 can be preset through off-line training.
[0130] In one embodiment, the MTS index may be defined as a single index that can simultaneously indicate a combination of horizontal and vertical transforms, or the MTS index may define a horizontal transform index and a vertical transform index separately.
[0131] In one embodiment, the MTS flag or the MTS index may be defined at at least one level of a sequence, a picture, a slice, a block, a coding unit, a transform unit, or a prediction unit. For example, the MTS flag or the MTS index may be defined at at least one level of a sequence parameter set (sps) or a transform unit.
[0132] In another embodiment, the transform combination (horizontal transform, vertical transform) corresponding to the transform index may be configured independently of the MTS flag, prediction mode, and / or block shape. For example, the transform combination may be at least one of DCT2, DST7, and / or DCT8. As a specific example, if the transform index is 0, 1, 2, 3, or 4, the transform combination may be (DCT2, DCT2), (DST7, DST7), (DCT8, DST7), (DST7, DCT8), or (DCT8, DCT8), respectively.
[0133] FIG. 7 is a flowchart showing an encoding process in which MTS (Multiple Transform Selection) is performed as an embodiment to which the present invention is applied.
[0134] Although the present specification basically describes an embodiment in which transforms are applied separately in the horizontal and vertical directions, the combination of transforms can also be configured as a non-separable transform.
[0135] Alternatively, it can be configured with a mixture of separable and non-separable transforms. In this case, if a non-separable transform is used, there is no need to select a transform by row / column or by horizontal / vertical direction. The combination of transforms shown in Figure 6 is used only when a separable transform is selected.
[0136] In addition, the method proposed in this specification can be applied regardless of whether it is a primary transform or a secondary transform. That is, there is no restriction that it must be applied to only one of the two, and it can be applied to both. Here, the primary transform may refer to a transform for initially transforming a residual block, and the secondary transform may refer to a transform for applying a transform to a block generated as a result of the primary transform.
[0137] First, the encoder determines a transform setting group corresponding to the current block (S710). Here, the transform setting group may refer to the transform setting group of FIG. 6, but the present invention is not limited thereto and may be configured with other transform combinations.
[0138] The encoder may perform a transform for a combination of available candidate transforms within the group of transforms (S720).
[0139] After performing the transforms, the encoder can determine or select a combination of transforms that has the smallest RD (Rate Distortion) cost (S730).
[0140] The encoder may encode an index of a transform combination corresponding to the selected transform combination (S740).
[0141] FIG. 8 is a flowchart showing a decoding process performed by MTS (Multiple Transform Selection) as an embodiment to which the present invention is applied.
[0142] First, the decoder may determine a transform setting group for the current block (S810).
[0143] The decoder may analyze (or acquire) a transform combination index from the video signal, where the transform combination index may correspond to any one of a plurality of transform combinations in the transform setting group (S820). For example, the transform setting group may include DST7 (Discrete Sine Transform type 7) and DCT8 (Discrete Cosine Transform type 8). The transform combination index may be referred to as an MTS index.
[0144] In one embodiment, the transform setting group may be set based on at least one of a prediction mode, a block size, or a block shape of the current block.
[0145] The decoder may derive a transform combination corresponding to the transform combination index (S830), where the transform combination includes a horizontal transform and a vertical transform and may include at least one of the DST-7 or DCT-8.
[0146] Furthermore, the combination of transformations may refer to the combination of transformations described in Fig. 6, but the present invention is not limited thereto, and may be configured using other combinations of transformations according to other embodiments of the present specification.
[0147] The decoder may perform an inverse transform of the current block based on the combination of transforms (S840). If the combination of transforms consists of a row (horizontal) transform and a column (vertical) transform, the row (horizontal) transform may be applied first, followed by the column (vertical) transform. However, the present invention is not limited thereto, and if the combination of transforms consists of non-separable transforms, the non-separable transform may be applied immediately.
[0148] In one embodiment, when the vertical transform or the horizontal transform is the DST-7 or DCT-8, the inverse transform of the DST-7 or the inverse transform of the DCT-8 may be applied to each column and then to each row.
[0149] In one embodiment, the vertical transformation or the horizontal transformation may be applied separately to each row and / or each column.
[0150] In one embodiment, the index of the transformation combination may be obtained based on an MTS flag indicating whether MTS is performed, i.e., the index of the transformation combination may be obtained if MTS is performed based on the MTS flag.
[0151] In one embodiment, the decoder may check whether the number of non-zero transform coefficients is greater than a threshold, and the index of the transform combination may be obtained if the number of non-zero transform coefficients is greater than a threshold.
[0152] In one embodiment, the MTS flag or the MTS index may be defined at at least one level of a sequence, a picture, a slice, a block, a coding unit, a transform unit, or a prediction unit.
[0153] In one embodiment, the inverse transform can only be applied if the width and height of the transform unit are both less than or equal to 32.
[0154] Meanwhile, in another embodiment, the process of determining the transform setting group and the process of analyzing the index of the transform combination may be performed simultaneously, or step S810 may be preset in the encoder and / or decoder and may be omitted.
[0155] FIG. 9 is a flowchart illustrating a process of encoding an MTS flag and an MTS index according to an embodiment of the present invention.
[0156] The encoder can determine whether a multiple transform selection (MTS) is applied to the current block (S910).
[0157] If MTS (Multiple Transform Selection) is applied, the encoder may encode with MTS flag=1 (S920).
[0158] The encoder may then determine an MTS index based on at least one of the prediction mode, horizontal transform, and vertical transform of the current block (S930). Here, the MTS index refers to an index indicating one of a plurality of transform combinations for each intra prediction mode, and the MTS index may be transmitted for each transform unit.
[0159] Once the MTS index is determined, the encoder may encode the MTS index (S940).
[0160] On the other hand, if the (MTS Multiple Transform Selection) is not applied, the encoder may encode the MTS flag to 0 (S950).
[0161] FIG. 10 is a flowchart illustrating a decoding process for applying a horizontal transform or a vertical transform to a row or a column based on an MTS flag and an MTS index, as an embodiment to which the present invention is applied.
[0162] The decoder can parse the MTS flag from the bitstream (S1010), where the MTS flag can indicate whether or not MTS (Multiple Transform Selection) of the current block is applied.
[0163] The decoder may determine whether a Multiple Transform Selection (MTS) of the current block is applied based on the MTS flag (S1020). For example, the decoder may determine whether the MTS flag is 1.
[0164] If the MTS flag is 1, the decoder may check whether the number of non-zero transform coefficients is greater than (or equal to) a threshold (S1030). For example, the threshold may be set to 2, which may be set differently depending on the block size or transform unit size.
[0165] If the number of non-zero transform coefficients is greater than a threshold, the decoder may analyze an MTS index (S1040). Here, the MTS index refers to an index indicating one of a plurality of transform combinations for each intra prediction mode or inter prediction mode, and the MTS index may be transmitted for each transform unit. Alternatively, the MTS index may refer to an index indicating one of the transform combinations defined in a pre-defined transform combination table, and the pre-defined transform combination table may refer to FIG. 6, but the present invention is not limited thereto.
[0166] The decoder may derive or determine horizontal and vertical transforms based on at least one of the MTS index or the prediction mode (S1050).
[0167] Alternatively, the decoder may derive a combination of transforms corresponding to the MTS index. For example, the decoder may derive or determine a horizontal transform and a vertical transform corresponding to the MTS index.
[0168] On the other hand, if the number of non-zero transform coefficients is not greater than the threshold, the decoder may apply a preset vertical inverse transform to each column (S1060). For example, the vertical inverse transform may be an inverse transform of DST7.
[0169] Then, the decoder may apply a preset horizontal inverse transform to each row (S1070). For example, the horizontal inverse transform may be the inverse transform of DST7. That is, if the number of non-zero transform coefficients is not greater than a threshold, a transform kernel preset in the encoder or decoder may be used. For example, a commonly used transform kernel may be used instead of one defined in the transform combination table shown in FIG. 6.
[0170] On the other hand, if the MTS flag is 0, the decoder may apply a preset vertical inverse transform to each column (S1080). For example, the vertical inverse transform may be an inverse transform of DCT2.
[0171] Then, the decoder may apply a preset horizontal inverse transform to each row (S1090). For example, the horizontal inverse transform may be an inverse transform of DCT2. That is, if the MTS flag is 0, a transform kernel preset in the encoder or decoder may be used. For example, a commonly used transform kernel may be used instead of one defined in the transform combination table shown in FIG. 6.
[0172] FIG. 11 shows a flowchart for performing an inverse transformation based on transformation-related parameters as an embodiment to which the present invention is applied.
[0173] A decoder to which the present invention is applied can acquire sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag (S1110). Here, sps_mts_intra_enabled_flag indicates whether tu_mts_flag is present in the residual coding syntax of an intra-coding unit. For example, if sps_mts_intra_enabled_flag = 0, tu_mts_flag is not present in the residual coding syntax of an intra-coding unit, and if sps_mts_intra_enabled_flag = 0, tu_mts_flag is present in the residual coding syntax of an intra-coding unit. And sps_mts_inter_enabled_flag indicates whether tu_mts_flag is present in the residual coding syntax of an inter-coding unit. For example, if sps_mts_inter_enabled_flag=0, then tu_mts_flag is not present in the residual coding syntax of the inter coding unit, and if sps_mts_inter_enabled_flag=0, then tu_mts_flag is present in the residual coding syntax of the inter coding unit.
[0174] The decoder may acquire tu_mts_flag based on sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag (S1120). For example, when sps_mts_intra_enabled_flag = 1 or sps_mts_inter_enabled_flag = 1, the decoder may acquire tu_mts_flag. Here, tu_mts_flag indicates whether multiple transform selection (hereinafter referred to as "MTS") is applied to the residual samples of the luma transform block. For example, when tu_mts_flag = 0, MTS is not applied to the residual samples of the luma transform block, and when tu_mts_flag = 1, MTS is applied to the residual samples of the luma transform block.
[0175] As another example, at least one of the embodiments of this document can be applied to the tu_mts_flag.
[0176] The decoder may obtain mts_idx based on tu_mts_flag (S1130). For example, when tu_mts_flag = 1, the decoder may obtain mts_idx, where mts_idx indicates which transform kernel is applied to the luma residual samples along the horizontal and / or vertical directions of the current transform block.
[0177] For example, for mts_idx, at least one of the embodiments in this document can be applied. As a specific example, at least one of the embodiments in FIG. 6 can be applied.
[0178] The decoder can derive a transform kernel corresponding to mts_idx (S1140). For example, the transform kernel corresponding to mts_idx can be defined as a horizontal transform and a vertical transform.
[0179] As another example, different transform kernels may be applied to the horizontal transform and the vertical transform, but the present invention is not limited thereto, and the same transform kernel may be applied to the horizontal transform and the vertical transform.
[0180] In one embodiment, mts_idx can be defined as shown in Table 1 below.
[0181] [Table 1]
[0182] The decoder can then perform an inverse transform based on the transform kernel (S1150).
[0183] A decoding process that performs a conversion process according to another embodiment of the present invention will now be described.
[0184] The decoder may check the transform size (nTbS) (S10), where the transform size (nTbS) may be a variable indicating the horizontal sample size of the scaled transform coefficients.
[0185] The decoder may check a transform kernel type (trType) (S20). Here, the transform kernel type (trType) may be a variable indicating a transform kernel type, and various embodiments of this document may be applied. The transform kernel type (trType) may include a horizontal transform kernel type (trTypeHor) and a vertical transform kernel type (trTypeVer).
[0186] Referring to Table 1, if the transform kernel type (trType) is 0, it indicates DCT2, if it is 1, it indicates DCT7, and if it is 2, it indicates DCT8.
[0187] The decoder can perform transform matrix multiplication based on at least one of a transform size (nTbS) or a transform kernel type (S30).
[0188] As another example, if the transformation kernel type is 1 and the transformation size is 4, a predetermined transformation matrix (1) may be applied when performing the transformation matrix multiplication.
[0189] As another example, if the transformation kernel type is 1 and the transformation size is 8, a transformation matrix (2) determined when performing the transformation matrix multiplication may be applied.
[0190] As another example, if the transformation kernel type is 1 and the transformation size is 16, the previously determined transformation matrix (3) can be applied when performing the transformation matrix multiplication.
[0191] As another example, if the transformation kernel type is 1 and the transformation size is 32, a predefined transformation matrix (4) can be applied.
[0192] Similarly, if the transformation kernel type is 2 and the transformation size is 4, 8, 16, or 32, predefined transformation matrices (5), (6), (7), and (8) can be applied, respectively.
[0193] Here, the predefined transformation matrices (1) to (8) may correspond to any one of various types of transformation matrices. For example, the transformation matrices of the type illustrated in FIG. 6 may be applied.
[0194] The decoder can derive transformed samples based on multiplication of a transform matrix (S40).
[0195] The above embodiments may be used individually, but the present invention is not limited thereto and may be used in combination with the above embodiments and other embodiments described herein.
[0196] FIG. 12 is a table showing allocation of a transform set for each intra prediction mode in NSST as an embodiment to which the present invention is applied.
[0197] Non-Separable Secondary Transform (NSST) The secondary transform unit can apply a secondary transform to the primary transformed signal, where the secondary transform can be predefined in a table in the encoder and / or decoder.
[0198] In one embodiment, the secondary transform may be a Non-Separable Secondary Transform (hereinafter referred to as "NSST") that may be conditionally applied. For example, the NSST may be applied only to intra-predicted blocks, and may have a set of applicable transforms for each prediction mode group.
[0199] Here, the prediction mode group may be set based on the symmetry of the prediction direction. For example, since prediction mode 52 and prediction mode 16 are symmetrical based on prediction mode 34 (diagonal direction), they form one group and the same transform set may be applied. In this case, when applying the transform of prediction mode 52, input data is transposed before application, because the transform set is the same as that of prediction mode 16.
[0200] Meanwhile, in the case of the planar mode and the DC mode, since there is no directional symmetry, each mode has a transformation set, and the transformation set may consist of two transformations. For the remaining directional modes, each transformation set may consist of three transformations. However, the present invention is not limited thereto, and each transformation set may consist of multiple transformations.
[0201] FIG. 13 shows a calculation flow diagram for the Givens rotation as an embodiment to which the present invention is applied.
[0202] In another embodiment, the Nsst may be applied only to the top-left 8x8 region rather than to the entire primary transformed block. For example, if the block size is 8x8 or more, 8x8 Nsst is applied, and if it is less than 8x8, 4x4 Nsst is applied. In this case, the block is divided into 4x4 blocks and then 4x4 Nsst is applied to each block.
[0203] In another embodiment, 4x4 Nsst can be applied even when 4xN / Nx4 (N>=16).
[0204] Both 8x8 Nsst and 4x4 Nsst are non-separable transforms according to the transform combination configuration described in this document, so 8x8 Nsst receives 64 data inputs and outputs 64 data, while 4x4 Nsst has 16 inputs and 16 outputs.
[0205] Both 8x8 Nsst and 4x4 Nsst are composed of a hierarchical combination of Gibbons rotations. The matrix corresponding to one Gibbons rotation is shown in Equation 4, and the matrix product is shown in Equation 5.
[0206]
number
[0207]
number
[0208] As shown in Figure 13, one Gibbons rotation rotates two data, so a total of 32 or 8 Gibbons rotations are required to process 64 (in the case of 8x8 NSST) or 16 (in the case of 4x4 NSST) data, respectively.
[0209] Therefore, a Gibbons rotation layer is formed by bundling 32 or 8 elements. The output data of one Gibbons rotation layer is transmitted as input data to the next Gibbons rotation layer with a predetermined permutation.
[0210] FIG. 14 shows one round configuration of 4x4 Nsst composed of a Gibbons rotation layer and permutation as an embodiment to which the present invention is applied.
[0211] If we look carefully at Figure 14, we can see that in the case of 4x4 Nsst, four Gibbons rotation layers are processed sequentially. The output data of the Gibbons rotation layer shown in Figure 14 undergoes a predetermined permutation (i.e., shuffling) and is transmitted as input data to the next Gibbons rotation layer.
[0212] As shown in FIG. 14, the permutation pattern is determined regularly, and in the case of 4x4 Nsst, four Gibbons rotation layers and their permutations form one round.
[0213] In the case of 8x8 NSST, six Gibbons rotation layers and their permutations form one round. 4x4 NSST undergoes the second round, and 8x8 NSST undergoes the fourth round. Different rounds use the same permutation pattern, but the applied Gibbons rotation angles are different. Therefore, the angle data of all Gibbons rotations that make up each transformation must be saved.
[0214] In the last step, one final substitution is further performed on the data output through the Gibbons rotation layer, and the substitution information is separately saved for each conversion. The substitution is finally performed in the forward NSST, and conversely, in the inverse NSST, first, the inverse substitution is applied.
[0215] In the case of the inverse Nsst, it is advisable to execute the Gibbons rotation layer and substitution applied in the forward Nsst in reverse order, and also rotate each angle of Gibbons rotation by taking a (-) value.
[0216] FIG. 15 is a block diagram for explaining the operations of forward reduced transform and inverse reduced transform as embodiments to which the present invention is applied.
[0217] Reduced Secondary Transform(RST)
[0218] When an orthogonal matrix representing one transformation has the form of NxN, the reduced transform (hereinafter referred to as "RT") is to leave only R out of N transformation basis vectors (R < N). The matrix of the forward RT for generating the transformation coefficients is given as in Equation 6 below.
[0219]
Equation
[0220] Since the matrix of the inverse RT is the transpose matrix of the forward RT matrix, when illustrating the application of the forward RT and the inverse RT, it is as shown in FIG. 15 above.
[0221] Assuming the case of applying the RT to the upper left 8x8 block of the transformation block after the primary transformation, the RT can be named as 8x8 reduced secondary transform (8x8 RST).
[0222] When the R value in Equation 6 is 16, the forward 8x8 RST has the form of a 16x64 matrix, and the backward 8x8 RST has the form of a 64x16 matrix.
[0223] Also, the same configuration of the transform set as in Figure 12 can be applied to the 8x8 RST. That is, a corresponding 8x8 RST can be applied based on the transform set in Figure 12.
[0224] In one embodiment, when one transform set in Figure 12 is composed of two or three transforms depending on the intra prediction mode, one of up to four transforms may be selected, including a case where no quadratic transform is applied, where one transform can be considered as an identity matrix.
[0225] When the four transforms are assigned indices 0, 1, 2, and 3, respectively, a syntax element called the Nsst index can be signaled for each transform block to specify the corresponding transform. That is, for an 8x8 top left block, the Nsst index can specify 8x8 Nsst, and the RST configuration can specify 8x8 RST. In addition, the 0th index can be assigned to the identity matrix, i.e., when no quadratic transform is applied.
[0226] When the forward 8x8 RST as in Equation 6 is applied, 16 valid transform coefficients are generated, so it can be said that 64 input data constituting an 8x8 area is reduced to 16 output data. From the perspective of a two-dimensional area, valid transform coefficients are filled in only about 1 / 4 of the area. Therefore, the 16 output data obtained by applying the forward 8x8 RST are filled in the upper left 4x4 area in Figure 16.
[0227] FIG. 16 is a diagram illustrating a process of performing a backward scan from the 64th to the 17th blocks based on the backward scan order, as an embodiment to which the present invention is applied.
[0228] 16 shows that the forward scan order starts from 1 and scans from the 17th coefficient to the 64th coefficient (in the forward scan order). However, in FIG. 16, the reverse scan is shown, which means that the reverse scan is performed from the 64th coefficient to the 17th coefficient.
[0229] 16, the upper left 4x4 region is a Region of Interest (ROI) to which valid transform coefficients are assigned, and the remaining region is empty, i.e., the remaining region may be assigned a value of 0 by default.
[0230] If there is a valid transform coefficient other than 0 outside the ROI region of FIG. 16, it means that the 8x8 RST is not applied, and in this case, the corresponding Nsst index coding can be omitted.
[0231] Conversely, if there are no non-zero transform coefficients outside the ROI region of Figure 16 (when 8x8 RST is applied, areas outside the ROI are assigned 0), it is possible that 8x8 RST has been applied, so the Nsst index can be coded.
[0232] In this way, the conditional Nsst index coding can be performed after the residual coding process because it must check whether or not there are any transform coefficients that are not zero.
[0233] The present invention deals with the design of RSTs and related optimization methods that can be applied to 4x4 blocks from the RST structure. The embodiments described herein can be applied not only to 4x4 RSTs, but also to 8x8 RSTs or other forms of transformation.
[0234] FIG. 17 shows three forward scan sequences of a transform coefficient block (transform block) as an embodiment to which the present invention is applied.
[0235] Embodiment 1: RST applicable to 4x4 blocks
[0236] The only non-separable transform that can be applied to a 4x4 block is a 16x16 transform. That is, if the data elements that make up the 4x4 block are arranged in row-first or column-first order, they become a 16x1 vector, and a non-separable transform can be applied.
[0237] A forward 16x16 transform is composed of 16 row-wise transform basis vectors, and applying the inner product of the 16x1 vector and each transform basis vector yields the transform coefficients of the transform basis vector. The process of obtaining the corresponding transform coefficients for all 16 transform basis vectors is the same as multiplying the input 16x1 vector by a 16x16 non-separable transform matrix.
[0238] The transform coefficients obtained by matrix multiplication have the form of a 16x1 vector, but each transform coefficient may have different statistical characteristics. For example, if a 16x1 transform coefficient vector is composed of the 0th element to the 15th element, the variance of the 0th element may be greater than the variance of the 15th element. That is, the earlier an element is, the greater its variance value and the greater its energy value.
[0239] By applying an inverse 16x16 non-separable transform from the 16x1 transform coefficients, the original 4x4 block signal can be recovered. If the forward 16x16 non-separable transform is an orthonormal transform, the inverse 16x16 transform can be obtained via the transpose matrix of the forward 16x16 transform.
[0240] By multiplying the inverse 16x16 non-separable transform matrix with the 16x1 transform coefficient vector, we obtain 16x1 vector data, and by arranging it in row-first or column-first order, we can restore the 4x4 block signal.
[0241] As mentioned above, the elements of a 16x1 transform coefficient vector may have different statistical properties.
[0242] If the transform coefficients located earlier (close to the 0th element) have more energy, a signal fairly close to the original can be restored even if an inverse transform is applied to only some of the first transform coefficients without using all of the transform coefficients. For example, if an inverse 16x16 non-separable transform is composed of 16 column basis vectors, a 16xL matrix can be constructed by leaving only L column basis vectors. Also, by leaving only the L most important transform coefficients (Lx1 vector) and multiplying the 16xL matrix by the Lx1 vector, a 16x1 vector with a small error compared to the original input 16x1 vector data can be restored.
[0243] As a result, since only L coefficients are used to restore data, there is no 16x1 transform coefficient vector when obtaining transform coefficients, but an Lx1 transform coefficient vector is obtained. That is, by selecting L row-direction transform vectors from a forward 16x16 non-separable transform matrix to construct an Lx16 transform, and then multiplying it by a 16x1 input vector, the important L transform coefficients can be obtained.
[0244] The value of L has a range of 1≦L<16. Generally, L of the 16 transformation basis vectors can be selected in any way. However, from the viewpoint of encoding and decoding, it may be advantageous in terms of coding efficiency to select transformation basis vectors that are highly important in terms of signal energy.
[0245] Embodiment 2: Setting the application area of 4x4 RST and arranging the conversion coefficients
[0246] The 4x4 RST can be applied to a secondary transform, and in this case, it can be applied secondarily to a block to which a primary transform such as DCT-type 2 has been applied. When the size of a block to which a primary transform has been applied is NxN, the size of a block to which a primary transform has been applied is generally larger than 4x4. Therefore, when applying the 4x4 RST to the NxN block, there are two methods:
[0247] Embodiment 2-1) Instead of applying 4x4 RST to the entire NxN region, it can be applied to only a portion of the region, for example, only to the upper left MxM region (M≦N).
[0248] In embodiment 2-2), the area to which the secondary transformation is applied is divided into 4x4 blocks, and then the 4x4 RST can be applied to each divided block.
[0249] As an embodiment, the above-mentioned embodiments 2-1) and 2-2) may be combined and applied. For example, only the upper left MxM region may be divided into 4x4 blocks, and then 4x4 RST may be applied.
[0250] In one embodiment, a secondary transformation is applied only to the upper left 8x8 region, and if the NxN block is equal to or greater than 8x8, an 8x8 RST is applied. If the NxN block is smaller than 8x8 (4x4, 8x4, 4x8), the NxN block is divided into 4x4 blocks as in embodiment 2-2, and a 4x4 RST can be applied to each block. Also, if the NxN block is 4xN / Nx4 (N>=16), a 4x4 RST can be applied.
[0251] When L (1≦L<16) transform coefficients are generated after applying 4x4 RST, there is a degree of freedom in how to arrange the L transform coefficients. However, since there is a predetermined order when processing the transform coefficients in the residual coding step, coding performance may vary depending on how the L transform coefficients are arranged in a two-dimensional block.
[0252] For example, in the case of HEVC residual coding, coding starts from the position farthest from the DC position in order to improve coding performance by taking advantage of the fact that the value of the quantized coefficient is closer to 0 than it is to 0 as it moves farther away from the DC position.
[0253] Therefore, it may be advantageous in terms of coding performance to arrange the L transform coefficients so that the more important coefficients having higher energy are coded later in the residual coding order.
[0254] Figure 17 shows three forward scan orders for 4x4 transform blocks (Coefficient Group (CG)) applied in HEVC. In residual coding, the scan order shown in Figure 17 is reversed (i.e., coding is performed from 16 to 1).
[0255] Since the three scan orders presented in Figure 17 are selected depending on the intra prediction mode, the present invention can be configured to determine the scan order for the L transform coefficients depending on the intra prediction mode as well.
[0256] FIG. 18 shows an embodiment to which the present invention is applied, in which a diagonal scan is applied to the upper left 4x8 block and a 4x4 RST is applied, and illustrates the positions of valid transform coefficients and the forward scan order for each 4x4 block.
[0257] When the upper left 4x8 block in Figure 17 is divided into 4x4 blocks according to the diagonal scan order and a 4x4 RST is applied to each block, if the value of L is 8 (i.e., only 8 transform coefficients out of 16 remain), the transform coefficients can be positioned as shown in Figure 18.
[0258] Only half of each 4x4 block can have transform coefficients, and the positions where an X appears can be assigned a value of 0 by default.
[0259] Therefore, L transform coefficients are arranged for each 4x4 block according to the scan order shown in Figure 17, and residual coding can be applied to the remaining (16 - L) positions of each 4x4 block, assuming that they are filled with 0s.
[0260] FIG. 19 shows an embodiment in which the present invention is applied, in which a diagonal scan is applied to the upper left 4x8 block and a 4x4 RST is applied, and valid transform coefficients of two 4x4 blocks are combined into one 4x4 block.
[0261] 19, L transform coefficients arranged in two 4x4 blocks can be merged into one. In particular, when the value of L is 8, the transform coefficients of the two 4x4 blocks are merged to completely fill one 4x4 block, leaving no transform coefficients remaining in the other 4x4 block.
[0262] Therefore, for such empty 4x4 blocks, most residual coding is unnecessary, and the corresponding coded_sub_block_flag can be coded to 0.
[0263] In addition, in an embodiment of the present invention, various methods can be applied to how to mix the transform coefficients of two 4x4 blocks. Although they may be mixed in any order, the present invention may provide the following method.
[0264] 1) The transform coefficients of two 4x4 blocks are alternately shuffled in the scan order. That is, in FIG. 18, the transform coefficients of the upper block are JPEG0007769157000005.jpg8170, and the transformation coefficients of the lower block are When the image is JPEG0007769157000006.jpg9170, They can be mixed one after the other, like JPEG0007769157000007.jpg8170, or JPEG0007769157000008.jpg9119 and You can also change the order of JPEG0007769157000009.jpg8117. You can set it so that JPEG0007769157000010.jpg9119 comes first.
[0265] 2) The transform coefficients of the first 4x4 block can be arranged first, followed by the transform coefficients of the second 4x4 block, i.e. You can connect and place it like this: JPEG0007769157000011.jpg9170. Or, You can also change the order, for example, JPEG0007769157000012.jpg8170.
[0266] Embodiment 3: Method for coding NSST index of 4x4 RST
[0267] When the 4x4 RST is applied as shown in FIG. 18, the (L+1)th to 16th transform coefficients can be filled with 0 values according to the scan order of the transform coefficients of each 4x4 block.
[0268] Therefore, if any one of the two 4x4 blocks has a non-zero value in the L + 1 to 16 positions, it can be determined that the 4x4 RST is not applied.
[0269] If the 4x4 RST also has a structure that selects and applies one of a set of prepared transformations like Nsst, it can signal a transformation index (which can be named Nsst index in this embodiment) indicating which transformation to apply.
[0270] Suppose a decoder knows the Nsst index through bitstream analysis and performs such analysis after residual decoding.
[0271] If residual decoding is performed and it is confirmed that there is at least one non-zero transform coefficient between L + 1 and 16, 4x4 RST is not applied, so the Nsst index can be set not to be analyzed.
[0272] Therefore, the Nsst index is selectively analyzed only when necessary, thereby reducing signaling costs.
[0273] 18, if a 4x4 RST is applied to multiple 4x4 blocks within a specific region (for example, the same 4x4 RST may be applied to all of them, or different 4x4 RSTs may be applied to each), a 4x4 RST to be applied to all of the 4x4 blocks can be specified via a single Nsst index. In this case, the same 4x4 RST may be specified, or a 4x4 RST to be applied to each of the 4x4 blocks may be specified.
[0274] Since the 4x4 RST and its application for all the 4x4 blocks are determined by one Nsst index, it can be checked during the residual decoding process whether a non-zero transform coefficient exists at positions L+1 to 16 for all the 4x4 blocks. If the check result shows that a non-zero transform coefficient exists at an impermissible position (L+1 to 16) even in one 4x4 block, it can be configured not to code the Nsst index.
[0275] The Nsst index may be signaled separately for luma blocks and chroma blocks, and in the case of chroma blocks, separate Nsst indices may be signaled for Cb and Cr, or one Nsst index may be shared.
[0276] If Cb and Cr share the same Nsst index, the 4x4 RST specified by the same Nsst index can be applied. In this case, the 4x4 RST for Cb and Cr may be the same, or they may have the same Nsst index but individual 4x4 RSTs.
[0277] To apply the above-mentioned conditional signaling to the shared Nsst index, it can be configured to check that there are non-zero transform coefficients from L + 1 to 16 for all 4x4 blocks of Cb and Cr, and not signal the Nsst index if there are non-zero transform coefficients.
[0278] As shown in Figure 19, even when combining transform coefficients of two 4x4 blocks, when 4x4 RST is applied, it can be determined whether the Nsst index is signaled after checking whether a non-zero transform coefficient exists in a position where no valid transform coefficient exists.
[0279] For example, as shown in Figure 19(b), since the L value is 8, when 4x4 RST is applied to one 4x4 block, if there are no valid transform coefficients (blocks marked with X), the coded_sub_block_flag of the block without valid transform coefficients can be checked. In this case, if the coded_sub_block_flag is 1, the Nsst index can be set not to be signaled.
[0280] Embodiment 4: Optimization method when coding of Nsst index is performed before residual coding
[0281] When coding of the Nsst index is performed before residual coding, whether or not to apply 4x4 RST is determined in advance, so residual coding can be omitted for positions where the transform coefficients are assigned to 0.
[0282] Here, whether or not 4x4 RST is applied can be determined through the Nsst index. For example, if the Nsst index is 0, 4x4 RST is not applied.
[0283] Alternatively, it can be signaled via another syntax element (e.g., Nsst flag). For example, if another syntax element is Nsst flag, the Nsst flag is first parsed to determine whether to apply 4x4 RST. If the value of the Nsst flag is 1, residual coding can be omitted at positions where no valid transform coefficients exist.
[0284] In one embodiment, when performing residual coding, the position of the last non-zero transform coefficient on a TU is coded first. If coding of the Nsst index is performed after coding of the position of the last non-zero transform coefficient, and it is determined that the position of the last non-zero transform coefficient is a position where a non-zero transform coefficient cannot occur when applying 4x4 RST, it is possible to configure not to code the Nsst index and not to apply 4x4 RST.
[0285] For example, when 4x4 RST is applied to the position indicated by X in Figure 18, no valid transform coefficient is located (e.g., a value of 0 can be satisfied), so if the last non-zero transform coefficient is located in the area indicated by X, coding of the Nsst index can be omitted. If the last non-zero transform coefficient is not located in the area indicated by X, coding of the Nsst index can be performed.
[0286] In one embodiment, if the Nsst index is conditionally coded after coding the position of the last non-zero transform coefficient to check whether to apply it to the 4x4 RST, the remaining residual coding portion can be processed using the following two methods.
[0287] 1) If 4x4 RST is not applied, general residual coding is maintained as is, that is, coding is performed under the assumption that a non-zero transform coefficient can exist at any position from the position of the last non-zero transform coefficient to DC.
[0288] 2) When applying 4x4 RST, since there are no transform coefficients for a particular position or a particular 4x4 block (e.g., the X position in Figure 18 can be filled with 0 by default), residual coding can be omitted for the corresponding position or block.
[0289] For example, when reaching the position indicated by X in Figure 18, coding of sig_coeff_flag can be omitted. Here, sig_coeff_flag represents a flag indicating whether a non-zero transform coefficient exists at the corresponding position.
[0290] When combining transform coefficients of two blocks as shown in Figure 19, for 4x4 blocks assigned to 0, the coding of coded_sub_block_flag can be omitted and the corresponding value can be induced to 0, and for the corresponding 4x4 blocks, all values can be induced to 0 without any additional coding.
[0291] When coding the Nsst index after coding the position of the last non-zero transform coefficient, if the x position (Px) and y position (Py) of the last non-zero transform coefficient are smaller than Tx and Ty, respectively, the Nsst index coding can be omitted and 4x4 RST can be configured not to be applied.
[0292] For example, when Tx=1 and Ty=1, it means that Nsst index coding is omitted when the last non-zero transform coefficient exists at the DC position.
[0293] The method of determining whether to perform Nsst index coding through comparison with such a threshold can be applied differently to luma and chroma. For example, different Tx and Ty may be applied to luma and chroma, respectively, or a threshold may be applied to luma but not to chroma, or vice versa.
[0294] The two methods mentioned above, i.e., the first method of omitting Nsst index coding when the last non-zero transform coefficient is located in an area where no valid transform coefficients exist, and the second method of omitting Nsst index coding when the X and Y coordinates of the last non-zero transform coefficient are each smaller than a certain threshold, can also be applied together.
[0295] For example, it is possible to first check the threshold value of the position coordinate of the last non-zero transform coefficient, and then check whether the last non-zero transform coefficient is located in an area where no valid transform coefficients exist. Alternatively, the order can be changed.
[0296] The method presented in this embodiment 4 can also be applied to 8x8 RST. That is, if the last non-zero transform coefficient is located in a non-upper left 4x4 region within the upper left 8x8 region, Nsst index coding can be omitted, and if not, Nsst index coding can be performed.
[0297] Also, if the X and Y coordinate values of the position of the last non-zero transform coefficient are all less than the threshold, coding of the Nsst index can be omitted, or the two methods can be applied together.
[0298] Embodiment 5: When RST is applied, different Nsst index coding and residual coding methods are applied to luma and chroma.
[0299] The schemes described in the third and fourth embodiments can be applied differently to luma and chroma, that is, the Nsst index coding and residual coding schemes can be applied differently to luma and chroma.
[0300] For example, the luma may be applied with the scheme of the fourth embodiment, and the chroma may be applied with the scheme of the third embodiment. Alternatively, the luma may be applied with the conditional Nsst index coding presented in the third or fourth embodiment, and the chroma may not be applied with the conditional Nsst index coding. Or vice versa.
[0301] FIG. 20 shows a flowchart for encoding a video signal based on a reduced secondary transform (RST) as an embodiment to which the present invention is applied.
[0302] The encoder may determine (or select) a forward secondary transform based on at least one of the prediction mode, block shape, and / or block size of the current block (S2010). In this case, the forward secondary transform candidates may include at least one of the embodiments of FIG. 6 and / or FIG. 12.
[0303] The encoder may determine an optimal forward quadratic transform through rate distortion optimization (RD optimization). The optimal forward quadratic transform may correspond to one of a plurality of transform combinations, and the plurality of transform combinations may be defined by a transform index. For example, for RD optimization, results of performing all of the forward quadratic transform, quantization, and residual coding for each candidate may be compared. In this case, a modifier such as cost = rate + λ·distortion or cost = distortion + λ·rate may be used, but the present invention is not limited thereto.
[0304] The encoder may signal a secondary transform index corresponding to the optimal forward secondary transform (S2020), where the secondary transform index may be any of the secondary transform indexes described in other embodiments herein.
[0305] For example, the secondary transform index may be configured as the transform set shown in FIG. 12. Since one transform set is composed of two or three transforms depending on the intra prediction mode, it may be configured to select one of up to four transforms, including cases where no secondary transform is applied. If the four transforms are assigned indexes 0, 1, 2, and 3, respectively, the secondary transform index may be signaled for each transform coefficient block to specify the transform to be applied. In this case, index 0 may be assigned to an identity matrix, i.e., when no secondary transform is applied.
[0306] In another embodiment, the signaling of the secondary transform index may be performed at any one of the following stages: 1) before residual coding, 2) during residual coding (after position coding of the last non-zero transform coefficient), or 3) after residual coding. This embodiment will be described in detail as follows.
[0307] 1) Signaling secondary transform index before residual coding The encoder can determine a forward quadratic transform.
[0308] The encoder may code a secondary transform index corresponding to the forward secondary transform.
[0309] The encoder can code the position of the last non-zero transform coefficient.
[0310] The encoder may perform residual coding of syntax elements other than the position of the last non-zero transform coefficient.
[0311] 2) Signaling secondary transform index during residual coding
[0312] The encoder can determine a forward quadratic transform.
[0313] The encoder can code the position of the last non-zero transform coefficient.
[0314] If the last non-zero transform coefficient is not located in a specific region, the encoder may code an index of a secondary transform corresponding to the forward secondary transform. Here, the specific region indicates a remaining region excluding positions where non-zero transform coefficients may exist when transform coefficients are arranged in a scan order when a reduced secondary transform is applied. However, the present invention is not limited thereto.
[0315] The encoder may perform residual coding of syntax elements other than the position of the last non-zero transform coefficient.
[0316] 3) Signaling secondary transform indexes after residual coding
[0317] The encoder can determine a forward quadratic transform.
[0318] The encoder can code the position of the last non-zero transform coefficient.
[0319] If the last non-zero transform coefficient is not located in a specific region, the encoder may perform residual coding of syntax elements other than the position of the last non-zero transform coefficient. Here, the specific region refers to a remaining region excluding positions where non-zero transform coefficients may exist when transform coefficients are arranged in scan order when a compacted secondary transform is applied. However, the present invention is not limited thereto.
[0320] The encoder may code a secondary transform index corresponding to the forward secondary transform.
[0321] On the one hand, the encoder can perform a forward first-order transformation on the current block (residual block) (S2030). Here, the forward first-order transformation can be similarly applied to the S2010 stage and / or the S2020 stage.
[0322] The encoder can perform a forward second-order transformation on the current block using the optimal forward second-order transformation (S2040). For example, the optimal forward second-order transformation can be a reduced second-order transformation. The reduced second-order transformation means a transformation in which N pieces of residual data (N×1 residual vector) are input and L pieces (L < N) of transformation coefficient data (L×1 transformation coefficient vector) are output.
[0323] As one embodiment, the reduced second-order transformation can be applied to a specific region of the current block. For example, when the current block is N×N, the specific region can mean the upper left N / 2×N / 2 region. However, the present invention is not limited thereto, and it can be set to be different based on at least one of the prediction mode, the shape of the block, or the block size. For example, when the current block is N×N, the specific region can mean the upper left M×M region (M≤N).
[0324] On the one hand, the encoder can generate a transformation coefficient block by quantizing the current block (S2050).
[0325] The encoder can perform entropy encoding on the transformation coefficient block to generate a bit stream.
[0326] FIG. 21 shows a flowchart for decoding a video signal based on a reduced second-order transformation (Reduced Secondary Transform, RST) as an embodiment to which the present invention is applied.
[0327] The decoder can obtain the index of the second-order transform from the bitstream (S2110). Here, the second-order transform index can be applied to other embodiments described in this specification. For example, the second-order transform index can include at least one of the embodiments of FIG. 6 and / or FIG. 12.
[0328] As another embodiment, the obtaining stage of the second-order transform index can be executed at any one of the following stages: 1) before residual decoding, 2) during residual decoding (after decoding the position of the last non-zero transform coefficient), or 3) after residual decoding.
[0329] The decoder can induce a second-order transform corresponding to the second-order transform index (S2120). At this time, the candidates for the second-order transform can include at least one of the embodiments of FIG. 6 and / or FIG. 12.
[0330] However, the steps of S2110 and S2120 are an embodiment, and the present invention is not limited thereto. For example, the decoder can induce a second-order transform based on at least one of the block shape and / or block size in the prediction mode of the current block without obtaining the index of the second-order transform.
[0331] On the other hand, the decoder can entropy-decode the bitstream to obtain a transform coefficient block and perform inverse quantization on the transform coefficient block (S2130)
[0332] The decoder can perform an inverse second-order transform on the inverse-quantized transform coefficient block (S2140). For example, the inverse second-order transform can be a reduced second-order transform. The reduced second-order transform means a transform in which N residual data (Nx1 residual vector) are input and L transform coefficient data (Lx1 transform coefficient vector) are output (L < N).
[0333] In one embodiment, the reduced secondary transform may be applied to a specific region of the current block. For example, if the current block is NxN, the specific region may refer to an N / 2xN / 2 region on the upper left side. However, the present invention is not limited thereto, and the specific region may be set differently based on at least one of a prediction mode, a block shape, or a block size. For example, if the current block is NxN, the specific region may refer to an MxM region (M≦N) or an MxL region (M≦N, L≦N) on the upper left side.
[0334] Then, the decoder may perform an inverse linear transform on the result of the inverse quadratic transform (S2150). The decoder generates a residual block through step S2150, and then adds the residual block to a predicted block to generate a reconstructed block.
[0335] FIG. 22 shows a structural diagram of a content streaming system as an embodiment to which the present invention is applied.
[0336] As shown in FIG. 22, a content streaming system to which the present invention is applied mainly includes an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0337] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0338] The bitstream is generated by an encoding method or a bitstream generating method to which the present invention is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0339] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. Here, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0340] The streaming server receives content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0341] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, smart glass, or a head mounted display (HMD)), a digital TV, a desktop computer, and a digital signage.
[0342] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.
[0343] As described above, the embodiments described in the present invention may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units shown in the drawings may be implemented and performed on a computer, processor, microprocessor, controller, or chip.
[0344] In addition, decoders and encoders to which the present invention is applied can be included in multimedia broadcasting transmitting / receiving devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, custom video (VoD) service providing devices, over-the-top (OTT) video devices, internet streaming service providing devices, three-dimensional (3D) video devices, image telephone video devices, medical video devices, etc., and can be used to process video signals and data signals. For example, over-the-top (OTT) video devices include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0345] Furthermore, a processing method according to the present invention can be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to the present invention can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media realized in the form of carrier waves (e.g., transmission via the Internet). The bitstream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0346] Furthermore, the embodiments of the present invention may be realized as a computer program product with program code, the program code being executed on a computer according to the embodiments of the present invention. The program code may be stored on a computer-readable carrier. [Industrial Applicability]
[0347] The above-described preferred embodiments of the present invention have been disclosed for illustrative purposes, and those skilled in the art will be able to improve, modify, substitute or add various other embodiments within the technical idea and scope of the present invention as disclosed in the appended claims.
Claims
1. A video decoding method performed by a decoder, comprising: deriving quantized transform coefficients for the current block from the bitstream; deriving transform coefficients by performing inverse quantization on the quantized transform coefficients; deriving a residual block based on performing an inverse non-separable transform on the transform coefficients based on an inverse non-separable transform matrix; generating a reconstructed block based on the residual block and a predicted block for the current block; The inverse non-separable transform is a transform in which L transform coefficients are input and N transform coefficients are output, where L is less than N, A method in which a non-separable transform index associated with the inverse non-separable transform matrix is obtained from the bitstream based on the case where the last non-zero transform coefficient is not present at positions from the (L+1)th to the Nth transform coefficients among the transform coefficients of the current block.
2. The inverse non-separable transform is applied to the upper left M×M region of the current block, The method of claim 1 , wherein M≦N.
3. The method described in claim 1, wherein a 4x4 inverse non-separable transform is applied to each of the divided 4x4 blocks within the current block based on which the inverse non-separable transform is performed.
4. A video encoding method performed by an encoder, comprising: deriving a residual block for the current block; deriving transform coefficients for the current block by performing a non-separable transform based on a non-separable transform matrix; performing quantization on the transform coefficients to derive quantized transform coefficients; encoding information about the quantized transform coefficients and non-separable transform indices associated with the non-separable transform matrix to generate a bitstream; The non-separable transform is a transform that receives N transform coefficients and outputs L transform coefficients, where L is less than N; A method configured to obtain the non-separable transform index from the bitstream based on the case where the last non-zero transform coefficient is not present at positions from the (L+1)th to the Nth transform coefficients among the transform coefficients of the current block.
5. A method for transmitting data for video, comprising: obtaining a bitstream for the video, The bitstream comprises: deriving a residual block for the current block; deriving transform coefficients for the current block by performing a non-separable quadratic transform based on a non-separable transform matrix; performing quantization on the transform coefficients to derive quantized transform coefficients; encoding information about the quantized transform coefficients and non-separable transform indices associated with the non-separable transform matrix to generate a bitstream; transmitting the data including the bitstream; A non-separable transform is a transform that takes N transform coefficients as input and outputs L transform coefficients, where L is less than N. A method configured to obtain the non-separable transform index based on the case where the last non-zero transform coefficient from the bitstream does not exist at the (L+1)th to Nth transform coefficient positions among the transform coefficients of the current block.
Citation Information
Patent Citations
Image conversion method and device, and image inverse conversion method and device
JP2013542664A
Method and device for encoding and decoding videos
US20120201303A1
Reduced size inverse transform for decoding and encoding
US20170034530A1
Non-separable secondary transform for video coding
US20170094313A1
Binarizing secondary transform index
US20170324643A1