Video signal encoding / decoding method and apparatus therefor

JP7686813B2Active Publication Date: 2025-06-02LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024014773
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-05
Filing Date
2024-02-02
Publication Date
2025-06-02
Estimated Expiration
2039-09-05

AI Technical Summary

Benefits of technology

【0015】 本発明の実施形態によれば、現在ブロックに適した変換を決定して適用することにより、変換効率を向上させることができる。

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To provide a video signal processing method and apparatus for improving transform efficiency.SOLUTION: A method for decoding a video signal may include the steps of: determining, among predefined secondary transform sets based on intra-prediction modes of a current block, a secondary transform set applied to the current block; obtaining a first syntax element indicating a secondary transform matrix applied to the current block in the determined secondary transform set; deriving a secondary inverse-transformed block by performing a secondary inverse transform on a left top region of the current block by using the secondary transform matrix specified by the first syntax element; and deriving a residual block of the current block by performing a primary inverse transform on the secondary inverse-transformed block by using a primary transform matrix of the current block.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method and apparatus for processing a video signal, and more particularly to a method and apparatus for encoding / decoding a video signal by performing a transformation. [Background technology]

[0002] Compression coding refers to a series of signal processing techniques for transmitting digitalized information over communication lines or storing it in a form suitable for storage media. Media such as video, images, and audio can be subject to compression coding, and the technology that performs compression coding on video in particular is called video compression.

[0003] Next-generation video content will have characteristics such as high spatial resolution, high frame rate, and high dimensionality of scene representation. Processing such content will bring about a huge increase in memory storage, memory access rate, and processing power.

[0004] Therefore, it is necessary to design coding tools to process next-generation video content more efficiently. In particular, video codec standards following the high efficiency video coding (HEVC) standard require more accurate prediction techniques as well as efficient conversion techniques for converting spatial domain video signals to frequency domain. Summary of the Invention [Problem to be solved by the invention]

[0005] It is an object of the present invention to provide a method and device for processing an image signal, which applies a transformation appropriate for a current block.

[0006] The technical problems to be solved by the present invention are not limited to those mentioned above, and other technical problems not mentioned should be clearly understood by a person having ordinary skill in the art to which the present invention pertains from the description below. [Means for solving the problem]

[0007] One aspect of the present invention is a method for decoding a video signal, the method including: determining a secondary transform set to be applied to the current block from among predefined secondary transform sets based on an intra prediction mode of the current block; obtaining a first syntax element indicating a secondary transform matrix to be applied to the current block within the determined secondary transform set; performing a secondary inverse transform on a top left region of the current block using the secondary transform matrix identified by the first syntax element to derive a secondary inverse transformed block; and performing a primary inverse transform on the secondary inverse transformed block using a primary transform matrix of the current block to derive a residual block of the current block.

[0008] Preferably, each of the predefined quadratic transformation sets may include two quadratic transformation matrices.

[0009] Preferably, the step of deriving the quadratic inverse transformed block may further include the step of determining an input length and an output length of the quadratic inverse transform based on a width and a height of the current block.

[0010] Preferably, if the height and width of the current block are each 4, the input length of the non-separable transform can be determined to be 8 and the output length to be 16.

[0011] Preferably, the method may further include the steps of parsing a second syntax element indicating a linear transformation matrix to be applied to a linear transformation of the current block, and determining whether a secondary transformation can be applied to the current block based on the second syntax element.

[0012] Preferably, the step of determining whether a secondary transform can be applied may be performed by determining that a secondary transform can be applied to the current block if the second syntax element indicates a predefined specific transform type.

[0013] Preferably, the predefined specific transform type can be defined as DCT2.

[0014] According to another aspect of the present invention, an apparatus for decoding a video signal includes a memory for storing the video signal, and a processor coupled to the memory, wherein the processor determines a secondary transform set to be applied to the current block from among predefined secondary transform sets based on an intra prediction mode of a current block, obtains a first syntax element indicating a secondary transform matrix to be applied to the current block within the determined secondary transform set, performs a secondary inverse transform on an upper left corner region of the current block using the secondary transform matrix specified by the first syntax element to derive a secondary inverse transformed block, and performs a primary inverse transform on the secondary inverse transformed block using a primary transform matrix of the current block to derive a residual block of the current block. Effect of the Invention

[0015] According to an embodiment of the present invention, the efficiency of transformation can be improved by determining and applying a transformation suitable for a current block.

[0016] In addition, according to an embodiment of the present invention, the transformations used in the primary and secondary transforms can be efficiently designed, thereby improving computational complexity and enhancing compression performance.

[0017] Additionally, embodiments of the present invention can significantly improve computational complexity by limiting the transform kernel of the primary transform to which the secondary transform is applied.

[0018] The effects obtained by the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those having ordinary skill in the art to which the present invention pertains from the following description. [Brief description of the drawings]

[0019] The accompanying drawings, which are included as part of the detailed description to facilitate understanding of the present invention, provide embodiments of the present invention and, together with the detailed description, explain the technical features of the present invention.

[0020] [Figure 1] As an embodiment to which the present invention is applied, an example of a video coding system will be described. [Diagram 2] 1 shows a schematic block diagram of an encoding device for encoding a video / image signal as an embodiment to which the present invention is applied. [Diagram 3] 1 shows a schematic block diagram of a decoding device for decoding a video signal as an embodiment to which the present invention is applied. [Figure 4] FIG. 1 is a structural diagram of a content streaming system as an embodiment to which the present invention is applied. [Figure 5a] 1 is a diagram for explaining a block division structure using QT (QuadTree, hereinafter referred to as "QT") as an embodiment to which the present invention can be applied. FIG. [Figure 5b]1 is a diagram for explaining a block division structure based on a Binary Tree (hereinafter, abbreviated as "BT") as an embodiment to which the present invention can be applied. FIG. [Figure 5c] 1 is a diagram for explaining a block division structure using a Ternary Tree (hereinafter, referred to as "TT") as an embodiment to which the present invention can be applied. FIG. [Figure 5d] 1 is a diagram for explaining a block division structure using an Asymmetric Tree (hereinafter, referred to as "AT") as an embodiment to which the present invention can be applied. [Figure 6] As an embodiment to which the present invention is applied, a schematic block diagram of a transform and quantization unit and an inverse quantization and inverse transform unit in an encoding device is shown. [Figure 7] 1 shows a schematic block diagram of an inverse quantization and inverse transform unit in a decoding device as an embodiment to which the present invention is applied. [Figure 8] 1 is a flowchart showing a process in which AMT (adaptive multiple transform) is performed. [Figure 9] 13 is a flowchart showing a decoding process in which AMT is performed. [Figure 10] 10 is a flowchart illustrating an MTS-based inverse transformation process according to an embodiment of the present invention. [Figure 11] FIG. 2 is a block diagram of an apparatus for performing MTS-based decoding according to an embodiment of the present invention. [Figure 12] 1 is an encoding / decoding flowchart in which a secondary transformation is applied as an embodiment to which the present invention is applied; [Figure 13] 1 is an encoding / decoding flowchart in which a secondary transformation is applied as an embodiment to which the present invention is applied; [Figure 14] As an embodiment to which the present invention is applied, a diagram for explaining Givens rotation is shown. [Figure 15]As an embodiment to which the present invention is applied, a configuration of one round in a 4x4 NSST (non-separable secondary transform) composed of a Givens rotation layer and permutation is shown. [Figure 16] As an embodiment to which the present invention is applied, the operation of a reduced secondary transform (RST) will be shown. [Figure 17] 13 is a diagram illustrating a process of performing a reverse scan from the 64th to the 17th blocks based on a reverse scan order, according to an embodiment of the present invention. [Figure 18] As an embodiment to which the present invention is applied, an example of an encoding flowchart using a single transform indicator (STI) is shown. [Figure 19] As an embodiment to which the present invention is applied, an example of an encoding flowchart using a unified transform indicator (UTI) will be shown. [Figure 20a] Another example of an encoding flowchart using UTI will be shown as an embodiment to which the present invention is applied. [Figure 20b] Another example of an encoding flowchart using UTI will be shown as an embodiment to which the present invention is applied. [Figure 21] As an embodiment to which the present invention is applied, an example of an encoding flowchart for performing conversion will be shown. [Figure 22] As an embodiment to which the present invention is applied, an example of a decoding flowchart for performing conversion is shown. [Diagram 23] As an embodiment to which the present invention is applied, an example of a detailed block diagram of the conversion unit 120 in the encoding device 100 is shown. [Figure 24] As an embodiment to which the present invention is applied, an example of a detailed block diagram of the inverse transform unit 230 in the decoding device 200 is shown. [Diagram 25] 1 shows a flow chart for processing a video signal in an embodiment to which the present invention is applied. [Figure 26] 2 is a flowchart illustrating a method for converting a video signal according to an embodiment of the present invention; [Figure 27] 2 is a flowchart illustrating a method for converting a video signal according to an embodiment of the present invention; [Figure 28] 1 shows an example of a block diagram of an apparatus for processing a video signal as an embodiment to which the present invention is applied. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0021] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. The detailed description disclosed below together with the accompanying drawings is intended to describe exemplary embodiments of the present invention, and is not intended to show the only embodiments in which the present invention can be implemented. The following detailed description includes specific details to provide a complete understanding of the present invention. However, those skilled in the art will appreciate that the present invention can be implemented without such specific details.

[0022] In some cases, in order to avoid obscuring the concept of the present invention, well-known structures and devices may be omitted or shown in the form of a block diagram focusing on the core functions of each structure and device.

[0023] Furthermore, the terms used in the present invention are currently common terms that are widely used whenever possible, but in certain cases, the applicant may use terms arbitrarily selected by the applicant to explain the invention. In such cases, the meanings of the terms will be clearly described in the detailed description of the relevant part, so it is made clear that the terms should not be interpreted simply based on the names of the terms used in the description of the present invention, but should be interpreted while also understanding the meanings of the terms.

[0024] Specific terms used in the following description are provided to aid in understanding the present invention, and the use of such specific terms may be changed to other forms without departing from the technical spirit of the present invention. For example, in the case of a signal, data, sample, picture, frame, block, etc., they may be appropriately substituted and interpreted in each coding process.

[0025] Hereinafter, in this specification, the term "processing unit" refers to a unit in which encoding / decoding processes such as prediction, transformation, and / or quantization are performed. In addition, the processing unit may be interpreted to include a unit of a luma component and a unit of a chroma component. For example, the processing unit may correspond to a block, a coding unit (CU), a prediction unit (PU), or a transform unit (TU).

[0026] In the following description, pixels or picture elements are generally referred to as samples. Using a sample can mean using a pixel value or picture element value.

[0027] Furthermore, the processing units are not necessarily limited to square blocks, but may be configured in the shape of polygons having three or more vertices.

[0028] In the following description, pixels or picture elements are generally referred to as samples. Using a sample can mean using a pixel value or picture element value.

[0029] FIG. 1 shows an example of a video coding system as an embodiment to which the present invention is applied.

[0030] The video coding system may include a source device 10 and a receiving device 20. The source device 10 may transfer encoded video / image information or data to the receiving device 20 in a file or streaming form via a digital storage medium or a network.

[0031] The source device 10 may include a video source 11, an encoding device 12, and a transmitter 13. The receiving device 20 may include a receiver 21, a decoding device 22, and a renderer 23. The encoding device 10 may be referred to as a video / video encoding device, and the decoding device 20 may be referred to as a video / video decoding device. The transmitter 13 may be included in the encoding device 12. The receiver 21 may be included in the decoding device 22. The renderer 23 may include a display unit or may be configured as a separate device or an external component of the display unit.

[0032] A video source may acquire video / video through a video / video capture, synthesis, or generation process. A video source may include a video / video capture device and / or a video / video generation device. A video / video capture device may include, for example, one or more cameras, a video / video archive containing previously captured video / video, etc. A video / video generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / video. For example, a virtual video / video may be generated through a computer, etc., in which case the video / video capture process may be substituted for the process in which the associated data is generated.

[0033] The encoding device 12 can encode the input video / image. The encoding device 12 can perform a series of steps such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0034] The transfer unit 13 can transfer the encoded video / image information or data output in the form of a bitstream to a receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transfer unit 13 can include elements for generating a media file through a predetermined file format and can include elements for transfer via a broadcast / communication network. The receiver 21 can extract the bitstream and transfer it to the decoding device 22.

[0035] The decoding device 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, prediction, etc., corresponding to the operations of the encoding device 12.

[0036] A renderer 23 can render the decoded video / image. The rendered video / image can be displayed via a display unit.

[0037] 2 is a schematic block diagram of an encoding device for encoding a video / image signal as an embodiment to which the present invention is applied. The encoding device 100 in FIG. 2 may correspond to the encoding device 12 in FIG.

[0038] The image division unit 110 may divide an input image (or picture, frame) input to the encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) based on a quad-tree binary-tree (QTBT) structure. For example, one coding unit may be divided into a plurality of coding units of a deeper depth based on a quad-tree structure and / or a binary tree structure. In this case, for example, the quad-tree structure may be applied first and the binary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present invention may be performed based on a final coding unit that is not further divided. In this case, the maximum coding unit may be used as the final coding unit immediately based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as necessary, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0039] The term unit is sometimes used interchangeably with terms such as block or area. In the general case, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally refer to a pixel or a pixel value, or to a pixel / pixel value of the luma component only, or to a pixel / pixel value of the chroma component only. A sample is used as a term to refer to a pixel or pel of a picture (or image).

[0040] The encoding device 100 may generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 180 or the intra prediction unit 185 from an input video signal (original block, original sample array), and the generated residual signal is transferred to the conversion unit 120. In this case, as shown in the figure, a unit in the encoder 100 that subtracts a prediction signal (predicted block, prediction sample array) from an input video signal (original block, original sample array) may be called a subtraction unit 115. The prediction unit may predict a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including a prediction sample of the current block. The prediction unit may determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit may generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transfer it to the entropy encoding unit 190. The prediction information can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0041] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from the current block depending on the prediction mode. In the intra prediction, the prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the fineness of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to the neighboring blocks.

[0042] The inter prediction unit 180 may derive a predicted block of the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between surrounding blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter prediction unit 180 may generate information indicating which candidate is used to derive a motion vector and / or a reference picture index of the current block by forming a motion information candidate list based on neighboring blocks. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode and a merge mode, the inter prediction unit 180 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by signaling the motion vector difference using the motion vector of the neighboring block as a motion vector predictor.

[0043] The prediction signal generated via the inter prediction unit 180 or the intra prediction unit 185 is used to generate a reconstructed signal or is used to generate a residual signal.

[0044] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to pixel blocks having the same square size, or may be applied to non-square variable size blocks.

[0045] The quantizer 130 quantizes the transform coefficients and transmits them to the entropy encoding unit 190, which may encode the quantized signal (information on the quantized transform coefficients) and output it as a bitstream. The information on the quantized transform coefficients may be called residual information. The quantizer 130 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information on the quantized transform coefficients based on the quantized transform coefficients in a one-dimensional vector form. The entropy encoding unit 190 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 190 may encode information required for video / image restoration (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., video / image information) may be transferred or stored in the form of a bitstream in network abstraction layer (NAL) units. The bitstream may be transferred via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storing unit (not shown) for storing the signal may be configured as an element inside / outside the encoding device 100, or the transmitting unit may be a component of the entropy encoding unit 190.

[0046] The quantized transform coefficients output from the quantizer 130 can be used to generate a prediction signal. For example, the quantized transform coefficients can be inverse quantized and inverse transformed through the inverse quantizer 140 and the inverse transformer 150 in the loop to restore a residual signal. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual of the block to be processed, as in the case where the skip mode is applied, the predicted block is used as the reconstructed block. The adder 155 can be called a reconstruction unit or a reconstructed block generator. The generated reconstructed signal is used for intra prediction of the next block to be processed in the current picture, and may be used for inter prediction of the next picture after filtering as described below.

[0047] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and transfer the modified reconstructed picture to the decoded picture buffer 170. Various filtering methods may include, for example, diblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering as described below in the description of each filtering method and transfer it to the entropy encoding unit 190. The information related to filtering may be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0048] The modified reconstructed picture transferred to the decoded picture buffer 170 is used as a reference picture by the inter prediction unit 180. This allows the encoding apparatus to avoid prediction mismatch between the encoding apparatus 100 and the decoding apparatus when inter prediction is applied, and also improves encoding efficiency.

[0049] The decoded picture buffer 170 may store the modified reconstructed picture for use as a reference picture from the inter predictor 180 .

[0050] 3 is a schematic block diagram of a decoding device for decoding a video signal as an embodiment to which the present invention is applied. A decoding device 200 in FIG. 3 may correspond to the decoding device 22 in FIG.

[0051] Referring to FIG. 3, the decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a decoded picture buffer (DPB) 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a prediction unit. That is, the prediction unit may include an inter prediction unit 180 and an intra prediction unit 185. The inverse quantization unit 220 and the inverse transform unit 230 may be collectively referred to as a residual processing unit. That is, the residual processing unit may include an inverse quantization unit 220 and an inverse transform unit 230. The entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the adder 235, the filtering unit 240, the inter prediction unit 260, and the intra prediction unit 265 may be configured by one hardware component (e.g., a decoder or a processor) according to an embodiment. Additionally, the decoded picture buffer 250 may be implemented by a single hardware component (eg, a memory or digital storage medium) depending on the embodiment.

[0052] When a bitstream including video / image information is input, the decoding apparatus 200 can restore an image corresponding to a process in which the video / image information from the encoding apparatus 100 of Fig. 2 has been processed. For example, the decoding apparatus 200 can perform decoding using a processing unit applied in the encoding apparatus 100. Thus, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure and / or a binary tree structure. The restored image signal decoded and output through the decoding apparatus 200 can be reproduced through a reproduction device.

[0053] The decoding device 200 may receive a signal output from the encoding device 100 of Fig. 2 in the form of a bitstream, and the received signal may be decoded through the entropy decoding unit 210. For example, the entropy decoding unit 210 may derive information (e.g., video / image information) required for image restoration (or picture restoration) by farsing (analyzing) the bitstream. For example, the entropy decoding unit 210 may decode information in the bitstream based on a coding method such as exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to residuals. In more detail, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information of the syntax element to be decoded and decode information of the neighboring and blocks to be decoded, or information of symbols / bins decoded in a previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method may update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin. Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to a prediction unit (inter prediction unit 260 and intra prediction unit 265), and residual values ​​entropy decoded from the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. In addition, information regarding filtering among the information decoded by the entropy decoding unit 210 may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the encoding device 100 may be further configured as an internal / external element of the decoding device 200, or the receiving unit may be a component of the entropy decoding unit 210.

[0054] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients to output transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients into a two-dimensional block shape. In this case, the rearrangement may be performed based on the coefficient scan order performed in the encoding apparatus 100. The inverse quantization unit 220 may perform inverse quantization of the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0055] The inverse transform unit 230 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0056] The prediction unit may perform prediction of the current block and generate a predicted block including a prediction sample of the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode.

[0057] The intra prediction unit 265 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 265 may also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.

[0058] The inter prediction unit 260 may derive a predicted block of the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between surrounding blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may configure a motion information candidate list based on the surrounding blocks and derive a motion vector and / or a reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information on the prediction may include information indicating the mode of inter prediction of the current block.

[0059] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstruction, sample array) by adding the acquired residual signal to a prediction signal (predicted block, predicted sample array) output from the inter prediction unit 260 or the intra prediction unit 265. As in the case where the skip mode is applied, if there is no residual for the block to be processed, the predicted block is used as the reconstructed block.

[0060] The adder 235 may be referred to as a reconstruction unit or a reconstruction block generator. The reconstruction signal generated may be used for intra prediction of the next block to be processed in the current picture, and may be used for inter prediction of the next picture after filtering as described below.

[0061] The filtering unit 240 may improve the subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transfer the modified reconstructed picture to the decoded picture buffer 250. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), an adaptive loop filter (ALF), a bilateral filter, etc.

[0062] The modified reconstructed picture transferred to the decoded picture buffer 250 is used as a reference picture by the inter predictor 260 .

[0063] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the encoding device 100 can also be applied in an identical or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of each decoding device.

[0064] FIG. 4 is a structural diagram of a content streaming system according to an embodiment of the present invention.

[0065] The content streaming system to which the present invention is applied can broadly include an encoding server 410 , a streaming server 420 , a Web server 430 , a media storage 440 , a user device 450 and a multimedia input device 460 .

[0066] The encoding server 410 compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and transmits the bitstream to the streaming server 420. As another example, if a multimedia input device 460 such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server 410 can be omitted.

[0067] The bitstream may be generated by an encoding method or a bitstream generating method to which the present invention is applied, and the streaming server 420 may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0068] The streaming server 420 transfers multimedia data to the user device 450 based on a user request via the web server 430, and the web server 430 acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server 430, the web server 430 transfers it to the streaming server 420, and the streaming server 420 transfers the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays a role of controlling commands / responses between each device in the content streaming system.

[0069] The streaming server 420 may receive content from the media storage 440 and / or the encoding server 410. For example, when content is received from the encoding server 410, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server 420 may store the bitstream for a certain period of time.

[0070] Examples of user devices 450 include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glass, head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, and the like.

[0071] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.

[0072] FIG. 5 is a diagram for explaining block division structures according to an embodiment to which the present invention can be applied, in which FIG. 5a is a QT (QuadTree, QT), FIG. 5b is a BT (Binary Tree, BT), FIG. 5c is a TT (Ternary Tree, TT), and FIG. 5d is an AT (Asymmetric Tree, AT).

[0073] In video coding, a block can be divided based on QT. A subblock divided by QT can be further divided recursively using QT. A leaf block that is no longer divided by QT can be divided by at least one of BT, TT, or AT. BT can have two types of division: horizontal BT (2NxN, 2NxN) and vertical BT (Nx2N, Nx2N). TT can have two types of division: horizontal TT (2Nx1 / 2N, 2NxN, 2Nx1 / 2N) and vertical TT (1 / 2Nx2N, Nx2N, 1 / 2Nx2N). AT can have four types of partitions: horizontal-up AT (2Nx1 / 2N, 2Nx3 / 2N), horizontal-down AT (2Nx3 / 2N, 2Nx1 / 2N), vertical-left AT (1 / 2Nx2N, 3 / 2Nx2N), vertical-right AT (3 / 2Nx2N, 1 / 2Nx2N). Each BT, TT, and AT can be further partitioned recursively using BT, TT, and AT.

[0074] Figure 5a shows an example of QT partitioning. Block A can be partitioned by QT into four sub-blocks (A0, A1, A2, A3). Sub-block A1 can be partitioned again by QT into four sub-blocks (B0, B1, B2, B3).

[0075] Figure 5b shows an example of BT partitioning. Block B3, which is no longer divided by QT, can be partitioned into vertical BT (C0, C1) or horizontal BT (D0, D1). Like block C0, each subblock can be further partitioned recursively, such as in the form of horizontal BT (E0, E1) or vertical BT (F0, F1).

[0076] Figure 5c shows an example of a TT partition. Block B3, which is no longer partitioned by QT, can be partitioned into vertical TTs (C0, C1, C2) or horizontal TTs (D0, D1, D2). Like block C1, each subblock can be further partitioned recursively, such as in the form of horizontal TTs (E0, E1, E2) or vertical TTs (F0, F1, F2).

[0077] Figure 5d shows an example of an AT partition. Block B3, which is no longer partitioned by QT, can be partitioned into vertical AT(C0,C1) or horizontal AT(D0,D1). Like block C1, each subblock can be further partitioned recursively, such as in the form of horizontal AT(E0,E1) or vertical TT(F0,F1).

[0078] Meanwhile, BT, TT, and AT divisions can be used together for division. For example, a subblock divided by BT can be divided by TT or AT. Also, a subblock divided by TT can be divided by BT or AT. A subblock divided by AT can be divided by BT or TT. For example, after horizontal BT division, each subblock can be divided into vertical BT, or after vertical BT division, each subblock can be divided into horizontal BT. In this case, the division order is different, but the final division form is the same.

[0079] In addition, when a block is divided, the order of searching the block can be defined in various ways. In general, searching a block from left to right and from top to bottom can mean the order of determining whether to divide an additional block of each divided sub-block, or the coding order of each sub-block if the block is no longer divided, or the search order when a sub-block refers to information of other adjacent blocks.

[0080] A transform may be performed for each processing unit (or transform block) divided according to the division structure shown in Figures 5a to 5d, and in particular, a transform matrix may be applied to each divided row and column. According to an embodiment of the present invention, different transform types are used depending on the length of the processing unit (or transform block) in the row or column direction.

[0081] The transform is applied to the residual block in order to decorrelate it as much as possible, concentrate the coefficients in low frequencies, and create a zero tail at the beginning of the block. In the JEM software, the transform part includes two main functions: core transform and secondary transform. The core transform consists of the DCT (discrete cosine transform) and DST (discrete sine transform) transform families that are applied to every row and column of the residual block. The secondary transform can then be additionally applied to the top left corner of the output of the core transform. Similarly, the secondary inverse transform and the inverse transform of the core inverse transform order can be applied. First, the secondary inverse transform can be applied to the top left corner of the coefficient block. The core inverse transform is then applied to the rows and columns of the output of the secondary inverse transform. The core transform or inverse transform can be referred to as the primary transform or inverse transform.

[0082] 6 and 7 show an embodiment to which the present invention is applied. FIG. 6 shows a schematic block diagram of the transform and quantization unit (120 / 130) and the inverse quantization and inverse transform unit (140 / 150) in the encoding device 100 of FIG. 2, and FIG. 7 shows a schematic block diagram of the inverse quantization and inverse transform unit (220 / 230) in the decoding device 200.

[0083] 6, the transform and quantization unit (120 / 130) may include a primary transform unit 121, a secondary transform unit 122, and a quantization unit 130. The inverse quantization and inverse transform unit (140 / 150) may include an inverse quantization unit 140, an inverse secondary transform unit 151, and an inverse primary transform unit 152.

[0084] Looking carefully at FIG. 7, the inverse quantization and inverse transform unit (220 / 230) may include an inverse quantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.

[0085] In the present invention, when performing a transformation, the transformation can be performed through multiple stages. For example, as shown in Figure 6, two stages of a primary transform and a secondary transform can be applied, or more transformation stages can be used based on the algorithm. Here, the primary transform can be referred to as a core transform.

[0086] The linear transform unit 121 may apply a linear transform to the residual signal, where the linear transform may be predefined in a table from the encoder and / or decoder.

[0087] The secondary transform unit 122 can apply a secondary transform to the primary transformed signal, where the secondary transform can be predefined in a table from the encoder and / or decoder.

[0088] In one embodiment, a non-separable secondary transform (NSST) may be conditionally applied as the secondary transform. For example, the NSST may be applied only for intra-predicted blocks, and may have a set of applicable transforms for each prediction mode group.

[0089] Here, the prediction mode group may be set based on the symmetry of the prediction direction. For example, since the prediction mode 52 and the prediction mode 16 are symmetric based on the prediction mode 34 (diagonal direction), they form one group and the same transform set may be applied. In this case, when applying the transform of the prediction mode 52, the input data is transposed before being applied, because the transform set is the same as that of the prediction mode 16.

[0090] On the other hand, Planar mode and DC mode have their own transformation sets, each consisting of two transformations, since there is no directional symmetry, while the remaining directional modes have three transformations per transformation set.

[0091] The quantization unit 130 may perform quantization on the quadrature-transformed signal.

[0092] The inverse quantization and inverse transformation unit (140 / 150) performs the above-described process inversely, and a duplicated description will be omitted.

[0093] FIG. 7 shows a schematic block diagram of the inverse quantization and inverse transform unit (220 / 230) in the decoding device 200.

[0094] Referring to FIG. 7, the inverse quantization and inverse transform unit (220 / 230) may include an inverse quantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.

[0095] The inverse quantization unit 220 uses the quantization step size information to obtain transform coefficients from the entropy decoded signal.

[0096] The inverse quadratic transform unit 231 executes an inverse quadratic transform on the transform coefficients. Here, the inverse quadratic transform refers to the inverse transform of the quadratic transform described in FIG.

[0097] The inverse primary transform unit 232 performs an inverse primary transform on the inverse secondary transformed signal (or block) to obtain a residual signal. Here, the inverse primary transform refers to an inverse transform of the primary transform described in FIG. 6.

[0098] In addition to the DCT-2 and 4x4 DST-4 applied in HEVC, the adaptive (or explicit) multiple transform (AMT or EMT) technique is used for residual coding of inter- and intra-encoded blocks. In addition to the HEVC transforms, a number of selected transforms from other DCT / DST families are used. The newly introduced transform matrices from JEM are DST-7, DCT-8, DST-1, and DCT-5. Table 1 below shows the selected DST / DCT basis functions.

[0099] [Table 1]

[0100] EMT can be applied to CUs with width and height less than or equal to 64, and whether EMT is applied can be controlled by a CU level flag. If the CU level flag is 0, DCT-2 is applied to the CU to encode the residual. For intra-CU luma coding blocks to which EMT is applied, two additional flags are signaled to identify the horizontal and vertical transforms to be used. Like HEVC, the residuals of blocks in JEM can be coded in transform skip mode. For intra residual coding, a mode-dependent transform candidate selection process is used, depending on the residual statistics of other intra prediction modes. Three transform subsets are defined as shown in Table 2 below, and the transform subset is selected based on the intra prediction mode as shown in Table 3.

[0101] [Table 2]

[0102] With the subset concept, a subset of transforms is first identified based on Table 2 by using intra prediction modes of CUs with CU-level EMT_CU_flag set to 1. Then, for each horizontal (EMT_TU_horizontal_flag) and vertical (EMT_TU_vertical_flag) transform, one of two transform candidates in the identified subset of transforms is selected based on explicit signaling using flags based on Table 3.

[0103] [Table 3]

[0104] [Table 4]

[0105] Table 4 shows a transform configuration group to which AMT (adaptive multiple transform) is applied as an embodiment to which the present invention is applied.

[0106] Looking carefully at Table 4, a transform configuration group is determined based on a prediction mode, and there can be a total of six groups (G0 to G5). G0 to G4 correspond to cases where intra prediction is applied, and G5 indicates a combination (or transform set, transform combination set) of transforms applied to a residual block generated by inter prediction.

[0107] A combination of transformations can be done by a horizontal transform (or row transform) applied to the rows of the 2D block and a vertical transform (or column transform) applied to the columns.

[0108] Here, each group of transform settings may have four transform combination candidates. The four transform combination candidates may be selected or determined via transform combination indexes of 0 to 3, and the transform combination indexes may be encoded and transmitted from the encoder to the decoder.

[0109] In one embodiment, residual data (or residual signals) acquired through intra prediction may have different statistical characteristics depending on the intra prediction mode. Therefore, a different transform may be applied to each intra prediction mode instead of a general cosine transform as shown in Table 4. In this specification, the transform type may be expressed as, for example, DCT-Type 2, DCT-II, DCT-2, etc.

[0110] Looking carefully at Table 4, there are cases where 35 intra prediction modes are used and cases where 67 intra prediction modes are used. For each transform setting group divided by the column of each intra prediction mode, multiple transform combinations can be applied. For example, multiple transform combinations can be composed of four (row-directional transform, column-directional transform) combinations. As a specific example, in group 0, DST-7 and DCT-5 can be applied in both the row (horizontal) and column (vertical) directions, and a total of four combinations are possible.

[0111] Since a total of four transform kernel combinations can be applied to each intra prediction mode, a transform combination index for selecting one of them can be transmitted for each transform unit. In this specification, the transform combination index can be referred to as an AMT index and can be represented by amt_idx.

[0112] In addition to the transform kernels presented in Table 4, there may be cases where DCT-2 is optimal for all row and column directions due to the characteristics of the residual signal. Therefore, by defining an AMT flag for each coding unit, transforms can be applied adaptively. Here, when the AMT flag is 0, DCT-2 is applied to all row and column directions, and when the AMT flag is 1, one of the four combinations can be selected or determined using the AMT index.

[0113] In one embodiment, when the AMT flag is 0, if the number of transform coefficients in one transform unit is less than 3, the transform kernel in Table 4 is not applied, and DST-7 can be applied in all row and column directions.

[0114] In one embodiment, the values ​​of the transform coefficients are parsed first, and if the number of transform coefficients is less than 3, the AMT index is not parsed and DST-7 is applied, thereby reducing the amount of additional information transmitted.

[0115] In one embodiment, AMT can be applied only if the width and height of the transform unit are both less than or equal to 32.

[0116] In one embodiment, Table 4 can be populated via off-line training.

[0117] In one embodiment, the AMT index can be defined as one index that can simultaneously refer to a combination of horizontal and vertical transforms, or the AMT index can be defined as separate horizontal and vertical transform indexes.

[0118] FIG. 8 is a flowchart showing a process in which AMT (adaptive multiple transform) is performed.

[0119] Although this specification primarily describes an embodiment of a separable transform in which transformations are applied separately to the horizontal and vertical directions, the combination of transforms can also be configured to be non-separable transforms.

[0120] Alternatively, a combination of transforms can be constructed by mixing separable and non-separable transforms. In this case, when a non-separable transform is used, row / column or horizontal / vertical transform selection is not necessary, and the transform combinations in Table 4 can be used only when a separable transform is selected.

[0121] In addition, the method proposed in this specification can be applied regardless of whether it is a linear transform or a secondary transform. That is, there is no restriction that it must be applied to only one of the two, and it can be applied to both. Here, the linear transform can refer to a transform for initially transforming a residual block, and the secondary transform can refer to a transform for applying a transform to a block generated as a result of the linear transform.

[0122] First, the encoding device 100 may determine a transform group corresponding to a current block (S805). Here, the transform group may refer to the transform group in Table 4, but the present invention is not limited thereto and may be configured with other transform combinations.

[0123] The encoding device 100 may perform a transform on a combination of candidate transforms to be used in the transform group (S810). As a result of the transform execution, the encoding device 100 may determine or select a combination of transforms with a minimum rate distortion (RD) cost (S815). The encoding device 100 may encode an index of the transform combination corresponding to the selected transform combination (S820).

[0124] FIG. 9 is a flow chart showing the decoding process in which AMT is performed.

[0125] First, the decoding apparatus 200 may determine a transform group for a current block (S905). The decoding apparatus 200 may analyze an index of a transform combination, where the index of the transform combination may correspond to any one of a plurality of transform combinations in the transform group (S910). The decoding apparatus 200 may derive a transform combination corresponding to the index of the transform combination (S915). Here, the transform combination may refer to the transform combination described in Table 4, but the present invention is not limited thereto. That is, configurations using other transform combinations are also possible.

[0126] The decoding apparatus 200 may perform an inverse transform on the current block based on the transform combination (S920). If the transform combination is composed of a row transform and a column transform, the row transform may be applied first, and then the column transform may be applied. However, the present invention is not limited thereto, and if the transform combination is composed of a non-separable transform, the non-separable transform may be applied directly.

[0127] Meanwhile, in another embodiment, the process of determining the transformation group and the process of analyzing the transformation combination index may be performed simultaneously.

[0128] According to an embodiment of the present invention, the above-mentioned term "AMT" can be redefined as "MTS (multiple transform set or multiple transform selection)". The MTS-related syntax and semantics described below are defined in the versatile video coding (VVC) standard document JVET-K1001-v4.

[0129] In an embodiment of the present invention, two MTS candidates are used for directional modes and four MTS candidates for non-directional modes as follows:

[0130] A) Non-directional mode (DC, Planner)

[0131] When the MTS index is 0, DST-7 is used for horizontal and vertical transformation.

[0132] When the MTS index is 1, DCT-7 is used for the vertical transform and DCT-8 is used for the horizontal transform.

[0133] When the MTS index is 2, a DCT-8 is used for the vertical transform and a DST-7 is used for the horizontal transform.

[0134] When the MTS index is 3, the DCT-8 is used for the horizontal and vertical transforms.

[0135] B) Modes that belong to the horizontal group mode

[0136] When the MTS index is 0, DST-7 is used for horizontal and vertical transformation.

[0137] When the MTS index is 1, a DCT-8 is used for the vertical transform and a DST-7 is used for the horizontal transform.

[0138] C) Modes belonging to the vertical group mode

[0139] When the MTS index is 0, DST-7 is used for horizontal and vertical transformation.

[0140] When the MTS index is 1, DCT-7 is used for the vertical transform and DCT-8 is used for the horizontal transform.

[0141] Here (in VTM 2.0 where 67 modes are used), the horizontal group modes include intra prediction modes 2 through 34, and the vertical modes include intra prediction modes 35 through 66.

[0142] In another embodiment of the present invention, three MTS candidates are used for every intra prediction mode.

[0143] When the MTS index is 0, DST-7 is used for horizontal and vertical transformation.

[0144] When the MTS index is 1, DCT-7 is used for the vertical transform and DCT-8 is used for the horizontal transform.

[0145] When the MTS index is 2, a DCT-8 is used for the vertical transform and a DST-7 is used for the horizontal transform.

[0146] In another embodiment of the present invention, two MTS candidates are used for directional prediction modes and three MTS candidates for non-directional prediction modes.

[0147] A) Non-directional mode (DC, Planner)

[0148] When the MTS index is 0, DST-7 is used for horizontal and vertical transformation.

[0149] When the MTS index is 1, DCT-7 is used for the vertical transform and DCT-8 is used for the horizontal transform.

[0150] When the MTS index is 2, a DCT-8 is used for the vertical transform and a DST-7 is used for the horizontal transform.

[0151] B) Prediction mode corresponding to horizontal group mode

[0152] When the MTS index is 0, DST-7 is used for horizontal and vertical transformation.

[0153] When the MTS index is 1, a DCT-8 is used for the vertical transform and a DST-7 is used for the horizontal transform.

[0154] C) Prediction mode corresponding to vertical group mode

[0155] When the MTS index is 0, DST-7 is used for horizontal and vertical transformation.

[0156] When the MTS index is 1, DCT-7 is used for the vertical transform and DCT-8 is used for the horizontal transform.

[0157] In another embodiment of the present invention, one MTS candidate (e.g., DST-7) is used for all intra modes. In this case, encoding time can be reduced by up to 40% with minor coding loss. Furthermore, one flag is used to indicate between DCT-2 and DST-7.

[0158] FIG. 10 is a flow chart illustrating an MTS-based inverse transformation process according to an embodiment of the present invention.

[0159] The decoding device 200 to which the present invention is applied may acquire sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag (S1005). Here, sps_mts_intra_enabled_flag indicates whether cu_mts_flag exists in the residual coding syntax of the intra coding unit. For example, if sps_mts_intra_enabled_flag=0, cu_mts_flag does not exist in the residual coding syntax of the intra coding unit, and if sps_mts_intra_enabled_flag=1, cu_mts_flag exists in the residual coding syntax of the intra coding unit. And sps_mts_inter_enabled_flag indicates whether cu_mts_flag exists in the residual coding syntax of the inter coding unit. For example, if sps_mts_inter_enabled_flag=0, then cu_mts_flag is not present in the residual coding syntax of the inter coding unit, and if sps_mts_inter_enabled_flag=1, then cu_mts_flag is present in the residual coding syntax of the inter coding unit.

[0160] The decoding device 200 may acquire cu_mts_flag based on sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag (S1010). For example, when sps_mts_intra_enabled_flag=1 or sps_mts_inter_enabled_flag=1, the decoding device 200 may acquire cu_mts_flag. Here, cu_mts_flag indicates whether MTS is applied to the residual sample of the luma transform block. For example, when cu_mts_flag=0, MTS is not applied to the residual sample of the luma transform block, and when cu_mts_flag=1, MTS is applied to the residual sample of the luma transform block.

[0161] The decoding apparatus 200 may acquire mts_idx based on cu_mts_flag (S1015). For example, when cu_mts_flag=1, the decoding apparatus 200 may acquire mts_idx, where mts_idx indicates which transform kernel is applied to luma residual samples along the horizontal and / or vertical directions of the current transform block.

[0162] For example, for mts_idx, at least one of the embodiments described in this specification can be applied.

[0163] The decoding apparatus 200 may derive a transform kernel corresponding to mts_idx (S1020). For example, the transform kernel corresponding to mts_idx may be defined by being divided into a horizontal transform and a vertical transform.

[0164] For example, when MTS is applied to the current block (i.e., cu_mts_flag=1), the decoding apparatus 200 may configure an MTS candidate based on the intra prediction mode of the current block. In this case, the decoding flowchart of Fig. 10 may further include a step of configuring an MTS candidate. Then, the decoding apparatus 200 may determine an MTS candidate to be applied to the current block from among the configured MTS candidates using mts_idx.

[0165] As another example, different transform kernels may be applied to the horizontal transform and the vertical transform, although the present invention is not limited thereto, and the same transform kernel may be applied to the horizontal transform and the vertical transform.

[0166] Then, the decoding device 200 can perform an inverse transform based on the transform kernel (S1025).

[0167] In addition, in this document, MTS can also be expressed as AMT or EMT, and similarly, mts_idx, AMT_idx, EMT_idx, AMT_TU_idx, EMT_TU_idx, etc., and the present invention is not limited to such expressions.

[0168] In addition, in the present invention, a case where MTS is applied and a case where MTS is not applied are described based on the MTS flag, but the present invention is not limited to such an expression. For example, whether or not MTS is applied may mean whether or not to use another transform type (or transform kernel) other than a predefined specific transform type (which may be called a base transform type, a default transform type, etc.). If MTS is applied, another transform type (e.g., any one of a plurality of transform types or a combination of two or more transform types) other than the base transform type is used for the transform, and if MTS is not applied, the base transform type may be used for the transform. In one embodiment, the base transform type may be set (or defined) to DCT2.

[0169] As one example, an MTS flag syntax indicating whether MTS is applied to the current transform block and an MTS index syntax indicating the transform type applied to the current block if MTS is applied may be transmitted separately from the encoder to the decoder, and as another example, a syntax (e.g., an MTS index) including both whether MTS is applied to the current transform block and the transform type applied to the current block if MTS is applied may be transmitted from the encoder to the decoder. That is, in the latter embodiment, a syntax (or syntax element) indicating the transform type applied to the current transform block (or unit) within the entire transform type group (or transform type set) including the above-mentioned basic transform type may be transmitted from the encoder to the decoder.

[0170] Therefore, despite this expression, the syntax (MTS index) indicating the transform type applied to the current transform block can include information on whether or not to apply MTS. In other words, since only the MTS index is signaled without the MTS flag in the latter embodiment, in this case, it can be interpreted that the MTS includes DCT2, but in the present invention, the case where DCT2 is applied may be described as the case where MTS is not applied, and the technical scope of MTS is not limited to the definition.

[0171] FIG. 11 is a block diagram of an apparatus for performing MTS-based decoding according to an embodiment of the present invention.

[0172] The decoding device 200 to which the present invention is applied may include a sequence parameter acquisition unit 1105 , an MTS flag acquisition unit 1110 , an MTS index acquisition unit 1115 , and a transformation kernel derivation unit 1120 .

[0173] The sequence parameter acquisition unit 1105 may acquire sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag. Here, sps_mts_intra_enabled_flag indicates whether cu_mts_flag exists in the residual coding syntax of an intra coding unit, and sps_mts_inter_enabled_flag indicates whether cu_mts_flag exists in the residual coding syntax of an inter coding unit. For a specific example, the description related to FIG. 10 may be applied.

[0174] The MTS flag acquisition unit 1110 may acquire cu_mts_flag based on sps_mts_intra_enabled_flag or sps_mts_inter_enabled_flag. For example, when sps_mts_intra_enabled_flag=1 or sps_mts_inter_enabled_flag=1, the MTS flag acquisition unit 1110 may acquire cu_mts_flag. Here, cu_mts_flag indicates whether MTS is applied to the residual sample of the luma transform block. For a specific example, the description related to FIG. 10 may be applied.

[0175] The MTS index acquirer 1115 may acquire mts_idx based on cu_mts_flag. For example, when cu_mts_flag=1, the MTS index acquirer 1115 may acquire mts_idx. Here, mts_idx indicates which transform kernel is applied to luma residual samples along the horizontal and / or vertical directions of the current transform block. For a specific example, the description of FIG. 10 may be applied.

[0176] The transform kernel deriving unit 1120 can derive a transform kernel corresponding to mts_idx, and the decoding device 200 can perform an inverse transform based on the induced transform kernel.

[0177] A mode-dependent non-separable secondary transform (MDNSST) is introduced. To keep the complexity low, MDNSST is applied only to the low frequency coefficients after the primary transform. A non-separable transform that is applied mainly to the low frequency coefficients can be called a low frequency non-separable transform (LFNST). If the width (W) and height (H) of the transform coefficient block are both equal to or greater than 8, then an 8x8 non-separable secondary transform is applied to the upper left 8x8 region of the transform coefficient block. Otherwise, if the width or height is less than 8, then a 4x4 non-separable secondary transform is applied, which can be performed on the upper left min(8,W)xmin(8,H) region of the transform coefficient block. Here, min(A,B) is a function that outputs the smaller of A and B. Also, WxH is the size of the block, W is the width, and H is the height.

[0178] There may be a total of 35x3 non-separable quadratic transforms for 4x4 and 8x8 block sizes, where 35 is the number of transform sets specified by the intra prediction mode, and 3 is the number of NSST candidates for each prediction mode. The mapping from intra prediction mode to transform sets may be defined as shown in Table 5 below. Also, according to an embodiment of the present invention, among the four non-separable transform sets, depending on the intra prediction mode.

[0179] [Table 5]

[0180] An NSST index (NSST idx) may be coded to indicate a transform kernel within a transform set. If NSST is not applied, an NSST index with a value of 0 may be signaled.

[0181] 12 and 13 are encoding / decoding flowcharts in which a secondary transformation is applied as an embodiment to which the present invention is applied.

[0182] In JEM, secondary transform (MDNSST) is not applied to blocks coded in transform skip mode. If the MDNSST index is signaled to a CU and is non-zero, MDNSST is not used for blocks of components coded in transform skip mode within the CU. The overall coding structure including coefficient encoding and NSST index coding is shown in Figures 12 and 13. A coded block flag (CBF) is encoded to determine whether to perform coefficient encoding and NSST coding. In Figures 12 and 13, the CBF flag may indicate a luma block cbf flag (cbf_luma flag) or a chroma block cbf flag (cbf_cb flag or cbf_cr flag). The transform coefficients when the CBF flag is 1 are coded.

[0183] Referring to FIG. 12, the encoding apparatus 100 checks whether the CBF is 1 (S1205). If the CBF is 0, the encoding apparatus 100 does not perform encoding of the transform coefficients and encoding of the NSST index. If the CBF is 1, the encoding apparatus 100 performs encoding of the transform coefficients (S1210). Thereafter, the encoding apparatus 100 determines whether to perform NSST index coding (S1215) and performs NSST index coding (S1220). If NSST index coding is not applied, the encoding apparatus 100 may end the transform procedure in a state where NSST is not applied, and perform a subsequent step (e.g., quantization).

[0184] Referring to FIG. 13, the decoding apparatus 200 checks whether the CBF is 1 (S1305). If the CBF is 0, the decoding apparatus 200 does not perform the decoding of the transform coefficients and the NSST index decoding. If the CBF is 1, the decoding apparatus 200 performs the decoding of the transform coefficients (S1310). Thereafter, the decoding apparatus 200 determines whether to perform NSST index coding (S1315) and analyzes the NSST index (S1320).

[0185] NSST is not applied to the entire block (TU in the case of HEVC) to which the primary transform is applied, but may be applied to the upper left 8x8 region or 4x4 region. For example, if the size of the block is 8x8 or more, 8x8 NSST may be applied, and if it is less than 8x8, 4x4 NSST may be applied. In addition, when 8x8 NSST is applied, 4x4 NSST may be applied for each 4x4 block. Both 8x8 NSST and 4x4 NSST may be determined according to the configuration of the transform set described above, and as they are non-separable transforms, 8x8 NSST may have 64 input data and 64 output data, and 4x4 NSST may have 16 inputs and 16 outputs.

[0186] 14 and 15 are embodiments to which the present invention is applied. FIG. 14 shows a diagram for explaining Givens rotation, and FIG. 15 shows the configuration of one round in 4x4 NSST composed of a Givens rotation layer and permutation.

[0187] Both the 8x8 NSST and the 4x4 NSST can be constructed as a hierarchical combination of Givens rotations. The matrix corresponding to one Givens rotation is the same as Equation 1, and the matrix multiplication is shown in Figure 14.

[0188]

number

[0189] In FIG. 14, tm and tn output by the Givens rotation can be calculated as shown in Equation 2.

[0190]

number

[0191] As shown in FIG. 14, one Givens rotation rotates two data, so 32 or 8 Givens rotations are required to process 64 data (in the case of 8x8 NSST) or 16 data (in the case of 4x4 NSST), respectively. Therefore, a bundle of 32 or 8 Givens rotations can form a Givens rotation layer. As shown in FIG. 15, the output data of one Givens rotation layer is transferred to the input data of the next Givens rotation layer through permutation (shuffle). As shown in FIG. 15, the permutation pattern is regularly defined, and in the case of 4x4 NSST, four Givens rotation layers and corresponding permutations form one round. 4x4 NSST is performed in two rounds, and 8x8 NSST is performed in four rounds. Different rounds use the same permutation pattern, but the applied Givens rotation angles are different. Therefore, it is necessary to store the angle data of all Givens rotations that make up each transformation.

[0192] In the last step, a final permutation is performed on the data output through the Givens rotation layer, and the information of the permutation is stored separately for each transformation. The permutation is performed at the end of the forward NSST, and the inverse permutation is applied first to the inverse NSST.

[0193] Reverse NSST rotates the Givens rotation layers and permutations applied in the forward NSST in the reverse order, and also takes negative (-) values ​​for the angles of each Givens rotation.

[0194] RST (Reduced secondary transform)

[0195] FIG. 16 shows the operation of RST as an embodiment to which the present invention is applied.

[0196] Assuming that an orthogonal matrix representing a single transformation has an NxN form, RT (reduced transform) leaves only R out of N transformation basis vectors (R < N). The matrix of the forward RT that generates the transform coefficients can be defined as in Mathematical Formula 3.

[0197] [Number]

[0198] Since the matrix of the inverse RT is the transpose matrix of the forward RT matrix, if the application of the forward RT and the inverse RT is illustrated, it can be the same as FIGS. 16a and 16b.

[0199] The RT applied to the upper left 8x8 block of the block of transform coefficients to which the first transform is applied may be called an 8x8 RST. In Equation 3, when the value of R is set to 16, the forward 8x8 RST has a 16x64 matrix form, and the backward 8x8 RST has a 64x16 form. In addition, the configuration of the transform set as shown in Table 5 may be applied to the 8x8 RST. That is, the 8x8 RST may be determined based on the transform set according to the intra prediction mode as shown in Table 5. Since one transform set is composed of two or three transforms according to the intra prediction mode, one of up to four transforms may be selected, including the case where the second transform is not applied (one transform may correspond to an identity matrix). When the four transforms are assigned indexes of 0, 1, 2, and 3, respectively, a syntax element corresponding to the NSST index may be signaled for each block of transform coefficients, thereby specifying the transform to be applied. For example, the 0th index can be assigned to the identity matrix, i.e., when no quadratic transformation is applied. In conclusion, for an 8x8 top-left block via an NSST index, an 8x8 NSST can be specified according to the JEM NSST, and an 8x8 RST can be specified according to the RST configuration.

[0200] FIG. 17 is a diagram illustrating a process of performing a reverse scan from the 64th to the 17th blocks based on a reverse scan order, as an embodiment to which the present invention is applied.

[0201] When the forward 8x8 RST as shown in Equation 3 is applied, 16 valid transform coefficients are generated, so that 64 input data constituting an 8x8 area is reduced to 16 output data, and from the perspective of a two-dimensional area, only about 1 / 4 of the area is filled with valid transform coefficients. Therefore, by applying the forward 8x8 RST, the 16 output data obtained are filled in the upper left area of ​​FIG. 17.

[0202] In FIG. 17, the 4x4 region at the top left corner is a ROI (region of interest) region filled with valid transform coefficients, and the remaining regions are empty. The empty regions may be filled with a value of 0 by default. If a valid non-zero transform coefficient is found outside the ROI region of FIG. 17, it is certain that the 8x8 RST is not applied, and therefore the coding of the NSST index may be omitted. Conversely, if a non-zero transform coefficient is not found outside the ROI region of FIG. 17 (when the 8x8 RST is applied and the region outside the ROI is filled with 0), the NSST index may be coded since the 8x8 RST may have been applied. Such conditional NSST index coding may be performed after the residual coding process since it is necessary to check whether or not a non-zero transform coefficient exists.

[0203] FIG. 18 shows an example of an encoding flowchart using a single transform indicator as an embodiment to which the present invention is applied.

[0204] In an embodiment of the present invention, a single transform indicator (STI) is introduced. Instead of using two transforms (primary and secondary) sequentially, a single transform can be applied when the single transform indicator is activated (STI coding == 1). Here, the single transform can be any kind of transform. For example, the single transform can be a separable transform or a non-separable transform. The single transform can be a transform approximated from a non-separable transform. A single transform index (ST_idx in FIG. 18) can be signaled when the single transform indicator is activated. Here, the single transform index can indicate a transform corresponding to the transform to be applied from among available transform candidates.

[0205] Referring to FIG. 18, the encoding apparatus 100 determines whether the CBF is 1 (S1805). If the CBF is 1, the encoding apparatus 100 determines whether STI coding is applied (S1810). If the STI coding is applied, the encoding apparatus 100 encodes an STI index (STI_Idx) (S1845) and performs coding of transform coefficients (S1850). If the STI coding is not applied, the encoding apparatus 100 encodes a flag (EMT_CU_Flag) indicating whether EMT (or MTS) is applied at the CU level (S1815). Thereafter, the encoding apparatus 100 performs coding of transform coefficients (S1820). Thereafter, the encoding apparatus 100 determines whether EMT is applied to a transform unit (TU) (S1825). If EMT is applied to the TU, the encoding apparatus 100 encodes an index (EMT_TU Idx) of a primary transform applied to the TU (S1830). Thereafter, encoding apparatus 100 determines whether an NSST is applied (S1835). If an NSST is applied, encoding apparatus 100 encodes an index (NSST_Idx) indicating the NSST to be applied (S1840).

[0206] In one example, when a condition for single transform coding is satisfied / activated (e.g., STI_coding == 1), a single transform index (ST_Idx) is not signaled and may be derived implicitly. ST_idx may be implicitly determined based on the block size and intra prediction mode. Here, ST_idx may indicate the transform (or transform kernel) to be applied to the current transform block.

[0207] A single transformation indicator can be activated (STI_coding == 1) if one or more of the following conditions are met:

[0208] 1) The block size corresponds to a predetermined value, such as 4 or 8.

[0209] 2) Block width == Block height (square block)

[0210] 3) The intra prediction mode is one of the pre-determined modes such as DC or Planar.

[0211] In another example, the STI coding flag may be signaled to indicate whether a single transform is applied. The STI coding flag may be signaled based on the STI coding value and the CBF. For example, the STI coding flag may be signaled when the CBF is 1 and STI coding is activated. Furthermore, the STI coding flag may be conditionally signaled taking into account the block size, the block shape (square block or non-square block), or the intra prediction mode.

[0212] Since the acquired information is used in the coefficient coding, ST_idx can be determined after the coefficient coding. In one example, ST_idx can be implicitly determined based on the block size, intra prediction mode, and the number of non-zero coefficients. In another example, ST_idx can be conditionally encoded / decoded based on the block size and / or the block shape and / or the intra prediction mode, and / or the number of non-zero coefficients. In another example, ST_idx signaling can be omitted depending on the distribution of non-zero coefficients (i.e., the location of the non-zero coefficients). In particular, if a non-zero coefficient is found in an area other than the upper left 4x4 area, the signaling of ST_idx can be omitted.

[0213] FIG. 19 shows an example of an encoding flowchart using a unified transform indicator (UTI) as an embodiment to which the present invention is applied.

[0214] In the embodiment of the present invention, a unified conversion indicator is introduced. The UTI includes a primary conversion indicator and a secondary conversion indicator.

[0215] Referring to FIG. 19, the encoding apparatus 100 determines whether the CBF is 1 (S1905). If the CBF is 1, the encoding apparatus 100 determines whether UTI coding is applied (S1910). If UTI coding is applied, the encoding apparatus 100 encodes a UTI index (UTI_Idx) (S1945) and performs coding of transform coefficients (S1950). If UTI coding is not applied, the encoding apparatus 100 encodes a flag (EMT_CU_Flag) indicating whether EMT (or MTS) is applied at the CU level (S1915). Thereafter, the encoding apparatus 100 performs encoding of transform coefficients (S1920). Thereafter, the encoding apparatus 100 determines whether EMT is applied to a transform unit (TU) (S1925). If EMT is applied to the TU, encoding apparatus 100 encodes an index (EMT_TU Idx) of a primary transform applied to the TU (S1930). Then, encoding apparatus 100 determines whether NSST is applied (S1935). If NSST is applied, encoding apparatus 100 encodes an index (NSST_Idx) indicating the NSST to be applied (S1940).

[0216] The UTI can be encoded for each predetermined unit (CTU or CU).

[0217] The UTI coding mode may depend on the following conditions:

[0218] 1) Block Size

[0219] 2) Block shape

[0220] 3) Intra Prediction Mode

[0221] How to derive / extract the core transform index from the UTI is predefined. How to derive / extract the secondary transform index from the UTI is predefined.

[0222] The syntax structure of the UTI is selectively used. The UTI may depend on the CU (or TU) size. For example, smaller CUs (TUs) may have a relatively narrow range of UTI indices. In one example, if a predefined condition is met (e.g., the block size is smaller than a predefined threshold), the UTI may only point to the core transform index.

[0223] [Table 6]

[0224] In another example, if no secondary transform is indicated to be used (e.g., secondary transform index == 0 or secondary transform has already been determined), the UTI index can be treated as a core transform index. Similarly, if the core transform index is known, the UTI index can be treated as a secondary transform index. In particular, a pre-determined core transform can be used, taking into account the intra prediction mode and block size.

[0225] 20a and 20b show another example of an encoding flowchart using UTI as an embodiment to which the present invention is applied.

[0226] In another example, the transform encoding structure uses UTI index coding as shown in Figures 20a and 20b, where the UTI index can be encoded before or after the coefficient encoding.

[0227] Referring to FIG. 20a, the encoding apparatus 100 checks whether the CBF is 1 (S2005). If the CBF is 1, the encoding apparatus 100 codes a UTI index (UTI_Idx) (S2010) and codes transform coefficients (S2015).

[0228] Referring to FIG. 20b, the encoding apparatus 100 checks whether the CBF is 1 (S2055). If the CBF is 1, the encoding apparatus 100 performs coding of the transform coefficients (S2060) and codes the UTI index (UTI_Idx) (S2065).

[0229] In another embodiment of the present invention, data hiding and implicit coding methods of transform indicator are introduced. Here, the transform indicator includes ST_idx, UTI_idx, EMT_CU_Flag, EMT_TU_Flag, NSST_idx, and a transform related index used to indicate the transform kernel. The above-mentioned transform indicator is not signaled, and the information can be inserted into the coefficient encoding process (can be extracted during the coefficient coding process). The coefficient encoding process can include the following parts:

[0230] - Last x position (Last_position_x), Last y position (Last_position_y)

[0231] - Group flag

[0232] - Significance map

[0233] - Flag indicating whether it is greater than 1 (Greater_than_1_flag)

[0234] - Flag indicating whether it is greater than 2 (Greater_than_2_flag)

[0235] - Remaining level coding

[0236] - Sign coding

[0237] For example, the transformation indicator information can be inserted into one or more of the coefficient coding processes described above. The following can be considered together for inserting the transformation indicator information:

[0238] - Pattern of Sign coding

[0239] - The absolute value of remaining level

[0240] The number of Greater_than_1_flag indicating whether it is greater than -1

[0241] - The value of Last_position_X and Last_position_Y

[0242] The above mentioned data hiding methods can be considered conditional. For example, the data hiding method can depend on the number of non-zero coefficients.

[0243] In yet another example, NSST_idx and EMT_idx can be dependent, for example, when EMT_CU_flag is 0 (or 1), NSST_idx may not be 0. In this case, NSST_idx-1 can be signaled instead of NSST_idx.

[0244] In another embodiment of the present invention, the mapping of NSST transform sets based on intra prediction modes is presented as shown in Table 7 below. As mentioned above, in the following description, NSST will be mainly described as an example of a non-separable transform, but other known terms for non-separable transforms (e.g., LFNST) may be used. For example, NSST set and NSST index are used in place of LFNST set and LFNST index. In addition, RST described in this document is used in place of RST or LFNST as an example of a non-separable transform (e.g., LFNST) using a non-square transform matrix having a reduced input length and / or a reduced output length with a square non-separable transform matrix applied to at least a part of a transform block (a 4x4 or 8x8 region on the upper left side or a remaining region excluding a 4x4 region on the lower right side in an 8x8 block).

[0245] [Table 7]

[0246] The NSST set numbers can be rearranged between 0 and 3 as shown in Table 8.

[0247] [Table 8]

[0248] In the NSST transform set, four transform sets are used (instead of 35) to reduce the memory space required.

[0249] Furthermore, for each transformation set, a different number of transformation kernels are used as follows:

[0250] Case A: Two available transform kernels are used for each transform set, and the NSST index ranges from 0 to 2. For example, if the NSST index is 0, a secondary transform (secondary inverse transform based on the decoder) may not be applied. If the NSST index is 1 or 2, a secondary transform may be applied. A transform set may include two transform kernels, and an index of 1 or 2 may be mapped to the two transform kernels.

[0251] [Table 9]

[0252] Referring to Table 9, two transform kernels are used for each of the 0th to 3rd non-separable transform (NSST or LFNST) sets.

[0253] Case B: Use two available transform kernels for transform set 0, and one each for the remaining transform sets. The available NSST indices for transform set 0 (DC, Planner) are 0 to 2. However, the NSST indices for other modes (transform sets 1, 2, 3) are 0 to 1.

[0254] [Table 10]

[0255] Referring to Table 10, two non-separable transform kernels are set for the non-separable transform (NSST) set corresponding to the 0th index, and one non-separable transform kernel is set for each of the non-separable transform (NSST) sets corresponding to the 1st, 2nd, and 3rd indexes.

[0256] Case C: Use one transform kernel for each transform set, and the NSST index range is 0 to 1.

[0257] [Table 11]

[0258] FIG. 21 shows an example of an encoding flowchart for performing conversion as an embodiment to which the present invention is applied.

[0259] The encoding apparatus 100 performs a primary transform on the residual block (S2105). The primary transform may be referred to as a core transform. In an embodiment, the encoding apparatus 100 may perform the primary transform using the above-mentioned MTS. The encoding apparatus 100 may also transmit an MTS index indicating a specific MTS from among MTS candidates to the decoding apparatus 200. In this case, the MTS candidates may be configured based on an intra prediction mode of the current block.

[0260] The encoding device 100 determines whether to apply a secondary transform (S2110). As an example, the encoding device 100 may determine whether to apply a secondary transform based on the linearly transformed residual transform coefficients. For example, the secondary transform may be NSST or RST.

[0261] The encoding device 100 determines the secondary transform (S2115). At this time, the encoding device 100 can determine the secondary transform based on the NSST (or RST) transform set specified according to the intra prediction mode.

[0262] Also, for example, prior to operation S2115, the encoding apparatus 100 may determine an area to which the secondary transformation is applied based on the size of the current block.

[0263] The encoding apparatus 100 performs the secondary transformation using the secondary transformation determined in operation S2115 (S2120).

[0264] FIG. 22 shows an example of a decoding flowchart for performing conversion as an embodiment to which the present invention is applied.

[0265] The decoding device 200 determines whether to apply the secondary inverse transform (S2205). For example, the secondary inverse transform may be NSST or RST. As an example, the decoding device 200 may determine whether to apply the secondary inverse transform based on the secondary transform flag received from the encoding device 100.

[0266] The decoding apparatus 200 determines a secondary inverse transform (S2210). At this time, the decoding apparatus 200 may determine a secondary inverse transform to be applied to the current block based on the NSST (or RST) transform set designated according to the above-described intra prediction mode.

[0267] Also, for example, the decoding apparatus 200 may determine an area to which the secondary inverse transform is applied based on the size of the current block prior to operation S2210.

[0268] The decoding apparatus 200 performs a secondary inverse transform on the dequantized residual block using the secondary inverse transform determined in step S2210 (S2215).

[0269] The decoding apparatus 200 performs a primary inverse transform on the secondary inverse transformed residual block (S2220). The primary inverse transform may be referred to as a core inverse transform. As an embodiment, the decoding apparatus 200 may perform the primary inverse transform using the above-mentioned MTS. As an example, the decoding apparatus 200 may determine whether MTS is applied to the current block prior to step S2220. In this case, the decoding flowchart of FIG. 22 may further include a step of determining whether MTS is applied.

[0270] As an example, when MTS is applied to the current block (i.e., cu_mts_flag=1), the decoding apparatus 200 may configure an MTS candidate based on the intra prediction mode of the current block. In this case, the decoding flowchart of Fig. 22 may further include a step of configuring an MTS candidate. Then, the decoding apparatus 200 may determine a first inverse transform to be applied to the current block using mts_idx indicating a specific MTS among the configured MTS candidates.

[0271] FIG. 23 shows an example of a detailed block diagram of the conversion unit 120 in the encoding device 100 as an embodiment to which the present invention is applied.

[0272] The encoding device 100 to which the embodiment of the present invention is applied may include a primary transform unit 2310, a secondary transform application determining unit 2320, a secondary transform determining unit 2330, and a secondary transform unit 2340.

[0273] The primary transform unit 2310 may perform a primary transform on the residual block. The primary transform may be referred to as a core transform. As an embodiment, the primary transform unit 2310 may perform the primary transform using the above-mentioned MTS. Also, the primary transform unit 2310 may transmit an MTS index indicating a specific MTS from among MTS candidates to the decoding apparatus 200. In this case, the MTS candidates may be configured based on an intra prediction mode of the current block.

[0274] The secondary transform application determining unit 2320 may determine whether to apply a secondary transform. For example, the secondary transform application determining unit 2320 may determine whether to apply a secondary transform based on a transform coefficient of the primary transformed residual block. For example, the secondary transform may be NSST or RST.

[0275] The secondary transform determination unit 2330 determines the secondary transform. At this time, the secondary transform determination unit 2330 may determine the secondary transform based on the NSST (or RST) transform set designated according to the intra prediction mode, as described above.

[0276] Also, for example, the secondary transformation determination unit 2330 may determine an area to which the secondary transformation is applied based on the size of the current block.

[0277] The secondary transform unit 2340 may perform the secondary transform using the determined secondary transform.

[0278] FIG. 24 shows an example of a detailed block diagram of the inverse transform unit 230 in the decoding device 200 as an embodiment to which the present invention is applied.

[0279] The decoding device 200 to which the present invention is applied includes a secondary inverse transform application determining unit 2410, a secondary inverse transform determining unit 2420, a secondary inverse transform unit 2430, and a primary inverse transform unit 2440.

[0280] The secondary inverse transform application determining unit 2410 may determine whether to apply the secondary inverse transform. For example, the secondary inverse transform may be NSST or RST. As an example, the secondary inverse transform application determining unit 2410 may determine whether to apply the secondary inverse transform based on a secondary transform flag received from the encoding apparatus 100. As another example, the secondary inverse transform application determining unit 2410 may determine whether to apply the secondary inverse transform based on a transform coefficient of a residual block.

[0281] The secondary inverse transform determination unit 2420 may determine a secondary inverse transform. In this case, the secondary inverse transform determination unit 2420 may determine a secondary inverse transform to be applied to a current block based on an NSST (or RST) transform set designated according to an intra prediction mode.

[0282] Also, for example, the secondary inverse transform determination unit 2420 may determine an area to which the secondary inverse transform is applied based on the size of the current block.

[0283] Also, as an example, the secondary inverse transform unit 2430 may perform a secondary inverse transform on the dequantized residual block using the determined secondary inverse transform.

[0284] The first inverse transform unit 2440 may perform a first inverse transform on the second inverse transformed residual block. As an embodiment, the first inverse transform unit 2440 may perform a first transform using the above-mentioned MTS. As an example, the first inverse transform unit 2440 may determine whether MTS is applied to the current block.

[0285] For example, when MTS is applied to the current block (i.e., cu_mts_flag=1), the primary inverse transform unit 2440 may configure MTS candidates based on the intra prediction mode of the current block, and may determine a primary transform to be applied to the current block using mts_idx indicating a specific MTS from among the configured MTS candidates.

[0286] 25 shows a flowchart for processing a video signal as an embodiment to which the present invention is applied. The flowchart in FIG. 25 can be executed by the decoding device 200 or the inverse transform unit 230.

[0287] First, the decoding apparatus 200 may determine whether to apply an inverse non-separable transform of the current block based on the non-separable transform index and the width and height of the current block. For example, the decoding apparatus 200 may determine to apply the non-separable transform when the non-separable transform index is not 0 and the width and height of the current block are each equal to or greater than 4. If the non-separable transform index is 0 or the width or height of the current block is less than 4, the decoding apparatus 200 may omit the inverse non-separable transform and perform an inverse linear transform.

[0288] In operation S2505, the decoding apparatus 200 determines a non-separable transform set index indicating a non-separable transform set used for non-separable transform of the current block from among predefined non-separable transform sets based on the intra prediction mode of the current block. The non-separable transform set index may be set to be assigned to each of four transform sets set according to the range of the intra prediction mode as shown in Table 7 or Table 8. That is, as shown in Table 7 or Table 8, if the intra prediction mode is 0 to 1, the non-separable transform set index may be determined to a first index value, if the intra prediction mode is 2 to 12 or 56 to 66, the non-separable transform set index may be determined to a second index value, if the intra prediction mode is 13 to 23 or 45 to 55, the non-separable transform set index may be determined to a third index value, and if the intra prediction mode is 24 to 44, the non-separable transform set index may be determined to a fourth index value.

[0289] Here, each of the predefined non-separable transformation sets may include two transformation kernels, as shown in Table 9. Also, each of the predefined non-separable transformation sets may include one or two transformation kernels, as shown in Table 10 or Table 11.

[0290] In operation S2510, the decoding apparatus 200 determines a transform kernel indicated by a non-separable transform index of a current block among transform kernels included in a non-separable transform set indicated by a non-separable transform set index, using a non-separable transform matrix. For example, two non-separable transform kernels may be set for each index value of the non-separable transform set index, and the decoding apparatus 200 may determine a non-separable transform matrix based on a transform kernel indicated by a non-separable transform index among two transform matrix kernels corresponding to the non-separable transform set index.

[0291] In operation S2515, the decoding apparatus 200 applies a non-separable transform matrix to an upper left region of the current block, which is determined according to the width and height of the current block. For example, if both the width and height of the current block are equal to or greater than 8, a non-separable transform may be applied to an 8x8 region at the upper left of the current block, and if the width or height of the current block is less than 8, a non-separable transform may be applied to a 4x4 region of the current block. The size of the non-separable transform may also be set to 8x8 or 4x4 depending on the region to which the non-separable transform is applied.

[0292] In addition, the decoding apparatus 200 may apply a horizontal transform and a vertical transform to the current block to which the non-separable transform has been applied, where the horizontal transform and the vertical transform may be determined based on a prediction mode applied to the current block and an MTS index for selecting a transform matrix.

[0293] Hereinafter, a method of combining and applying a primary transform and a secondary transform will be described. That is, in an embodiment of the present invention, a method of efficiently designing transforms used in the primary transform and the secondary transform is proposed. Here, the methods proposed in FIG. 1 to FIG. 25 may be applied, and the overlapping description will be omitted.

[0294] As described above, the primary transform refers to a transform that is first applied to a residual block with reference to an encoder. When a secondary transform is applied, the encoder performs a secondary transform on the primary transformed residual block. On the other hand, when a secondary transform is applied, a secondary inverse transform is performed prior to the primary inverse transform with reference to a decoder. The decoder can derive a residual block by performing a primary inverse transform on a secondary inverse transformed transform coefficient block.

[0295] Also, as described above, a non-separable transform may be used as the secondary transform, and may be applied only to low-frequency coefficients in a specific region on the upper left side to maintain low complexity. Such a secondary transform applied to low-frequency coefficients may be referred to as a non-separable secondary transform (NSST), low frequency non-separable transform (LFNST), or reduced secondary transform (RST). Also, the primary transform may be referred to as a core transform.

[0296] In one embodiment of the present invention, the primary transform candidates used in the primary transform and the secondary transform kernels used in the secondary transform may be predefined in various combinations. In this specification, the primary transform candidates used in the primary transform may be called MTS candidates, but are not limited to this name. As an example, the primary transform candidates may be a combination of transform kernels (or transform types) applied to the horizontal and vertical directions, respectively, and the transform kernels may be any one of DCT2, DST7, and / or DCT8. In other words, the primary transform candidates may be at least one combination of DCT2, DST7, and / or DCT8. A specific example will be described below.

[0297] -Combination A

[0298] In combination A, the primary transform candidates and the secondary transform kernels are defined according to the intra prediction mode, as shown in Table 12 below.

[0299] [Table 12]

[0300] Referring to Table 12, as an example (Case 1), when the intra prediction mode has directionality, two primary transform candidates are used, and when it does not have directionality (e.g., DC, planar mode), four primary transform candidates are used. Here, the secondary transform candidates may include two transform kernels regardless of the directionality of the intra prediction mode. That is, as described above, a plurality of secondary transform kernel sets are predefined according to the intra prediction mode, and each of the predefined plurality of secondary transform kernel sets includes two transform kernels.

[0301] Also, as an example (Case 2), when the intra prediction mode has directionality, two primary transform candidates are used, and when it does not have directionality, four primary transform candidates are used. Here, the secondary transform candidates may include one transform kernel when the intra prediction mode has directionality, and may include two transform kernels when it does not have directionality.

[0302] Also, as an example (Case 3), if the intra prediction mode has directionality, two primary transform candidates are used, and if it does not have directionality, four primary transform candidates are used, where the secondary transform candidates may include one transform kernel regardless of the directionality of the intra prediction mode.

[0303] -Combination B

[0304] In combination B, the primary transform candidates and the secondary transform kernels are defined according to the intra prediction mode, as shown in Table 13 below.

[0305] [Table 13]

[0306] Referring to Table 13, as an example (Case 1), three primary transform candidates are used regardless of the direction of the intra prediction mode. Here, the secondary transform candidates may include two transform kernels regardless of the direction of the intra prediction mode. That is, as described above, a plurality of secondary transform kernel sets are predefined according to the intra prediction mode, and each of the predefined plurality of secondary transform kernel sets may include two transform kernels.

[0307] Also, as an example (Case 2), three primary transform candidates are used regardless of the directionality of the intra prediction mode, where the secondary transform candidates include one transform kernel if the intra prediction mode has directionality, and two transform kernels if the intra prediction mode does not have directionality.

[0308] Also, as an example (Case 3), three primary transform candidates are used regardless of the directionality of the intra prediction mode, and the secondary transform candidates may include one transform kernel regardless of the directionality of the intra prediction mode.

[0309] -Combination C

[0310] In combination C, the primary transform candidates and the secondary transform kernels are defined according to the intra prediction mode, as shown in Table 14 below.

[0311] [Table 14]

[0312] Referring to Table 14, as an example (Case 1), when the intra prediction mode has directionality, two primary transform candidates are used, and when it does not have directionality (e.g., DC, planar mode), three primary transform candidates are used. Here, the secondary transform candidates may include two transform kernels regardless of the directionality of the intra prediction mode. That is, as described above, a plurality of secondary transform kernel sets are predefined according to the intra prediction mode, and each of the predefined plurality of secondary transform kernel sets may include two transform kernels.

[0313] Also, as an example (Case 2), when the intra prediction mode has directionality, two primary transform candidates are used, and when it does not have directionality, three primary transform candidates are used. Here, the secondary transform candidates may include one transform kernel when the intra prediction mode has directionality, and may include two transform kernels when it does not have directionality.

[0314] Also, as an example (Case 3), if the intra prediction mode has directionality, two primary transform candidates are used, and if it does not have directionality, three primary transform candidates are used, where the secondary transform candidates may include one transform kernel regardless of the directionality of the intra prediction mode.

[0315] The above description has focused on the case where a plurality of primary conversion candidates are used. Below, a description will be given of an example of a combination of primary conversion and secondary conversion when a fixed primary conversion candidate is used.

[0316] -Combination D

[0317] In combination D, the primary transform candidates and secondary transform kernels are defined according to the intra prediction mode, as shown in Table 15 below.

[0318] [Table 15]

[0319] Referring to Table 15, in an embodiment, one primary transform candidate is fixedly used regardless of the intra prediction mode. For example, the fixed primary transform candidate may be at least one combination of DCT2, DST7, and / or DCT8.

[0320] As an example (Case 1), one primary transform candidate is fixedly used regardless of the intra prediction mode, and the secondary transform candidate may include two transform kernels regardless of the direction of the intra prediction mode. That is, as described above, a plurality of secondary transform kernel sets are predefined according to the intra prediction mode, and each of the predefined plurality of secondary transform kernel sets may include two transform kernels.

[0321] As another example (Case 2), one primary transform candidate is fixedly used regardless of the intra prediction mode, where the secondary transform candidate may include one transform kernel if the intra prediction mode is directional and may include two transform kernels if the intra prediction mode is not directional.

[0322] Also, as an example (Case 3), one primary transform candidate may be fixedly used regardless of the intra prediction mode, where the secondary transform candidate may include one transform kernel regardless of the directionality of the intra prediction mode.

[0323] -Combination E

[0324] In combination E, the primary transform candidates and secondary transform kernels are defined according to the intra prediction mode, as shown in Table 16 below.

[0325] [Table 16]

[0326] Referring to Table 16, the secondary transform is defined only when DCT2 is applied as the primary transform. In other words, the secondary transform can be applied when MTS is not applied (i.e., when DCT2 is applied as the primary transform). As described in FIG. 10, in this specification, the case where MTS is applied and the case where it is not applied are described separately, but the present invention is not limited to such an expression. For example, whether or not MTS is applied may mean whether or not to use another transform type (or transform kernel) other than a predefined specific transform type (which may be called a base transform type, a default transform type, etc.). If MTS is applied, another transform type (e.g., any one or a combination of two or more transform types) other than the base transform type may be used for the transform, and if MTS is not applied, the base transform type may be used for the transform. In one embodiment, the base transform type may be set (or defined) to DCT2.

[0327] As an example (Case 1), when DCT2 is applied to the primary transform, a secondary transform can be applied, where the secondary transform candidates include two transform kernels regardless of the direction of the intra prediction mode. That is, as described above, a plurality of secondary transform kernel sets are predefined according to the intra prediction mode, and each of the predefined plurality of secondary transform kernel sets includes two transform kernels.

[0328] Also, as an example (Case 2), when DCT2 is applied to the primary transform, a secondary transform can be applied, where the secondary transform candidates may include one transform kernel if the intra prediction mode has directionality, and two transform kernels if it does not have directionality.

[0329] Also, as an example (Case 3), when DCT2 is applied to the primary transform, a secondary transform can be applied, where the secondary transform candidates can include one transform kernel regardless of the directionality of the intra prediction mode.

[0330] FIG. 26 is a flowchart illustrating a method for converting a video signal according to an embodiment of the present invention.

[0331] 26, for convenience of explanation, a decoder will be mainly described, but the present invention is not limited thereto, and the conversion method for a video signal according to the present embodiment can be substantially equally applied to an encoder. The flowchart of FIG. 26 is performed by the decoding device 200 or the inverse conversion unit 230.

[0332] The decoding apparatus 200 parses a first syntax element indicating a primary transform kernel to be applied to a primary transform of a current block (S2601).

[0333] The decoding apparatus 200 determines whether a secondary transform can be applied to the current block based on the first syntax element (S2602).

[0334] If a secondary transform can be applied to the current block, the decoding apparatus 200 parses a second syntax element indicating a secondary transform kernel to be applied to the secondary transform of the current block (S2603).

[0335] The decoding apparatus 200 performs a secondary inverse transform on a specific region at the upper left side of the current block using the secondary transform kernel indicated by the second syntax element, thereby deriving a secondary inverse transformed block (S2604).

[0336] The decoding apparatus 200 derives a residual block of the current block by performing a primary inverse transform on the secondary inverse transformed block using a primary transform kernel indicated by the first syntax element (S2605).

[0337] As described above, step S2602 is performed by determining that a secondary transform is applicable to the current block if the first syntax element indicates a predefined first transform kernel, where the first transform kernel is defined as DCT2.

[0338] Also, as described above, the decoding apparatus 200 may determine a secondary transform kernel set to be used for secondary transform of the current block from among predefined secondary transform kernel sets based on an intra prediction mode of the current block, and the second syntax element may indicate a secondary transform kernel to be applied to secondary transform of the current block from the determined secondary transform kernel set.

[0339] Also, as mentioned above, each of the predefined secondary transformation kernel sets may include two transformation kernels.

[0340] In one embodiment of the present invention, an example of a syntax structure in which a Multiple Transform Set (MTS) is used will be described.

[0341] As an example, Table 17 below shows an example of a syntax structure of a sequence parameter set.

[0342] [Table 17]

[0343] Referring to Table 17, whether or not MTS is enabled according to an embodiment of the present invention can be signaled via a sequence parameter set syntax. Here, sps_mts_intra_enabled_flag indicates whether or not an MTS flag or MTS index is present in a lower level syntax (e.g., residual coding syntax, transform unit syntax) for an intra coding unit. And, sps_mts_inter_enabled_flag indicates whether or not an MTS flag or MTS index is present in a lower level syntax for an inter coding unit.

[0344] As another example, Table 18 below shows an example of a transform unit syntax structure.

[0345] [Table 18]

[0346] Referring to Table 18, cu_mts_flag indicates whether MTS is applied to the residual samples of the luma transform block. For example, if cu_mts_flag=0, MTS is not applied to the residual samples of the luma transform block, and if cu_mts_flag=1, MTS is applied to the residual samples of the luma transform block.

[0347] As described above, in the present invention, the case where MTS is applied and the case where it is not applied are described based on the MTS flag, but the present invention is not limited to such an expression. For example, whether or not MTS is applied may mean whether or not a transform type (or a transform kernel) other than a predefined specific transform type (which may be called a base transform type, a default transform type, etc.) is used. If MTS is applied, a transform type other than the base transform type (e.g., any one of a plurality of transform types, or a combination of two or more transform types) may be used for the transform, and if MTS is not applied, the base transform type may be used for the transform. In one embodiment, the base transform type may be set (or defined) to DCT2.

[0348] As an example, an MTS flag syntax indicating whether MTS is applied to the current transform block and an MTS index syntax indicating the transform type applied to the current block if MTS is applied can be transmitted from the encoder to the decoder separately, and as another example, a syntax (e.g., an MTS index) including whether MTS is applied to the current transform block and, if MTS is applied, all the transform types applied to the current block can be transmitted from the encoder to the decoder. That is, in the latter embodiment, a syntax (or syntax element) indicating the transform type applied to the current transform block (or unit) within the entire transform type group (or transform type set) including the above-mentioned basic transform type can be transmitted from the encoder to the decoder.

[0349] Therefore, despite the expression, the syntax (MTS index) indicating the transform type applied to the current transform block may include information on whether MTS is applicable. In other words, in the latter embodiment, only the MTS index may be signaled without the MTS flag, and in this case, it may be interpreted that the MTS includes DCT2, but in the present invention, the case where DCT2 is applied may be described as the case where MTS is not applied, and the technical scope of MTS is not limited to the definition.

[0350] As another example, Table 19 below shows an example of a residual unit syntax structure.

[0351] [Table 19-1]

[0352] [Table 19-2]

[0353] [Table 19-3]

[0354] [Table 19-4]

[0355] [Table 19-5]

[0356] Referring to Table 19, the transform_skip_flag and / or mts_idx syntax (or syntax element) may be signaled via a residual syntax. However, this is only an example and the present invention is not limited thereto. For example, the transform_skip_flag and / or mts_idx syntax may be signaled via a transform unit syntax.

[0357] The following describes a specific embodiment of a secondary transform matrix that can be used for the above-mentioned secondary transform. As described above, the secondary transform can be called a non-separable secondary transform (NSST), a low frequency non-separable transform (LFNST), or a reduced secondary transform (RST).

[0358] As described above, in an embodiment of the present invention, four transform sets (or secondary transform sets) may be used to improve memory efficiency when applying secondary transforms. In one embodiment, the encoder / decoder may assign indices 0, 1, 2, and 3 to the four transform sets, respectively.

[0359] Also, as described above, each transform set may include a predefined number of transform kernels. In one embodiment, four transform sets used for the secondary transform may be predefined in the encoder and decoder, and each transform set may include one or two transform matrices (or transform types, transform kernels).

[0360] The following Table 20 shows an example of a transformation that can be applied to an 8x8 region.

[0361] [Table 20-1]

[0362] [Table 20-2]

[0363] [Table 20-3]

[0364] [Table 20-4]

[0365] [Table 20-5]

[0366] [Table 20-6]

[0367] [Table 20-7]

[0368] Referring to Table 20, an example is shown in which the coefficients of the transformation matrix are multiplied by a scaling value of 128. In Table 20, the first input [4] in the g_aiNsst8×8[4][2]

[16]

[64] array means the number of transformation sets (each transformation set can be divided into indexes 0, 1, 2, and 3), the second input [2] means the number of transformation matrices that make up each transformation set, and the third and fourth inputs

[16]

[64] mean the rows and columns of the 16×64 RST (Reduced Secondary Transform).

[0369] Table 20 assumes that the transform set includes two transform matrices, but if the transform set includes one transform matrix, it can be set to use a specific order of transform matrices for each transform set in Table 20. For example, if the transform set includes one transform matrix, the encoder / decoder can use a transform matrix predefined in each transform set in Table 20, i.e., the first or second transform matrix.

[0370] When applying the RST of Table 20, the encoder / decoder can be configured (or defined, set) to output 16 transform coefficients, or can be configured to output only m transform coefficients by applying only m×64 parts of the 16×64 matrix. For example, the encoder / decoder can be configured to output only 8 transform coefficients using only 8×64 matrices from the top by setting m=8. In this way, the amount of calculations can be reduced by half by applying a reduced quadratic transform. In one embodiment, the encoder / decoder can apply an 8×64 matrix to an 8×8 transform unit (TU) to reduce the amount of calculations in the worst case.

[0371] The following Table 21 shows an example of a transformation that can be applied to a 4x4 region.

[0372] [Table 21-1]

[0373] [Table 21-2]

[0374] [Table 21-3]

[0375] Referring to Table 21, an example is shown in which the coefficients of the transformation matrix are multiplied by a scaling value of 128. In Table 21, the [4] of the first input of the g_aiNsst4x4[4][2]

[16]

[16] array means the number of transformation sets (each transformation set can be divided into indexes 0, 1, 2, and 3), the [2] of the second input means the number of transformation matrices constituting each transformation set, and the

[16]

[16] of the third and fourth inputs means the rows and columns of a 16x16 RST (Reduced Secondary Transform).

[0376] Table 21 assumes that the transform set includes two transform matrices, but if the transform set includes one transform matrix, it can be set to use a specific order of transform matrices for each transform set in Table 21. For example, if the transform set includes one transform matrix, the encoder / decoder can use a transform matrix predefined in each transform set in Table 21, i.e., the first or second transform matrix.

[0377] When applying the RST of Table 21, the encoder / decoder can be configured (or defined, set) to output 16 transform coefficients, or can be configured to output only m transform coefficients by applying only m×16 parts of a 16×16 matrix. For example, the encoder / decoder can be configured to output only 8 transform coefficients using only an 8×16 matrix from the top by setting m=8. In this way, the amount of calculation can be reduced by half by applying a reduced quadratic transform. In one embodiment, the encoder / decoder can apply an 8×64 matrix to an 8×8 transform unit (TU) to reduce the amount of calculation in the worst case.

[0378] In one embodiment, the secondary transform may be applied to the top left 4x4, 4x8, or 8x4 region (i.e., TU) according to a predefined condition, or may be applied only to the top left 4x4 region. In the case of a 4x8 TU and an 8x4 TU, the encoder / decoder may divide the TU into two 4x4 regions and apply a specified transform to each divided region. If the secondary transform is defined to be applied only to the 4x4 region, only the transform defined in Table 21 may be applied (or used).

[0379] Meanwhile, in Tables 20 and 21, the coefficients of the transformation matrices are defined assuming that the scaling value is 128, but the present invention is not limited thereto. For example, Tables 20 and 21 can be defined as shown in Tables 22 and 23 below by setting the scaling value to 256.

[0380] [Table 22-1]

[0381] [Table 22-2]

[0382] [Table 22-3]

[0383] [Table 22-4]

[0384] [Table 22-5]

[0385] [Table 22-6]

[0386] [Table 22-7]

[0387] [Table 22-8]

[0388] [Table 23-1]

[0389] [Table 23-2]

[0390] [Table 23-3]

[0391] As described above, in an embodiment of the present invention, four transform sets (or secondary transform sets) may be used to improve memory efficiency when applying secondary transforms. In one embodiment, the encoder / decoder may assign indices 0, 1, 2, and 3 to the four transform sets, respectively.

[0392] Also, as described above, each transform set may include a predefined number of transform kernels. In one embodiment, four transform sets used for the secondary transform may be predefined in the encoder and decoder, and each transform set may include one or two transform matrices (or transform types, transform kernels).

[0393] Various examples of various quadratic transformation sets and transformation matrices (or transformation types, transformation kernels) applicable to quadratic transformation will be described below. Although various transformation matrices different from those in Tables 20 to 23 may be defined in detail, in this embodiment, for convenience of explanation, examples will be described focusing on non-directional modes (e.g., DC mode, planar mode) along with a generalized method of constructing a quadratic transformation set.

[0394] First, examples of secondary transformations that can be applied to a 4×4 region will be described in detail. Among the following examples of secondary transformation sets available for secondary transformation, the first and fourth examples can be applied to an embodiment in which the transformation matrices are each composed of two transformation matrices. The transformation matrices of the second and third examples can be applied to an embodiment in which the transformation sets are each composed of one transformation matrix.

[0395] In particular, the first example may be applied to the above-mentioned combination D and Case 1 of the embodiment described in Table 15, and may also be applied to combination A and Case 1 of the embodiment described in Table 12, combination B and Case 1 of the embodiment described in Table 13, combination C and Case 1 of the embodiment described in Table 14, or combination E and Case 1 of the embodiment described in Table 16.

[0396] In particular, the second example transformation array (i.e., transformation set) may be applied to Case 3 of the embodiment described in combination D and Table 15 above, and further to Case 3 of combination A and the embodiment described in Table 12, Case 3 of combination B and the embodiment described in Table 13, Case 3 of combination C and the embodiment described in Table 14, or Case 3 of combination E and the embodiment described in Table 16.

[0397] Although the above-mentioned combinations A, B, C, D, and E only deal with cases where the number of MTS candidates is three or less, it is also possible to configure the first transform to apply all four MTS candidates to all intra prediction modes. The following first to fourth examples can also be used when all four MTS candidates are applied, and in particular, the transform sequence of the fourth example can be more suitable for the case where four MTS candidates are applied.

[0398] The fifth to seventh exemplary transform arrays below correspond to the case where 35 transform sets are applied. The transform sets may be applied to each intra prediction mode in the case where they are mapped as shown in Table 24 below.

[0399] [Table 24]

[0400] In Table 24, NSST set index represents a transformation set index. When the mapping method of Table 24 is applied, it can be applied to the above-mentioned combinations A to E. That is, it can be applied to the fifth to eighth examples in the same manner as the above-mentioned method for each combination.

[0401] The fifth and eighth example transformation arrays may be applied to embodiments in which each transformation set is composed of two transformation matrices, and the sixth and seventh example transformation arrays may be applied to embodiments in which each transformation set is composed of one transformation matrix.

[0402] In particular, the fifth example may be applied to the above-mentioned combination D and Case 1 of the embodiment described in Table 15, and may also be applied to combination A and Case 1 of the embodiment described in Table 12, combination B and Case 1 of the embodiment described in Table 13, combination C and Case 1 of the embodiment described in Table 14, or combination E and Case 1 of the embodiment described in Table 16.

[0403] In particular, the sixth and seventh exemplary transformation arrays (i.e., transformation sets) may be applied to Case 3 of the embodiment described in combination D and Table 15 above, and further to Case 3 of combination A and the embodiment described in Table 12, Case 3 of combination B and the embodiment described in Table 13, Case 3 of combination C and the embodiment described in Table 14, or Case 3 of combination E and the embodiment described in Table 16.

[0404] Although the above-mentioned combinations A, B, C, D, and E only deal with cases where the number of MTS candidates is three or less, it is also possible to configure the first transform to apply all four MTS candidates to all intra prediction modes. The following fifth to eighth examples can also be used when all four MTS candidates are applied, and in particular, the transform sequence of the eighth example can be more suitable for the case where four MTS candidates are applied.

[0405] Among the following first to eighth example transformation arrays, all of the transformation examples that can be applied to a 4×4 region correspond to transformation matrices multiplied by a scaling value of 128. The following example transformation arrays can be commonly expressed as a g_aiNsst4×4[N1][N2]

[16]

[16] array, where N1 represents the number of transformation sets. Here, N1 can be 4 or 35, and can be divided into indexes 0, 1, ..., N1-1. N2 represents the number of transformation matrices that make up each transformation set (i.e., 1 or 2), and

[16]

[16] represents a 16×16 transformation matrix.

[0406] Similarly, in the following example, when a transform set is composed of one transform, it can be set to use a specific order of transform matrices for each transform set. For example, when a transform set includes one transform matrix, the encoder / decoder can use the predefined transform matrix, i.e., the first or second transform matrix, in each transform set.

[0407] To reduce the worst-case computational complexity, the encoder / decoder can apply an 8×16 matrix to a 4×4 TU. The following exemplary transformations that can be applied to 4×4 TUs can be applied to 4×4 TUs, 4×M TUs, and M×4 TUs (M>4), and when applied to 4×M TUs and M×4 TUs, the transformations can be applied by dividing the 4×4 regions into separate regions, or can be applied only to the top-left 4×8 or 8×4 region. Also, the transformations can be applied only to the top-left 4×4 region.

[0408] In one embodiment, to reduce the worst case computational complexity, the following method may be applied.

[0409] For example, for a block with width W and height H, if W>=8 and H>=8, the encoder / decoder may apply a transform array (or transform matrix, transform kernel) that can be applied to an 8x8 region to the top left 8x8 region (e.g., a 16x64 matrix). If W=8 and H=8, the encoder / decoder may apply only the 8x64 portion of the 16x64 matrix. In this case, the input for the secondary transform may be generated with 8 transform coefficients, and the remaining coefficients in the region may be considered to be 0.

[0410] Also, for example, for a block with width W and height H, if one of W and H is less than 8 (i.e., 4), the encoder / decoder can apply a transform array that can be applied to a 4×4 region. If W=4 and H=4, the encoder / decoder can apply only an 8×16 portion of a 16×16 matrix. In this case, the input for the secondary transform can be generated with 8 transform coefficients, and the remaining coefficients in the region can be considered to be 0.

[0411] In one embodiment, if (W, H)=(4, 8) or (8, 4), the encoder / decoder can apply the quadratic transform only to the top left 4×4 region. If W or H is greater than 8, the encoder / decoder can apply the quadratic transform only to the top left two 4×4 blocks. That is, the encoder / decoder can apply the transform matrix specified for two 4×4 blocks only up to the top left 4×8 or 8×4 region at most.

[0412] First Example

[0413] A first example can be defined as shown in Table 25 below. Four transformation sets can be defined, and each transformation set can be composed of two transformation matrices.

[0414] [Table 25]

[0415] Second Example

[0416] A second example can be defined as shown in Table 26 below. Four transformation sets can be defined, and each transformation set can be composed of one transformation matrix.

[0417] [Table 26]

[0418] Third Example

[0419] A third example can be defined as shown in Table 27 below. Four transformation sets can be defined, and each transformation set can be composed of one transformation matrix.

[0420] [Table 27]

[0421] Fourth Example

[0422] A fourth example can be defined as shown in Table 28 below. Four transformation sets can be defined, and each transformation set can be composed of two transformation matrices.

[0423] [Table 28]

[0424] Fifth Example

[0425] The fifth example can be defined as shown in Table 29 below. Thirty-five transformation sets can be defined, and each transformation set can be composed of two transformation matrices.

[0426] [Table 29]

[0427] Sixth Example

[0428] The sixth example can be defined as shown in Table 30 below. Thirty-five transformation sets can be defined, and each transformation set can be composed of one transformation matrix.

[0429] [Table 30]

[0430] Seventh Example

[0431] The seventh example can be defined as shown in Table 31 below. Thirty-five transformation sets can be defined, and each transformation set can be composed of one transformation matrix.

[0432] [Table 31]

[0433] Example No. 8

[0434] The eighth example can be defined as shown in Table 32 below. Thirty-five transformation sets can be defined, and each transformation set can be composed of two transformation matrices.

[0435] [Table 32]

[0436] The following describes examples of secondary transformations that can be applied to an 8x8 region. Among the following examples of secondary transformation sets available for the secondary transformation, the ninth and twelfth examples can be applied to an embodiment in which each transformation set is composed of two transformation matrices. The tenth and eleventh examples can be applied to an embodiment in which each transformation set is composed of one transformation matrix.

[0437] In particular, the ninth example may be applied to the above-mentioned combination D and Case 1 of the embodiment described in Table 15, and may also be applied to combination A and Case 1 of the embodiment described in Table 12, combination B and Case 1 of the embodiment described in Table 13, combination C and Case 1 of the embodiment described in Table 14, or combination E and Case 1 of the embodiment described in Table 16.

[0438] In particular, the tenth example transformation array (i.e., transformation set) may be applied to Case 3 of the embodiment described in combination D and Table 15 above, and further to Case 3 of combination A and the embodiment described in Table 12, Case 3 of combination B and the embodiment described in Table 13, Case 3 of combination C and the embodiment described in Table 14, or Case 3 of combination E and the embodiment described in Table 16.

[0439] Although the above-mentioned combinations A, B, C, D, and E only deal with cases where the number of MTS candidates is three or less, it is also possible to configure the first transform to apply all four MTS candidates to all intra prediction modes. The following ninth to twelfth examples can also be used when all four MTS candidates are applied, and in particular, the transform sequence of the twelfth example can be more suitable for the case where four MTS candidates are applied.

[0440] The following thirteenth to sixteenth exemplary transform arrays correspond to the case where 35 transform sets are applied, and can be applied when mapping the transform sets as shown in Table 24 for each intra prediction mode.

[0441] In Table 24, NSST set index represents a transformation set index. When the mapping method of Table 24 is applied, it can be applied to the above-mentioned combinations A to E. That is, it can be applied to the 13th to 16th examples in the same manner as the above-mentioned method for each combination.

[0442] The thirteenth and sixteenth example transformation arrays may be applied to embodiments in which each transformation set is composed of two transformation matrices, and the fourteenth and fifteenth example transformation arrays may be applied to embodiments in which each transformation set is composed of one transformation matrix.

[0443] In particular, the thirteenth example may be applied to the above-mentioned combination D and Case 1 of the embodiment described in Table 15, and may also be applied to combination A and Case 1 of the embodiment described in Table 12, combination B and Case 1 of the embodiment described in Table 13, combination C and Case 1 of the embodiment described in Table 14, or combination E and Case 1 of the embodiment described in Table 16.

[0444] In particular, the 14th and 15th example transformation arrays (i.e., transformation sets) may be applied to Case 3 of the embodiment described in combination D and Table 15 above, and further to Case 3 of combination A and the embodiment described in Table 12, Case 3 of combination B and the embodiment described in Table 13, Case 3 of combination C and the embodiment described in Table 14, or Case 3 of combination E and the embodiment described in Table 16.

[0445] Although the above-mentioned combinations A, B, C, D, and E only deal with cases where the number of MTS candidates is three or less, it is also possible to configure the first transform to apply all four MTS candidates to all intra prediction modes. The following thirteenth to sixteenth examples can also be used when all four MTS candidates are applied, and in particular, the eighth example transform sequence can be further adapted to the case where four MTS candidates are applied.

[0446] Among the following eighth to sixteenth exemplary transform arrays, the transform examples that can be applied to an 8×8 region all correspond to transform matrices multiplied by a scaling value of 128. The following exemplary transform arrays can be commonly expressed as a g_aiNsst8×8[N1][N2]

[16]

[64] array, where N1 represents the number of transform sets. Here, N1 can be 4 or 35 and can be divided into indexes 0, 1, ..., N1-1. N2 represents the number of transform matrices constituting each transform set (i.e., 1 or 2), and

[16]

[64] represents a 16×64 Reduced Secondary Transform (RST).

[0447] Similarly, in the following example, when a transform set is composed of one transform, it can be set to use a specific order of transform matrices for each transform set. For example, when a transform set includes one transform matrix, the encoder / decoder can use the predefined transform matrix, i.e., the first or second transform matrix, in each transform set.

[0448] When the RST is applied, 16 transform coefficients are output, but when only m×64 part of the 16×64 matrix is ​​applied, it can be configured to output only m transform coefficients. For example, by setting m=8 and multiplying only the top 8×64 matrix to output only 8 transform coefficients, the amount of calculation can be reduced by half.

[0449] Ninth Example

[0450] The ninth example can be defined as shown in Table 33 below. Four transformation sets can be defined, and each transformation set can be composed of two transformation matrices.

[0451] [Table 33-1]

[0452] [Table 33-2]

[0453] [Table 33-3]

[0454] Tenth Example

[0455] The tenth example can be defined as shown in Table 34 below. Four transformation sets can be defined, and each transformation set can be composed of one transformation matrix.

[0456] [Table 34]

[0457] Example No. 11

[0458] An eleventh example can be defined as shown in Table 35 below. Four transformation sets can be defined, and each transformation set can be composed of one transformation matrix.

[0459] [Table 35]

[0460] Example No. 12

[0461] The twelfth example can be defined as shown in Table 36 below. Four transformation sets can be defined, and each transformation set can be composed of two transformation matrices.

[0462] [Table 36-1]

[0463] [Table 36-2]

[0464] 13th Example

[0465] The thirteenth example can be defined as shown in Table 37 below. Thirty-five transformation sets can be defined, and each transformation set can be composed of two transformation matrices.

[0466] [Table 37-1]

[0467] [Table 37-2]

[0468] Fourteenth Example

[0469] The fourteenth example can be defined as shown in Table 38 below. Thirty-five transformation sets can be defined, and each transformation set can be composed of one transformation matrix.

[0470] [Table 38]

[0471] Fifteenth Example

[0472] The fifteenth example can be defined as shown in Table 39 below. Thirty-five transformation sets can be defined, and each transformation set can be composed of one transformation matrix.

[0473] [Table 39]

[0474] Example 16

[0475] The 16th example can be defined as shown in Table 40 below. Thirty-five transformation sets can be defined, and each transformation set can be composed of two transformation matrices.

[0476] [Table 40-1]

[0477] [Table 40-2]

[0478] FIG. 27 is a flowchart illustrating a method for converting a video signal according to an embodiment to which the present invention is applied.

[0479] For convenience of explanation, the following description will be focused on a decoder as shown in Fig. 27, but the present invention is not limited thereto, and the conversion method for a video signal according to the present embodiment can be substantially similarly applied to an encoder. The flowchart of Fig. 27 can be performed by the decoding device 200 or the inverse conversion unit 230.

[0480] The decoding apparatus 200 determines a secondary transform set to be applied to a current block from among predefined secondary transform sets based on an intra prediction mode of the current block (S2701).

[0481] The decoding apparatus 200 obtains a first syntax element indicating a secondary transform matrix to be applied to the current block within the determined secondary transform set (S2702).

[0482] Decoding device 200 derives a quadratic inverse transformed block by performing a quadratic inverse transform on the upper left corner region of the current block using the quadratic transform matrix specified by the first syntax element (S2703).

[0483] The decoding apparatus 200 performs a primary inverse transform on the secondarily inverse transformed block using a primary transform matrix of the current block, thereby deriving a residual block of the current block (S2704).

[0484] As mentioned above, each of the predefined quadratic transformation sets may include two quadratic transformation matrices.

[0485] As described above, step S2704 may further include determining an input length and an output length of the quadratic inverse transform based on the width and height of the current block. As described above, if the height and width of the current block are each 4, the input length of the non-separable transform may be determined to be 8 and the output length may be determined to be 16.

[0486] As described above, the decoding apparatus 200 may parse a second syntax element indicating a linear transformation matrix to be applied to the linear transformation of the current block, and may determine whether a secondary transformation may be applied to the current block based on the second syntax element.

[0487] As mentioned above, the step of determining whether a secondary transform can be applied may be performed by determining that a secondary transform can be applied to the current block if the second syntax element indicates a predefined specific transform type.

[0488] As mentioned above, the predefined specific transform type may be defined as DCT2.

[0489] 28 shows an example of a block diagram of an apparatus for processing a video signal as an embodiment to which the present invention is applied. The video signal processing apparatus of FIG. 28 may correspond to the encoding apparatus of FIG. 1 or the decoding apparatus of FIG. 2.

[0490] An image processing device 2800 for processing an image signal includes a memory 2820 for storing an image signal, and a processor 2810 coupled to the memory for processing the image signal.

[0491] The processor 2810 according to the embodiment of the present invention may be configured with at least one processing circuit for processing a video signal, and may process the video signal by executing a command for encoding or decoding the video signal. That is, the processor 2810 may encode original video data or decode an encoded video signal by executing the encoding or decoding method described above.

[0492] In addition, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable storage medium. Multimedia data having a data structure according to the present invention can also be stored in a computer-readable storage medium. The computer-readable storage medium includes any kind of storage device and distributed storage device in which computer-readable data is stored. The computer-readable storage medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable storage medium also includes a medium realized in the form of a carrier wave (e.g., transmission through the Internet). Also, a bit stream generated by the encoding method can be stored in a computer-readable storage medium or transferred via a wired or wireless communication network.

[0493] Furthermore, the embodiments of the present invention may be realized as a computer program product by program code, which may be executed on a computer according to the embodiments of the present invention. The program code may be stored on a carrier readable by the computer.

[0494] As described above, the embodiments described in the present invention may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip.

[0495] In addition, the decoder and encoder to which the present invention is applied may be included in real-time communication devices such as multimedia broadcasting transmitting / receiving devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video interactive devices, video communications, etc., mobile streaming devices, storage media, camcorders, video on demand (VoD) service providing devices, OTT video (Over the top video) devices, Internet streaming service providing devices, three-dimensional (3D) video devices, video phones, video devices, and medical video devices, and are used to process video signals or data signals. For example, OTT video (Over the top video) devices may include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0496] In addition, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable storage medium. Multimedia data having a data structure according to the present invention can also be stored in a computer-readable storage medium. The computer-readable recording medium includes any kind of storage device and distributed storage device in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium realized in the form of a carrier wave (e.g., transmission through the Internet). In addition, a bit stream generated by the encoding method can be stored in a computer-readable recording medium or transferred via a wired or wireless communication network.

[0497] Furthermore, the embodiments of the present invention may be realized as a computer program product by program code, which may be executed on a computer according to the embodiments of the present invention. The program code may be stored on a carrier readable by the computer.

[0498] The above-described embodiments are combinations of the components and features of the present invention in a predetermined form. Each component or feature should be considered as optional unless otherwise explicitly stated. Each component or feature may be implemented in a form not combined with other components or features. It is also possible to combine some components and / or features to configure an embodiment of the present invention. The order of operations described in the embodiments of the present invention may be changed. Some configurations or features of any embodiment may be included in other embodiments, or may be replaced with corresponding configurations or features of other embodiments. It is obvious that claims that do not have an explicit reference relationship in the claims may be combined to configure an embodiment, or may be included as a new claim by amendment after filing.

[0499] Embodiments of the present invention may be implemented in various ways, such as hardware, firmware, software, or a combination thereof. In a hardware implementation, an embodiment of the present invention may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc.

[0500] In the case of implementation by firmware or software, an embodiment of the present invention may be implemented in the form of a module, procedure, function, etc. that performs the functions or operations described above. The software code may be stored in a memory and driven by a processor. The memory may be located inside or outside the processor and may transmit and receive data to and from the processor by various means known in the art.

[0501] It is obvious to those skilled in the art that the present invention can be embodied in other specific forms without departing from the essential features of the present invention. Therefore, the above detailed description should not be interpreted as limiting in all respects, but should be regarded as illustrative. The scope of the present invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present invention are included in the scope of the present invention. [Industrial Applicability]

[0502] The above-described preferred embodiments of the present invention have been disclosed for illustrative purposes, and those skilled in the art may improve, modify, substitute or add various other embodiments within the technical idea and technical scope of the present invention disclosed in the appended claims below.

Claims

1. A video decoding method performed by a decoding device, comprising: Obtaining a syntax element for a secondary transform matrix to be applied to the current block within a secondary transform set; performing an inverse secondary transform on the secondary transform coefficients of the current block based on the secondary transform matrix specified by the syntax element to obtain linear transform coefficients, the inverse secondary transform being a non-separable transform; performing a primary inverse transform on the primary transform coefficients based on a primary transform matrix of the current block to obtain a residual block of the current block; The secondary transform set is determined based on an intra prediction mode of the current block among four secondary transform sets; the secondary transformation matrix is ​​one of two secondary transformation matrices included in the secondary transformation set, A method for decoding an image, wherein the number of input coefficients of the inverse secondary transform and the number of output coefficients of the inverse secondary transform are determined to be 8 and 16, respectively, based on the width and height of the current block being equal to 4.

2. A method for encoding video by an encoding device, comprising: deriving a prediction sample of the current block based on an intra-prediction mode of the current block; deriving residual samples of the current block based on the prediction samples; performing a linear transform on the residual samples to obtain linear transform coefficients of the current block; performing a quadratic transform on the linearly transformed coefficients based on a quadratic transform matrix to obtain quadratic transform coefficients, the quadratic transform being a non-separable transform; generating a syntax element for the quadratic transformation matrix in a quadratic transformation set; The secondary transform set is determined based on an intra prediction mode of the current block among four secondary transform sets; the secondary transformation matrix is ​​one of two secondary transformation matrices included in the secondary transformation set, A method for encoding an image, wherein the number of input coefficients of the quadratic transform and the number of output coefficients of the quadratic transform are determined to be 16 and 8, respectively, based on the width and height of the current block being equal to 4.

3. A method for transmitting data for a video, comprising the steps of: obtaining a bitstream for the image, the bitstream being generated by: deriving prediction samples of the current block based on an intra prediction mode of the current block; deriving residual samples of the current block based on the prediction samples; performing a primary transform on the residual samples to obtain primary transform coefficients of the current block; performing the secondary transform on the primary transformed coefficients based on a secondary transform matrix to obtain secondary transform coefficients, the secondary transform being a non-separable transform; and generating syntax elements for the secondary transform matrix in a secondary transform set; transmitting the data including the bitstream; The secondary transform set is determined based on an intra prediction mode of the current block among four secondary transform sets; the secondary transformation matrix is ​​one of two secondary transformation matrices included in the secondary transformation set, A method for transmitting data, wherein the number of input coefficients of the quadratic transform and the number of output coefficients of the quadratic transform are determined to be 16 and 8, respectively, based on the width and height of the current block being equal to 4.