Video signal processing method and device

JP2025083568A5Active Publication Date: 2025-09-04LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025043546
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-09-02
Filing Date
2025-03-18
Publication Date
2025-09-04
Estimated Expiration
2039-09-02

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently processing next-generation video content with high spatial resolution, high frame rate, and high dimensionality, requiring more complex and computationally intensive methods.

Method used

The method involves applying a non-separable transform based on the size of the current block, determining the appropriate transform matrix, and applying it to the block, which includes specific steps for different block sizes and intra prediction modes.

Benefits of technology

This approach provides a video coding method with high coding efficiency and low complexity, effectively handling the increased demands of next-generation video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To realize a video signal processing method and device which apply transformation having high coding efficiency and low complexity.SOLUTION: There are provided video signal processing method and device according to an embodiment of the present invention. A decoding method of an image signal according to an embodiment of the present invention includes the steps of: determining an input length and an output length of non-separable transformation on the basis of height and width of a current block; determining a non-separable transformation matrix corresponding to the input length and the output length of the non-separable transformation, and applying the non-separable transformation matrix to the current block; and when each of the height and width of the current block is 4, the input length of the non-separable transformation is determined to be 8, and the output length is determined to be 16.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and an apparatus for processing video signals, and more particularly to a method and an apparatus for encoding or decoding video signals by performing conversion.

Background Art

[0002] Compression encoding (encoding) means a series of signal processing techniques for transferring digitized information via a communication line or storing it in a form suitable for a storage medium. Media such as video, video, and audio can be the target of compression encoding, and in particular, the technique of performing compression encoding on video is called video compression.

[0003] Next-generation video content will have characteristics such as high spatial resolution, high frame rate, and high dimensionality of scene representation. Processing such content will result in a huge increase in terms of memory storage, memory access rate, and processing power.

[0004] Therefore, it is necessary to design coding tools for more efficiently processing next-generation video content. In particular, video codec standards after the HEVC (High Efficiency Video Coding) standard require efficient conversion techniques for converting video signals in the spatial domain to the frequency domain, along with prediction techniques having higher accuracy.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Embodiments of the present invention seek to provide an image signal processing method and apparatus that apply a transform having high coding efficiency and low complexity.

[0006] The technical problems to be solved by the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned should be clearly understood by those with ordinary knowledge in the technical field to which the present invention belongs from the following description.

Means for Solving the Problems

[0007] The method for decoding an image signal according to an embodiment of the present invention includes steps of determining an input length and an output length of a non-separable transform based on a height and a width of a current block, determining a non-separable transform matrix corresponding to the input length and the output length of the non-separable transform, and applying the non-separable transform matrix to the current block. When the height and the width of the current block are both 4, the input length of the non-separable transform is determined to be 8 and the output length is determined to be 16.

[0008] Also, when the height and the width of the current block do not respectively correspond to 4, the input length and the output length of the non-separable transform are each determined to be 16.

[0009] Also, the step of applying the non-separable transform matrix to the current block includes a step of applying the non-separable transform matrix to a 4×4 region in the upper left side of the current block when the height and the width do not respectively correspond to 4 and the product of the width and the height is smaller than a threshold value.

[0010] Also, the step of applying the non-separable transform matrix to the current block includes a step of applying the non-separable transform matrix to a 4×4 region in the upper left side of the current block and a 4×4 region located on the right side of the 4×4 region in the upper left side when the height and the width do not respectively correspond to 4 and the width is greater than or equal to the height.

[0011] In addition, the step of applying the non-separable transform matrix to the current block does not apply when the height and width are each 4, and when the product of the width and height is greater than or equal to a threshold value and the width is smaller than the height, the step of applying the non-separable transform matrix to the upper left 4×4 region of the current block and the 4×4 region located below the upper left 4×4 region is included.

[0012] In addition, the step of determining the non-separable transform matrix includes a step of determining a non-separable transform set index based on the intra prediction mode of the current block, a step of determining a non-separable transform kernel corresponding to the non-separable transform index within the non-separable transform set having the non-separable transform set index, and a step of determining a non-separable transform matrix from the non-separable transform kernel based on the input length and the output length.

[0013] An image signal decoding apparatus according to another embodiment of the present invention includes a memory for storing an image signal and a processor coupled to the memory. The processor is configured to determine the input length and the output length of the non-separable transform based on the height and width of the current block, determine a non-separable transform matrix corresponding to the input length and the output length of the non-separable transform, and apply the non-separable transform matrix to the current block. When the height and width of the current block are each 4, the length of the non-separable transform is determined to be 8 and the output is 16.

Advantages of the Invention

[0014] According to an embodiment of the present invention, by applying a transform based on the size of a current block, it is possible to provide a video coding method and apparatus having high coding efficiency and low complexity.

[0015] The effects obtained by the present invention are not limited to the effects mentioned above, and other effects not mentioned should be clearly understood by those of ordinary skill in the technical field to which the present invention pertains from the following description.

Brief Description of the Drawings

[0016]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21A

Figure 21B

Figure 22

Figure 23

Figure 24

Figure 25A

Figure 25B

Figure 26A

Figure 26B

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Embodiments for Carrying Out the Invention

[0017] The accompanying drawings included in part of the detailed description to assist in understanding the present invention provide embodiments of the present invention and explain the technical features of the present invention together with the detailed description.

[0018] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. The detailed description disclosed below together with the attached drawings is intended to explain exemplary embodiments of the present invention, and is not intended to show the only embodiments in which the present invention can be implemented. The following detailed description includes specific details to provide a complete understanding of the present invention. However, those skilled in the art will understand that the present invention can be implemented without such specific details.

[0019] In some cases, to avoid obscuring the concept of the present invention, well-known structures and devices may be omitted or may be shown in the form of block diagrams centered on the core functions of each structure and device.

[0020] In some cases, to avoid obscuring the concept of the present invention, well-known structures and devices may be omitted or may be shown in the form of block diagrams centered on the core functions of each structure and device.

[0021] The specific terms used in the following description are provided to assist in understanding the present invention, and the use of such specific terms can be changed to other forms without departing from the technical idea of the present invention. For example, in the case of signals, data, samples, pictures, frames, blocks, etc., they may be appropriately substituted and interpreted in each coding process.

[0022] Hereinafter, in this specification, the "processing unit" means a unit in which encoding / decoding processing such as prediction, conversion, and / or quantization is executed. Also, the processing unit can be interpreted in a sense that includes a unit of a luma component and a unit of a chroma component. For example, the processing unit can correspond to a block, a coding unit (CU), a prediction unit (PU), or a transform unit (TU).

[0023] Also, the processing unit can be interpreted as a unit of a luma component or a unit of a chroma component. For example, the processing unit can correspond to a CTB, a CB, a PU, or a TB of a luma component. Alternatively, the processing unit can correspond to a CTB, a CB, a PU, or a TB of a chroma component. Also, without being limited thereto, the processing unit may be interpreted in a sense that includes a unit of a luma component and a unit of a chroma component.

[0024] Also, the processing unit is not necessarily limited to a square block, and may be configured in a polygon shape having three or more vertices.

[0025] Also, hereinafter, in this specification, pixels, picture elements, or coefficients (conversion coefficients or conversion coefficients after a primary conversion) and the like are collectively referred to as samples. And using a sample means using a pixel value, a picture element value, or a coefficient (conversion coefficient or conversion coefficient after a primary conversion) and the like.

[0026] Hereinafter, regarding an encoding / decoding method for a still image or a moving image, a design and an application method of a reduced secondary transform (RST) considering the worst-case computational complexity will be described.

[0027] Embodiments of the present invention provide an image and video compression method and apparatus. The compressed data has the form of a bitstream, and the bitstream can be stored in various forms of storage, and can also be streamed via a network and transmitted to a terminal having a decoder. In the terminal, when a display device is attached, the image decoded by the display device may be displayed, or the bitstream data may simply be stored. The methods and apparatuses proposed in the embodiments of the present invention can be applied to both the encoder and the decoder, can be applied to all apparatuses that generate or receive a bitstream, and can be applied regardless of whether or not to output via a display device in the terminal.

[0028] The image compression apparatus is composed of a prediction unit, a conversion and quantization unit, and an entropy coding unit. The schematic block diagrams of the encoding apparatus and the decoding apparatus are as shown in FIGS. 1 and 2. Among them, in the conversion and quantization unit, a prediction signal is subtracted from the original signal to obtain a residual signal, which is then converted into a frequency domain signal by a conversion such as DCT (Discrete Cosine Transform)-2, and then quantization is applied to significantly reduce the number of non-zero signals to enable image compression.

[0029] FIG. 1 shows a schematic block diagram of an encoding apparatus in an embodiment to which the present invention is applied, where video / image signal encoding is performed.

[0030] The image segmentation unit 110 divides an input image (or picture, frame) input to the encoding device 100 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit is recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a quad-tree binary-tree (QTBT) structure. For example, one coding unit is divided into a plurality of coding units with a deeper depth based on a quad-tree structure and / or a binary-tree structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure may be applied later. Alternatively, the binary-tree structure may be applied first. The coding procedure according to the present invention is performed based on the final coding unit that cannot be further divided. In this case, the largest coding unit may be immediately used as the final coding unit based on the coding efficiency according to the image characteristics, etc., or, if necessary, the coding unit is recursively divided into coding units with a deeper depth so that a coding unit of an optimal size is used as the final coding unit. Here, the coding procedure includes procedures such as prediction, conversion, and restoration described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit are respectively divided or partitioned from the final coding unit described above. The prediction unit is a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0031] The unit may, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block represents a set of samples or transform coefficients consisting of M columns and N rows. Samples generally represent pixels or pixel values, and may represent only the pixel / pixel values of the luma component, or only the pixel / pixel values of the chroma component. Samples can be used as terms corresponding to pixels (or pels) for one picture (or image).

[0032] The encoding device 100 subtracts the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 180 or the intra prediction unit 185 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 120. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 100 may be called the subtraction unit 115. The prediction unit performs prediction on the block to be processed (hereinafter referred to as the current block), and generates a predicted block including the prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit generates various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmits it to the entropy encoding unit 190. The information related to prediction is encoded in the entropy encoding unit 190 and output in the form of a bitstream.

[0033] The intra prediction unit 185 predicts the current block by referring to samples within the current picture. The samples to be referred to are located either in the neighborhood of the current block or away from it depending on the prediction mode. In intra prediction, the prediction mode includes a plurality of non-directional modes and a plurality of directional modes. The non-directional modes include, for example, the DC mode and the planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is just an example, and a greater or lesser number of directional prediction modes may be used depending on the setting. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0034] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information includes a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral blocks include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The above temporal neighboring blocks may be called by names such as collocated reference blocks and collocated CU (colCU), and the reference picture including the temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on the peripheral blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the peripheral blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted.In the case of the Motion Vector Prediction (MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and by signaling the motion vector difference, the motion vector of the current block can be indicated.

[0035] The prediction signal generated by the inter prediction unit 180 or the intra prediction unit 185 is used to generate a restored signal or to generate a residual signal.

[0036] The transform unit 120 applies a transform technique to the residual signal to generate transform coefficients. For example, the transform technique includes at least one of DCT, DST (Discrete Sine Transform), KLT (Karhunen - Loeve Transform), GBT (Graph - Based Transform), or CNT (Conditionally Non - linear Transform). Here, GBT means the transform obtained from the graph when representing the relationship information between pixels as a graph. CNT means the transform obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the transform process may be applied to pixel blocks having the same size of a square or to blocks of variable sizes that are not square.

[0037] The quantization unit 130 quantizes the transform coefficients and transmits them to the entropy encoding unit 190. The entropy encoding unit 190 encodes the quantized signal (information regarding the quantized transform coefficients) and outputs it as a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. The quantization unit 130 can also reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and generate the information regarding the quantized transform coefficients based on the one-dimensional vector-form quantized transform coefficients. The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding). The entropy encoding unit 190 can also encode, together or separately, information necessary for video / image restoration (such as, for example, the values of syntax elements) in addition to the quantized transform coefficients. The encoded information (such as, for example, video / image information) is transmitted or stored in units of NAL (Network Abstraction Layer) units in the form of a bitstream. The above bitstream is transmitted via a network or stored in a digital storage medium. Here, the network includes a broadcast network and / or a communication network, etc., and the digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 may be configured as internal / external elements of the encoding device 100, or the transmission unit may be a component of the entropy encoding unit 190.

[0038] The quantized transform coefficients output from the quantization unit 130 can be utilized to generate a prediction signal. For example, the quantized transform coefficients can restore the residual signal by applying inverse quantization and inverse transformation through the inverse quantization unit 140 and the inverse transformation unit 150 within the loop. The addition unit 155 generates a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 may be referred to as a restoration unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next processing target block within the current picture, or may be used for inter prediction of the next picture after passing through filtering as described later.

[0039] The filtering unit 160 can apply filtering to the reconstructed signal to improve the subjective / objective image quality. For example, the filtering unit 160 applies various filtering methods to the reconstructed picture to generate a modified reconstructed picture and transmits the modified reconstructed picture to the decoded picture buffer 170. Various filtering methods include, for example, deblock filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 generates various information related to filtering as described later and transmits it to the entropy encoding unit 190 in the description of each filtering method. The information related to filtering is encoded in the entropy encoding unit 190 and output in the form of a bitstream.

[0040] The corrected decoded picture transmitted to the decoded picture buffer 170 is used as a reference picture in the inter prediction unit 180. The encoding device 100 can thereby avoid prediction mismatches between the encoding device 100 and the decoding device when inter prediction is applied, and can also improve the encoding efficiency.

[0041] The decoded picture buffer 170 can store the corrected restored picture for use as a reference picture in the inter prediction unit 180.

[0042] FIG. 2 shows a schematic block diagram of a decoding device in an embodiment to which the present invention is applied, in which an image signal is decoded.

[0043] As shown in FIG. 2, the decoding device 200 includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a decoded picture buffer (DPB) 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a prediction unit. That is, the prediction unit includes the inter prediction unit 180 and the intra prediction unit 185. The inverse quantization unit 220 and the inverse transform unit 230 may be collectively referred to as a residual processing unit. That is, the residual processing unit includes the inverse quantization unit 220 and the inverse transform unit 230. The above-described entropy decoding unit 210, inverse quantization unit 220, inverse transform unit 230, addition unit 235, filtering unit 240, inter prediction unit 260, and intra prediction unit 265 are configured by one hardware component (for example, a decoder or a processor) according to an embodiment. Also, the decoded picture buffer 250 is configured by one hardware component (for example, a memory or a digital storage medium) according to an embodiment.

[0044] When a bitstream including video / image information is input, the decoding device 200 can restore an image corresponding to the process in which the video / image information was processed in the encoding device 100 of FIG. 2. For example, the decoding device 200 performs decoding using the processing units applied in the encoding device 100. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit is divided from a coding tree unit or a maximum coding unit by a quadtree structure and / or a binary tree structure. Then, the restored image signal decoded and output by the decoding device 200 is reproduced by a reproducing device.

[0045] The decoding device 200 receives the signal output from the encoding device 100 of FIG. 2 in the form of a bitstream, and the received signal is decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 parses the bitstream to derive information (e.g., video / image information) necessary for image restoration (or, picture restoration). For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding and decoding target blocks, or the information of the symbol / bin decoded in the previous stage, predicts the occurrence probability of the bin according to the determined context model, and performs arithmetic decoding (decoding) of the bin to generate a symbol corresponding to the value of each syntax element. Here, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin. Among the information decoded in the entropy decoding unit 210, the information regarding prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual value for which entropy decoding is performed in the entropy decoding unit 210, that is, the quantized transform coefficient and related parameter information, are input to the inverse quantization unit 220. Also, the information regarding filtering among the information decoded in the entropy decoding unit 210 is provided to the filtering unit 240. On the other hand, the receiving unit (not shown) that receives the signal output from the encoding device 100 may be further configured as an internal / external element of the decoding device 200, or the receiving unit may be a component of the entropy decoding unit 210.

[0046] In the inverse quantization unit 220, the quantized transform coefficients are inverse quantized to output transform coefficients. The inverse quantization unit 220 rearranges the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement (sequencing) can be performed based on the coefficient scan order performed in the encoding apparatus 100. The inverse quantization unit 220 performs inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain transform coefficients.

[0047] The inverse transform unit 230 obtains a residual signal (residual block, residual sample array) by inverse transforming the transform coefficients.

[0048] The prediction unit performs prediction on the current block and generates a predicted block including prediction samples for the current block. The prediction unit determines whether intra prediction or inter prediction is to be applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode.

[0049] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the vicinity (neighbor) of the current block or at a distance depending on the prediction mode. In intra prediction, the prediction mode includes all of a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 265 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0050] The inter prediction unit 260 derives a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter prediction mode, motion information is predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, neighboring blocks include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 constructs a motion information candidate list based on information regarding prediction of neighboring blocks, and derives a motion vector and / or a reference picture index of the current block based on the received candidate selection information. Inter prediction is performed based on various prediction modes, and the information regarding prediction includes information indicating the mode of inter prediction for the current block.

[0051] The addition unit 235 generates a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to a predicted signal (predicted block, predicted sample array) output from the inter prediction unit 260 or the intra prediction unit 265. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0052] The addition unit 235 may be referred to as a restoration unit or a restored block generation unit. The generated restored signal may be used for intra prediction of the next processing target block within the current picture, or may be used for inter prediction of the next picture after passing through filtering as described later.

[0053] The filtering unit 240 can improve the subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 applies various filtering methods to the restored picture to generate a modified restored picture, and sends the modified restored picture to the decoded picture buffer 250. Various filtering methods include, for example, deblocking filtering, Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), and bilateral filter.

[0054] The modified decoded picture sent to the decoded picture buffer 250 can be used as a reference picture by the inter prediction unit 260.

[0055] In this document, the embodiments described in the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the encoding device 100 are also applied identically or correspondingly to the filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of the decoding device 200, respectively.

[0056] FIG. 3 is an embodiment to which the present invention can be applied. FIG. 3A illustrates a block partitioning structure by QuadTree (QT), FIG. 3B illustrates a block partitioning structure by Binary Tree (BT), FIG. 3C illustrates a block partitioning structure by Ternary Tree (TT), and FIG. 3D illustrates a block partitioning structure by Asymmetric Tree (AT).

[0057] In video coding, one block can be divided based on QT. Also, one subblock divided by QT may be recursively further divided using QT. A leaf block that is no longer divided by QT is divided by at least one of the BT, TT, or AT methods. BT can have two forms of division: horizontal BT (2N×N, 2N×N) and vertical BT (N×2N, N×2N). TT can have two forms of division: horizontal TT (2N×1 / 2N, 2N×N, 2N×1 / 2N) and vertical TT (1 / 2N×2N, N×2N, 1 / 2N×2N). AT can have four forms of division: horizontal-up AT (2N×1 / 2N, 2N×3 / 2N), horizontal-down AT (2N×3 / 2N, 2N×1 / 2N), vertical-left AT (1 / 2N×2N, 3 / 2N×2N), and vertical-right AT (3 / 2N×2N, 1 / 2N×2N). Each of the BT, TT, and AT may be recursively further divided using BT, TT, and AT.

[0058] Figure 3A shows an example of QT division. Block A is divided into four subblocks (A0, A1, A2, A3) by QT. Subblock A1 is again divided into four subblocks (B0, B1, B2, B3) by QT.

[0059] Figure 3B shows an example of BT division. Block B3, which is no longer divided by QT, is divided into vertical BT (C0, C1) or horizontal BT (D0, D1). Each subblock, like block C0, is recursively further divided in the form of horizontal BT (E0, E1) or vertical BT (F0, F1).

[0060] Figure 3C shows an example of TT division. The block B3 that cannot be further divided by QT is divided into vertical TT (C0, C1, C2) or horizontal TT (D0, D1, D2). Each sub-block, like block C1, is recursively further divided, for example, in the form of horizontal TT (E0, E1, E2) or vertical TT (F0, F1, F2).

[0061] Figure 3D shows an example of AT division. The block B3 that cannot be further divided by QT is divided into vertical AT (C0, C1) or horizontal AT (D0, D1). Each sub-block, like block C1, can be recursively further divided, for example, in the form of horizontal AT (E0, E1) or vertical TT (F0, F1).

[0062] On the other hand, BT, TT, and AT divisions may all be used. For example, the sub-blocks divided by BT can be divided by TT or AT. Also, the sub-blocks divided by TT can be divided by BT or AT. The sub-blocks divided by AT can be divided by BT or TT. For example, after horizontal BT division, each sub-block can be divided vertically by BT, or after vertical BT division, each sub-block can also be divided horizontally by BT. In this case, although the division order is different, the finally divided shapes are the same.

[0063] Also, when a block is divided, the order of exploring the blocks can be defined in various ways. Generally, the exploration is performed from left to right and from the upper end to the lower end. Exploring a block means the order of determining whether additional block division of each divided sub-block is possible, or when a block cannot be further divided, it means the encoding order of each sub-block, or the exploration order when referring to the information of other adjacent blocks in the sub-block.

[0064] Conversions can be performed separately for each processing unit (or conversion block) divided by a divided structure such as FIGS. 3A to 3D. In particular, the conversion matrix can be applied by dividing in the row direction and the column direction. According to an embodiment of the present invention, different conversion types can be used according to the length of the processing unit (or conversion block) in the row direction or the column direction.

[0065] FIGS. 4 and 5 are embodiments to which the present invention is applied. FIG. 4 shows a schematic block diagram of the conversion and quantization units 120 / 130 and the inverse quantization and inverse conversion units 140 / 150 in the encoding device 100 of FIG. 1, and FIG. 5 shows a schematic block diagram of the inverse quantization and inverse conversion units 220 / 230 in the decoding device 200.

[0066] As shown in FIG. 4, the conversion and quantization units 120 / 130 include a primary transform unit 121, a secondary transform unit 122, and a quantization unit 130. The inverse quantization and inverse conversion units 140 / 150 include an inverse quantization unit 140, an inverse secondary transform unit 151, and an inverse primary transform unit 152.

[0067] As shown in FIG. 5, the inverse quantization and inverse conversion units 220 / 230 include an inverse quantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.

[0068] In the present invention, when performing conversion, the conversion is carried out through a plurality of stages. For example, as shown in FIG. 4, two stages of primary transform and secondary transform can be applied, and more conversion stages can also be used depending on the algorithm. Here, the primary transform may also be referred to as a core transform.

[0069] The primary conversion unit 121 applies a primary conversion to the residual signal, where the primary conversion can already be (pre-)defined as a table in the encoder and / or decoder.

[0070] The secondary conversion unit 122 applies a secondary conversion to the signal that has been primarily converted, where the secondary conversion can already be defined as a table in the encoder and / or decoder.

[0071] In one embodiment, a non-separable secondary transform (NSST) can be conditionally applied as the secondary conversion. For example, NSST is only applied when it is a prediction block within the screen, and can have a set of conversions applicable for each prediction mode group.

[0072] Here, the prediction mode group is set based on the symmetry with respect to the prediction direction. For example, prediction mode 52 and prediction mode 16 are symmetric with respect to prediction mode 34 (diagonal direction), so they can form one group and the same transform set can be applied. Here, when applying the conversion for prediction mode 52, it is applied after transposing the input data, because the prediction mode 16 and the conversion set are the same.

[0073] On the other hand, in the case of the planar mode and the DC mode, since there is no symmetry with respect to the direction, each has its own conversion set, and the conversion set can be composed of two conversions. For the remaining directional modes, each conversion set can be composed of three conversions.

[0074] The quantization unit 130 performs quantization on the secondarily converted signal.

[0075] The inverse quantization and inverse conversion units 140 / 150 perform the above-described process in reverse, and redundant descriptions are omitted.

[0076] FIG. 5 shows a schematic block diagram of the inverse quantization and inverse conversion units 220 / 230 in the decoding device 200.

[0077] As shown in FIG. 5, the inverse quantization and inverse conversion units 220 / 230 include an inverse quantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.

[0078] The inverse quantization unit 220 obtains conversion coefficients from the entropy-decoded signal using quantization step size information.

[0079] In the inverse secondary transform unit 231, an inverse secondary transform is performed on the conversion coefficients. Here, the inverse secondary transform represents the inverse transform of the secondary transform described in FIG. 4.

[0080] The inverse primary transform unit 232 performs an inverse primary transform on the inversely secondarily transformed signal (or block) to obtain a residual signal. Here, the inverse primary transform represents the inverse transform of the above-described primary transform described in FIG. 4.

[0081] FIG. 6 shows a flowchart for encoding a video signal by primary conversion and secondary conversion, which is an embodiment to which the present invention is applied. Each operation shown in FIG. 6 is performed by the conversion unit 120 of the encoding device 100.

[0082] The encoding device 100 determines (or selects) a forward secondary conversion based on at least one of the prediction mode, block shape, and / or block size of the current block (S610).

[0083] The encoding device 100 can determine an optimal forward secondary conversion by rate-distortion optimization. The optimal forward secondary conversion corresponds to one of a plurality of conversion combinations, and the plurality of conversion combinations are defined by a conversion index. For example, for rate-distortion optimization, the encoding device 100 can compare the results of performing all of forward secondary conversion, quantization, residual coding, etc. for each candidate.

[0084] The encoding device 100 signals a secondary conversion index corresponding to the optimal forward secondary conversion (S620). Here, the secondary conversion index can be applied to other embodiments described in this specification.

[0085] On the other hand, the encoding device 100 performs a forward primary conversion on the current block (residual block) (S630).

[0086] The encoding device 100 performs a forward secondary conversion on the current block using the optimal forward secondary conversion (S640). On the other hand, the forward secondary conversion can be the RST described below. The RST means a conversion in which N pieces of residual data (N×1 residual vector) are input and R pieces of conversion coefficient data (R×1 conversion coefficient vector) are output (R < N).

[0087] As one embodiment, RST can be applied to a specific region of the current block. For example, when the current block is N×N, the specific region may mean the upper left N / 2×N / 2 region. However, the present invention is not limited to this, and it is set to be different depending on at least one of the prediction mode, block shape, or block size. For example, when the current block is N×N, the specific region may mean the upper left M×M region (M≦N).

[0088] On the other hand, the encoding device 100 generates a transform coefficient block by performing quantization on the current block (S650).

[0089] The encoding device 100 can perform entropy encoding on the transform coefficient block to generate a bitstream.

[0090] FIG. 7 shows a flowchart for decoding a video signal by inverse second transformation and inverse first transformation, which is an embodiment to which the present invention is applied. Each operation shown in FIG. 7 is performed by the inverse transformation unit 230 of the decoding device 200.

[0091] The decoding device 200 acquires a second transformation index from the bitstream (S710).

[0092] The decoding device 200 derives a second transformation corresponding to the second transformation index (S720).

[0093] However, steps S710 and S720 are one embodiment, and the present invention is not limited to this. For example, the decoding device 200 can derive the second transformation based on at least one of the prediction mode, block shape, and / or block size of the current block without acquiring the second transformation index.

[0094] On the other hand, the decoder 200 entropy-decodes the bitstream to obtain a transform coefficient block, and performs inverse quantization on the transform coefficient block (S730).

[0095] The decoder 200 performs an inverse direction secondary transform on the inverse quantized transform coefficient block (S740). For example, the inverse direction secondary transform can be an inverse direction RST. The inverse direction RST is the transpose matrix of the RST described in FIG. 6, and means a transform in which R transform coefficient data (Rx1 transform coefficient vector) are input and N residual data (Nx1 residual vector) are output.

[0096] As one embodiment, the reduced secondary transform can be applied to a specific region of the current block. For example, when the current block is N×N, the specific region can mean the upper left N / 2×N / 2 region. However, the present invention is not limited thereto, and is set to vary depending on at least one of a prediction mode, a block shape, or a block size. For example, when the current block is N×N, the specific region can mean the upper left M×M region (M≦N) or M×L (M≦N, L≦N).

[0097] Then, the decoder 200 performs an inverse direction primary transform on the result of the inverse direction secondary transform (S750).

[0098] The decoder 200 generates a residual block in step S750, and generates a restored block by adding the residual block and the prediction block.

[0099] FIG. 8 shows an example of a transform configuration group to which AMT (Adaptive Multiple Transform) according to an embodiment of the present invention is applied.

[0100] According to FIG. 8, the transform configuration group is determined based on a prediction mode, and the number of groups can be six in total (G0 to G5). G0 to G4 correspond to the case where intra prediction is applied, and G5 indicates a transform combination (or transform set, transform combination set) applied to a residual block generated by inter prediction.

[0101] One transformation combination is composed of a horizontal transform (or row transform) applied to the rows of the corresponding two-dimensional block and a vertical transform (or column transform) applied to the columns.

[0102] Here, each of all the transformation setting groups includes four transformation combination candidates. The four transformation combination candidates are selected or determined by the transformation combination indexes from 0 to 3, and the transformation combination index is transmitted from the encoding device 100 to the decoding device 200 by the encoding procedure.

[0103] As one embodiment, the residual data (or residual signal) obtained by intra prediction has different statistical characteristics according to the intra prediction mode. Therefore, other transforms other than the general cosine transform can be applied according to the intra prediction mode as shown in FIG. 8. In this document, the transform type may be expressed, for example, as DCT-Type 2, DCT-II, DCT-2.

[0104] As shown in FIG. 8, the transform set configurations for the case where 35 intra prediction modes are used and the case where 67 intra prediction modes are used are respectively illustrated. A plurality of transformation combinations can be applied for each transformation setting group divided in the intra prediction mode column. For example, a plurality of transformation combinations (row direction transform, column direction transform) are composed of four combinations. More specifically, in group 0, DST-7 and DCT-5 can be applied to all in the row (horizontal) direction and the column (vertical) direction, so four combinations are possible.

[0105] For each intra prediction mode, a total of four combinations of transform kernels can be applied, so a transform combination index for selecting one of them is transmitted for each transform unit. In this document, the transform combination index is referred to as the AMT index and may be expressed as amt_idx.

[0106] In addition to the transform kernels shown in FIG. 8, due to the characteristics of the residual signal, DCT-2 may be optimal for both the row direction and the column direction. Therefore, adaptive transformation can be performed by defining an AMT flag for each coding unit. Here, when the AMT flag is 0, DCT-2 is applied to both the row direction and the column direction, and when the AMT flag is 1, one of the four combinations can be selected or determined by the AMT index.

[0107] As one embodiment, when the AMT flag is 0 and the number of transform coefficients for one transform unit is less than 3, instead of applying the transform kernel of FIG. 8, DST-7 is applied to both the row direction and the column direction.

[0108] As one embodiment, when the number of transform coefficients is first purged and the number of transform coefficients is less than 3, the amount of additional information transmitted can be reduced by applying DST-7 without purging the AMT index.

[0109] As one embodiment, AMT can be applied only when the width and height of the transform unit are both 32 or less.

[0110] As one embodiment, FIG. 8 may be pre-set by off-line training.

[0111] As one embodiment, the AMT index can be defined by one index that can simultaneously instruct a combination of a horizontal transform and a vertical transform. Alternatively, the AMT index can be separately defined by a horizontal transform index and a vertical transform index.

[0112] A technique of applying a selected transform among a plurality of transform kernels (e.g., DCT-2, DST-7, DCT-8) like the AMT described above may be referred to as MTS (Multiple Transform Selection) or EMT (Enhanced Multiple Transform), and the AMT index may be referred to as an MT index.

[0113] FIG. 9 shows an example of a flowchart of encoding to which AMT according to an embodiment of the present invention is applied. The operations shown in FIG. 9 are performed by the transform unit 120 of the encoding device 100.

[0114] This document basically describes an embodiment in which transforms are applied separately to the horizontal direction and the vertical direction, but the transform combination can also be composed of a non-separable transform.

[0115] Also, it can be composed of a mixture of a separable transform and a non-separable transform. In this case, when a non-separable transform is used, transform selection by row / column or selection by horizontal / vertical direction becomes unnecessary, and the transform combination in FIG. 8 above is used only when a separable transform (separable transform) is selected.

[0116] Also, the method proposed in this specification can be applied regardless of whether it is a primary transform or a secondary transform. That is, there is no restriction that it must be applied to only one of the two, and it can be applied to both. Here, the primary transform means a transform for first transforming a residual block, and the secondary transform may mean a transform for applying a transform to the block generated as a result of the primary transform.

[0117] First, the encoding device 100 determines a conversion setting group corresponding to the current block (S910). Here, the conversion setting group can also be configured in a combination as shown in FIG. 8.

[0118] The encoding device 100 performs conversion on the combinations of candidate conversions available within the conversion setting group (S920).

[0119] As a result of the conversion execution, the encoding device 100 determines or selects the conversion combination with the smallest RD (Rate Distortion) cost (S930).

[0120] The encoding device 100 encodes the conversion combination index corresponding to the selected conversion combination (S940).

[0121] FIG. 10 shows an example of a flowchart of decoding to which AMT according to an embodiment of the present invention is applied. The operations shown in FIG. 10 are performed by the inverse conversion unit 230 of the decoding device 200.

[0122] First, the decoding device 200 determines a conversion setting group for the current block (S1010). The decoding device 200 parses (or acquires) the conversion combination index from the video signal, where the conversion combination index corresponds to any one of a plurality of conversion combinations within the conversion setting group (S1020). For example, the conversion setting group includes DCT-2, DST-7, or DCT-8.

[0123] The decoding device 200 derives the conversion combination corresponding to the conversion combination index (S1030). Here, the conversion combination is composed of a horizontal conversion and a vertical conversion, and includes at least one of DCT-2, DST-7, or DCT-8. Also, the conversion combination may use the conversion combination described in FIG. 8.

[0124] The decoding device 200 performs inverse transformation on the current block based on the derived transformation combination (S1040). When the transformation combination is composed of row (horizontal) transformation and column (vertical) transformation, the row (horizontal) transformation can be applied first, and then the column (vertical) transformation can be applied. However, the present invention is not limited to this, and conversely, when it is applied in reverse or is composed of non-separable transformation, the non-separable transformation can be immediately applied.

[0125] In one embodiment, when the vertical transformation or the horizontal transformation is DST-7 or DCT-8, the inverse transformation of DST-7 or the inverse transformation of DCT-8 is applied for each column and then for each row. Also, different transformations are applied for each row and / or each column in the vertical transformation or the horizontal transformation.

[0126] In one embodiment, the transformation combination index can be obtained based on the AMT flag indicating whether AMT is performed. That is, the transformation combination index can be obtained only when AMT is performed by the AMT flag. Also, the decoding device 200 checks whether the number of non-zero coefficients is greater than a threshold. Here, the transformation combination index can be purged only when the number of non-zero coefficients is greater than the threshold.

[0127] In one embodiment, the AMT flag or the AMT index is defined at at least one level of sequence, picture, slice, block, coding unit, transform unit, or prediction unit.

[0128] On the other hand, as another embodiment, the process of determining the transformation setting group and the process of purging the transformation combination index can be performed simultaneously. Alternatively, step S1010 can be already set and omitted in the encoding device 100 and / or the decoding device 200.

[0129] FIG. 11 shows an example of a flowchart for encoding an AMT flag and an AMT index according to an embodiment of the present invention. The operation in FIG. 11 is performed by the conversion unit 120 of the encoding device 100.

[0130] The encoding device 100 determines whether AMT is applied to the current block (S1110).

[0131] If AMT is applied, the encoding device 100 encodes with AMT flag = 1 (S1120).

[0132] Then, the encoding device 100 determines the AMT index based on at least one of the prediction mode, horizontal transform, and vertical transform of the current block (S1130). Here, the AMT index indicates an index indicating any one of a plurality of transform combinations for each intra prediction mode, and the AMT index is transmitted for each transform unit.

[0133] When the AMT index is determined, the encoding device 100 encodes the AMT index (S1140).

[0134] On the other hand, if AMT is not applied, the encoding device 100 encodes with AMT flag = 0 (S1150).

[0135] FIG. 12 shows an example of a flowchart for decoding for performing a transform based on the AMT flag and the AMT index.

[0136] The decoding device 200 parses the AMT flag from the bitstream (S1210). Here, the AMT flag indicates whether AMT is applied to the current block.

[0137] The decoding device 200 checks whether AMT is applied to the current block based on the AMT flag (S1220). For example, it checks whether the AMT flag is 1.

[0138] When the AMT flag is 1, the decoding device 200 parses the AMT index (S1230). Here, the AMT index means an index indicating any one of a plurality of conversion combinations for each intra prediction mode, and the AMT index can be transmitted for each conversion unit. Alternatively, the AMT index means an index indicating any one of the conversion combinations defined in the already set conversion combination table. Although FIG. 8 may illustrate the already set conversion combination table, the present invention is not limited thereto.

[0139] The decoding device 200 derives or determines horizontal conversion and vertical conversion based on at least one of the AMT index or the prediction mode (S1240).

[0140] Alternatively, the decoding device 200 derives the conversion combination corresponding to the AMT index. For example, the decoding device 200 derives or determines the horizontal conversion and the vertical conversion corresponding to the AMT index.

[0141] On the other hand, when the AMT flag is 0, the decoding device 200 applies the already set inverse vertical conversion for each column (S1250). For example, the inverse vertical conversion can be an inverse conversion of DCT-2.

[0142] Then, the decoding device 200 applies the already set inverse horizontal conversion for each row (S1260). For example, the inverse horizontal conversion can be an inverse conversion of DCT-2. That is, when the AMT flag is 0, the already set conversion kernel in the encoding device 100 or the decoding device 200 is used. For example, instead of being defined in a conversion combination table as shown in FIG. 8, a frequently used conversion kernel may be used.

[0143] NSST (Non-Separable Secondary Transform)

[0144] The secondary transformation means applying the transformation kernel again with the result of applying the primary transformation as the input. The primary transformation includes DCT-2, DST-7 in HEVC, and the AMT described above. The non-separable transform does not sequentially apply the N×N transformation kernel to the row direction and the column direction, but regards the N×N two-dimensional residual block as an N 2 ×1 vector, and then applies the N 2 ×N 2 transformation kernel only once.

[0145] That is, NSST refers to a non-separable square matrix applied to a vector composed of the coefficients of the transformation block. Also, the embodiments of this document will mainly describe NSST as an example of the non-separable transform applied to the upper left region (low-frequency region) determined by the block size. However, the embodiments of the present invention are not limited to the term NSST, and any type of non-separable transform can be applied to the embodiments of the present invention. For example, the non-separable transform applied to the upper left region (low-frequency region) determined by the block size is called LFNST (Low Frequency Non-Separable Transform). In this document, the M×N transform (or transform matrix) means a matrix composed of M rows and N columns.

[0146] In NSST, after dividing the two-dimensional block data obtained by applying the primary transformation into M×M blocks, for each M×M block, M 2 ×M 2Apply a non-separable transform. The value of M can be 4 or 8. Instead of applying NSST to all regions of the two-dimensional block obtained by the primary transform, it is also possible to apply it only to some regions. For example, NSST can be applied only to the top-left 8×8 block. Also, the non-separable 64×64 transform can be applied only to the top-left 8×8 region when both the width and height of the two-dimensional block obtained by the primary transform are 8 or more. In the remaining cases, the 16×16 non-separable transform can be applied separately to each of the divided 4×4 blocks.

[0147] M 2 ×M 2 The non-separable transform can also be applied in the form of a matrix product, but for reducing the computational complexity and memory requirements, it can be approximated by a combination of a Givens rotation layer and a permutation layer. FIG. 13 shows one Givens rotation. It can be explained by one angle of one Givens rotation as shown in FIG. 13.

[0148] FIGS. 13 and 14 are embodiments to which the present invention is applied. FIG. 13 shows a diagram for explaining a Givens rotation, and FIG. 14 shows the configuration of one round in a 4×4 NSST composed of a Givens rotation layer and a permutation.

[0149] Both the 8×8 NSST and the 4×4 NSST can be composed of a hierarchical combination of Givens rotations. The matrix corresponding to one Givens rotation is as shown in Equation 1, and when representing the matrix product as a diagram, it is as shown in FIG. 13.

[0150] <Equation 1>

Number

[0151] <Equation 2>

Number

[0152] As a final step, one final replacement is further performed on the data output through the Givens rotation layer, and the information regarding the replacement is stored separately for each transformation. This replacement is performed at the end of the forward NSST, and for the inverse NSST, the inverse replacement is applied first.

[0153] The inverse NSST performs the Givens rotation layer and replacement applied in the forward NSST in reverse order, and rotates by taking the negative (-) value for the angle of each Givens rotation.

[0154] FIG. 15 shows an example of the configuration of a set of non-separable transforms for each intra prediction mode according to an embodiment of the present invention.

[0155] Intra prediction modes to which the same NSST or set of NSSTs is applied can form a group. FIG. 15 classifies 67 intra prediction modes into 35 groups. For example, mode 20 and mode 48 both belong to group 20 (hereinafter, the mode group).

[0156] For each mode group, a plurality of NSSTs other than one NSST can be configured as a set. Each set includes the case where no NSST is applied. For example, when three different NSSTs can be applied to one mode group, it can be configured to select one of four cases including the case where no NSST is applied. Here, an index is transmitted in TU units to distinguish one of the four cases. It is also possible to configure the number of NSSTs to be different for each mode group. For example, mode groups 0 and 1 are signaled to select one of three cases each including the case where no NSST is applied.

[0157] Embodiment 1: RST Applicable to 4×4 Blocks

[0158] A non-separable transform applicable to a single 4×4 block is a 16×16 transform. That is, when the data elements constituting the 4×4 block are arranged in a column in row-first or column-first order, they become a 16×1 vector, and the non-separable transform can be applied to the 16×1 vector. The forward 16×16 transform is composed of 16 row-direction transform basis vectors. When taking the inner product of the above 16×1 vector and each transform basis vector, the transform coefficient for the transform basis vector is obtained. The process of obtaining all the transform coefficients for the 16 transform basis vectors seems to be multiplying a 16×16 non-separable transform matrix by the input 16×1 vector. The transform coefficients obtained by the matrix product have the form of a 16×1 vector, but the statistical characteristics may be different for different transform coefficients. For example, if the 16×1 transform coefficient vector is composed of elements from the 0th to the 15th, the variance of the 0th element may be larger than the variance of the 15th element. That is, the more forward the element is, the larger the variance value can be and the larger the energy value it can have.

[0159] When applying the inverse 16×16 non-separable transform to the 16×1 transform coefficient, the original 4×4 block signal can be restored (when ignoring the effects such as quantization and integerization calculations). If the forward 16×16 non-separable transform is an orthonormal transform, the inverse 16×16 transform can be obtained by taking the transpose of the matrix for the forward 16×16 transform. Briefly, multiplying the inverse 16×16 non-separable transform matrix by the 16×1 transform coefficient vector gives data in the form of a 16×1 vector. When arranging it in the row-first or column-first order applied initially, the 4×4 block signal can be restored.

[0160] As described above, the elements forming the 16×1 conversion coefficient vector may have different statistical characteristics. As in the previous example, when the conversion coefficients arranged in the front (close to the 0th element) have larger energy, even if the inverse conversion is applied to some of the conversion coefficients that appear first without using all the conversion coefficients, a signal quite close to the original signal can be restored. For example, assuming that the inverse 16×16 non-separable conversion is composed of 16 column basis vectors, an L×16 matrix is formed by leaving only L column basis vectors, and after leaving only the more important L conversion coefficients from among the conversion coefficients (L×1 vector, which can appear first as in the previous example), multiplying the 16×L matrix by the L×1 vector can restore a 16×1 vector with a small error from the original input 16×1 vector data. As a result, since only L coefficients intervene in data restoration, an L×1 conversion coefficient vector instead of a 16×1 conversion coefficient vector may be obtained when obtaining the conversion coefficients. That is, in the forward 16×16 non-separable conversion matrix, after selecting L of the row-direction conversion vectors to form an L×16 conversion and multiplying by a 16×1 input vector, L important conversion coefficients can be obtained.

[0161] Embodiment 2: Application Region Setting of 4×4 RST and Arrangement of Transformation Coefficients

[0162] 4×4 RST can be applied as a secondary conversion, and at this time, it can be applied secondarily to the block to which a primary conversion such as DCT-type2 is applied. When the size of the block to which the primary conversion is applied is N×N, it is usually larger than 4×4. Therefore, when applying 4×4 RST to an N×N block, the following two methods can be considered.

[0163] 1) Instead of applying 4×4 RST to the entire N×N region, it can be applied only to a part of the region. For example, it can be applied only to the top-left M×M region (M <= N).

[0164] 2) After dividing the region to which the secondary conversion is applied into 4×4 blocks, 4×4 RST can be applied to each divided block.

[0165] The above methods 1) and 2) can be applied in combination. For example, after dividing only the upper left M×M region into 4×4 blocks, 4×4 RST can be applied.

[0166] As a specific embodiment, secondary conversion is applied only to the upper left 8×8 region. When the N×N block is larger than or equal to 8×8, 8×8 RST is applied. When the N×N block is smaller than 8×8 (4×4, 8×4, 4×8), after dividing it into 4×4 blocks as in 2) above, 4×4 RST can be applied to each of them.

[0167] After applying 4×4 RST, if L transform coefficients (1 <= L < 16) are generated, there is a degree of freedom regarding how to arrange the L transform coefficients. However, when reading and processing the transform coefficients in the residual coding part, since there is a defined order, the coding performance may vary depending on how the above L transform coefficients are arranged in a 2-dimensional block. Residual coding in the HEVC (High Efficiency Video Coding) standard starts coding from the position farthest from the DC position at the DC position. This is to improve the coding performance by taking advantage of the fact that the value of the coefficient after quantization is 0 or close to 0 as the distance from the DC position increases. Therefore, arranging the L transform coefficients so that they have high energy and the more important coefficients are coded later in the order of residual coding may be advantageous in terms of coding performance.

[0168] FIG. 16 shows three forward scan orders for the transform coefficients or transform coefficient blocks applied in the HEVC standard, where (a) is a diagonal scan, (b) is a horizontal scan, and (c) is a vertical scan.

[0169] FIG. 16 illustrates three forward scan orders for transform coefficients or transform coefficient blocks (4×4 blocks, Coefficient Group (CG)) applied in the HEVC standard, and the residual coding is performed in the reverse order of the scan order of (a), (b), or (c) (i.e., coded in the order from 16 to 1). Since the three scan orders shown in (a), (b), and (c) are selected according to the intra-prediction mode, it can be configured to determine the scan order according to the intra-prediction mode in the same way for the above L transform coefficients.

[0170] The L value has a range of 1 <= L < 16. Generally, L out of 16 transform basis vectors can be selected in any way. However, from the viewpoints of encoding and decoding, it may be advantageous from the viewpoint of encoding efficiency to select transform basis vectors that are more important in terms of signal energy as in the example presented above.

[0171] FIGS. 17 and 18 are embodiments to which the present invention is applied. FIG. 17 shows the positions of transform coefficients when forward diagonal scanning is applied when 4×4 RST is applied to a 4×8 block, and FIG. 18 shows an example when the valid transform coefficients of two 4×4 blocks are merged into one block.

[0172] When the upper left 4×8 block is divided into 4×4 blocks and 4×4 RST is applied according to the diagonal scan order of (a), when the L value is 8 (i.e., only 8 out of 16 transform coefficients are left), the transform coefficients are located as shown in FIG. 17, but only half of each 4×4 block can have transform coefficients, and a default value of 0 is filled (padded) at the positions where X is shown. Therefore, assuming that L transform coefficients are arranged for each 4×4 block according to the scan order presented in (a) and the remaining (16 - L) positions of each 4×4 block are filled with 0, the residual coding (e.g., residual coding in HEVC) can be applied.

[0173] Also, as shown in FIG. 18, L transform coefficients arranged in two 4×4 blocks can be configured into one block. In particular, when the L value is 8, since the transform coefficients of two 4×4 blocks completely fill one 4×4 block, no transform coefficients remain in other blocks. Therefore, for a 4×4 block with empty transform coefficients, residual coding is unnecessary. In the case of HEVC, a flag (coded_sub_block_flag) indicating whether application of residual coding for the block is possible is coded as 0. The combination methods for the positions of the transform coefficients of two 4×4 blocks are diverse. For example, they can be combined in any order, but the following methods may also be applied.

[0174] 1) Combine the transform coefficients of two 4×4 blocks alternately in scan order. That is, in FIGS. 8A, 8B, and 8C, when the transform coefficient for the upper block is TIFF2025083568000004.tif556 and the transform coefficient for the lower block is TIFF2025083568000005.tif552, they can be combined alternately one by one as in TIFF2025083568000006.tif568. Also, TIFF2025083568000007.tif55 and TIFF2025083568000008.tif55 can be arranged in a changed order TIFF2025083568000009.tif1286.

[0175] 2) The transform coefficients for the first 4×4 block can be arranged first, and then the transform coefficients for the second 4×4 block can be arranged. That is, they can be concatenated and arranged as in TIFF2025083568000010.tif667. Naturally, the order can also be changed as in TIFF2025083568000011.tif669.

[0176] Embodiment 3: Method for Coding NSST (Non-Separable Secondary Transform) Index for 4×4 RST

[0177] When 4×4RST is applied as shown in FIG. 17, values of 0 are filled from the (L + 1)-th position to the 16-th position in accordance with the transform coefficient scan order for each 4×4 block. Therefore, if there is a non-zero value among the positions from the (L + 1)-th position to the 16-th position in even one of the two 4×4 blocks, it is derived that the case is one where 4×4RST is not applied. When 4×4RST has a structure that applies a selected transform from among a set of transforms prepared like JEM (Joint Experiment Model) NSST, an index (hereinafter referred to as the NSST index) for what transform to apply is signaled.

[0178] In a certain decoder, the NSST index is understood by bit stream parsing, and the bit stream parsing can be performed after residual coding. In this case, if there is a non-zero transform coefficient between the (L + 1)-th position and the 16-th position by residual decoding, the decoder surely knows that 4×4RST is not applied, and thus does not parse the NSST index. Therefore, by selectively parsing the NSST index only when necessary, the signaling cost is reduced.

[0179] As shown in FIG. 17, when 4×4 RST is applied to a plurality of 4×4 blocks within a specific region (at this time, the same 4×4 RST may be applied to all of them, or different 4×4 RSTs may be applied to each of them), the (same or different) 4×4 RSTs applied to all 4×4 blocks are specified by one NSST index. Since it is determined whether 4×4 RST and the application of 4×4 RST can be performed on all 4×4 blocks by one NSST index, as a result of investigating whether there are non-zero conversion coefficients at positions from the (L + 1)-th to the 16-th positions for all 4×4 blocks during the residual decoding process, if there are non-zero conversion coefficients at positions not allowed in the 4×4 block (from the (L + 1)-th position to the 16-th position), the encoding device 100 can be set not to code the NSST index.

[0180] The encoding device 100 can also separately signal each NSST index for the luminance block and the chrominance block. In the case of the chrominance block, it can also separately signal a different NSST index for the Cb component and the Cr component, or it can also use one common NSST index. When one NSST index is used, the signaling of the NSST index is also performed only once. When one NSST index is shared for the Cb component and the Cr component, the 4×4 RST indicated by the same NSST index is applied. In this case, the 4×4 RSTs for the Cb component and the Cr component themselves may be the same, or although the NSST index is the same, individual 4×4 RSTs may be set for the Cb component and the Cr component. When the NSST index shared for the Cb component and the Cr component is used, it is checked whether there are non-zero conversion coefficients at positions from the (L + 1)-th to the 16-th positions for all 4×4 blocks for the Cb component and the Cr component for the above-mentioned conditional signaling. If non-zero conversion coefficients are found at positions from the (L + 1)-th to the 16-th positions, the signaling for the NSST index may be omitted.

[0181] As shown in FIG. 18, even when the transform coefficients for two 4×4 blocks are merged into one 4×4 block, when the 4×4 RST is applied, the encoding device 100 can determine whether it is possible to signal the NSST index after checking whether a non-zero transform coefficient appears at a position where there is no valid transform coefficient. In particular, since the L value is 8 as shown in FIG. 18, when there is no valid transform coefficient in one 4×4 block when the 4×4 RST is applied (the block indicated by X in FIG. 18(b)), a flag (coded_sub_block_flag) regarding whether the residual coding of the block is possible is checked, and if it is 1, the NSST index can be set not to be signaled. As described above, in the following description, NSST is mainly described as an example of non-separable transform, but other known terms (for example, LFNST) may be used for non-separable transform. For example, the NSST set and the NSST index may be used in place of the LFNST set and the LFNST index. Also, the RST described in this document can be used in place of LFNST as an example of non-separable transform (for example, LFNST) that uses a non-square (rectangular) transform matrix having a reduced input length and / or a reduced output length in a square non-separable transform matrix applied to at least a part of the transform block (the remaining area excluding the lower right 4×4 area in the upper left 4×4, 8×8 area or 8×8 block).

[0182] Embodiment 4: Optimization Method for the Case of Performing Coding for 4×4 Index Before Residual Coding

[0183] When coding for the NSST index is performed before residual coding, since it is determined in advance whether 4×4RST can be applied, residual coding for positions filled with conversion coefficients of 0 may be omitted. Here, whether 4×4RST can be applied can be configured to be determined by the NSST index value (for example, when the NSST index is 0, 4×4RST is not applied), or it can also be signaled whether 4×4RST can be applied by a separate syntax element (for example, an NSST flag). For example, assuming that the separate syntax element is an NSST flag, the decoding device 200 parses the NSST flag first to determine whether 4×4RST can be applied, and then, when the NSST flag value is 1, for positions where no valid conversion coefficient can exist as described above, residual coding (decoding) can be omitted.

[0184] In the case of HEVC, during the execution of residual coding, coding is performed at the position of the last non-zero coefficient in the first TU. Coding for the NSST index is performed after coding for the position of the last non-zero coefficient. When assuming the application of 4×4RST, if the position of the last non-zero coefficient is a position where no non-zero coefficient can exist, the decoding device 200 can be set not to apply 4×4RST without coding the NSST index. For example, in the case of the position indicated as X in FIG. 17, when 4×4RST is applied, since no valid conversion coefficient is located (a value of 0 can be filled), if the last non-zero coefficient is located in the area indicated as X, the decoding device 200 can omit coding for the NSST index. If the last non-zero coefficient is not located in the area indicated as X, the decoding device 200 can perform coding for the NSST index.

[0185] By conditionally coding the NSST index after coding for the positions of non-zero coefficients, when it is determined whether 4×4RST can be applied, hereinafter, the part where the remaining residual coding is performed can be processed in the following two ways.

[0186] 1) When 4×4RST is not applied, general residual coding is performed. That is, coding is performed under the assumption that non-zero transform coefficients may exist at any position from the position of the last non-zero coefficient to DC.

[0187] 2) When 4×4RST is applied, since the transform coefficient does not exist for a specific position or a specific 4×4 block (for example, the X position in FIG. 17) (filled with 0 by default), residual coding may be omitted for the position or block. For example, when reaching the position marked as X while scanning in accordance with the scan order in FIG. 17, coding for the flag (sig_coeff_flag) regarding whether a non-zero coefficient exists at the position in the HEVC standard can be omitted. When the transform coefficients of two blocks are merged into one block as shown in FIG. 18, coding for the flag (for example, code_sub_block_flag in the HEVC standard) indicating whether residual coding of the 4×4 block filled with 0 is possible is omitted, and the corresponding value can be derived as 0, and the corresponding 4×4 block can be filled (charged) with all 0 values without separate coding.

[0188] When coding the NSST index after coding for the position of the last non-zero coefficient, if the x-position (Px) and y-position (Py) of the last non-zero coefficient are each smaller than Tx and Ty, the coding of the NSST index can be omitted and the 4×4RST can be set not to be applied. For example, when Tx = 1, Ty = 1 and the last non-zero coefficient exists at the position of DC, the NSST index coding is omitted. The method for determining whether NSST index coding is possible for comparison with such a threshold can be applied differently to the luminance component and the chrominance components. For example, different Tx and Ty may be applied to the luminance component and the chrominance components respectively, or a threshold may be applied to the luminance component and no threshold may be applied to the chrominance components. Conversely, a threshold may be applied to the chrominance components and no threshold may be applied to the luminance component.

[0189] The two methods described above (when the last non-zero coefficient is located in an area where no valid transform coefficient exists, omit the NSST index coding; when the X coordinate and Y coordinate for the last non-zero coefficient are each smaller than the threshold, omit the NSST index coding) may be applied simultaneously. For example, after first checking the threshold for the position coordinates of the last non-zero coefficient, it is possible to check whether the last non-zero coefficient is located in an area where no valid transform coefficient exists, and the order of the two methods can be changed.

[0190] The method presented in Embodiment 4) can also be applied to 8×8RST. That is, when the last non-zero coefficient is located in the upper left 4×4 non-area within the upper left 8×8 area, the coding for the NSST index can be omitted, and otherwise, the coding for the NSST index can be performed. Also, when the values of the X and Y coordinates for the position of the last non-zero coefficient are all less than a certain threshold, the coding for the NSST index can be omitted. The two methods can also be applied simultaneously.

[0191] Embodiment 5: Applying Different NSST Index Coding and Residual Coding Methods to Luminance Component and Chrominance Component Respectively When Applying RST

[0192] In Embodiment 3 and Embodiment 4, the methods described can be applied differently to the luminance component and the chrominance components. That is, the NSST index coding and the residual coding method can be applied differently to the luminance component and the chrominance components. For example, the method described in Embodiment 4 can be applied to the luminance component, and the method described in Embodiment 3 can be applied to the chrominance components. Also, conditional NSST index coding proposed in Embodiment 3 or Embodiment 4 can be applied to the luminance component, and conditional NSST index coding may not be applied to the luminance component, and vice versa (conditional NSST index coding is applied to the chrominance components and not to the luminance component) is also possible.

[0193] Embodiment 6

[0194] In one embodiment of the present invention, a mixed NSST transform set (MNTS) for applying various NSST conditions in the process of applying NSST and a method of constructing the corresponding MNTS are provided.

[0195] According to JEM, depending on the size of the pre-selected sub-block, the 4×4 NSST set includes only 4×4 kernels, and the 8×8 NSST set includes only 8×8 kernels. Embodiments of the present invention additionally propose a method of constructing a mixed NSST set as follows.

[0196] - The size of the NSST kernel available in the NSST set is not fixed, and an NSST kernel having one or more variable sizes can be included in the NSST set (for example, both a 4×4 NSST kernel and an 8×8 NSST kernel are included in one NSST set).

[0197] - The number of NSST kernels available within an NSST set can be variable rather than fixed (e.g., the first set contains 3 kernels and the second set contains 4 kernels).

[0198] - The order of NSST kernels may be defined such that it is not fixed and varies by NSST set (e.g., in the first set, NSST kernels 1, 2, 3 are mapped to NSST indices 1, 2, 3 respectively, while in the second set, NSST kernels 3, 2, 1 are mapped to NSST indices 1, 2, 3 respectively).

[0199] More specifically, an example of a method for constructing a mixed NSST transform set is as follows.

[0200] - The priority of NSST kernels available for use in an NSST transform set can be determined by the size of the NSST kernels (e.g., 4×4 NSST and 8×8 NSST).

[0201] For example, when the block is large, the 8×8 NSST kernel may be more important than the 4×4 NSST kernel, so an NSST index with a lower value is assigned to the 8×8 NSST kernel.

[0202] - The priority of NSST kernels available for use in an NSST transform set can be determined by the order of the NSST kernels.

[0203] For example, a given 4×4 NSST first kernel may have priority over a 4×4 NSST second kernel.

[0204] Since the NSST index is encoded and transmitted, by assigning a higher priority (smaller index) to more frequently occurring NSST kernels, the NSST index can be signaled with fewer bits.

[0205] Tables 1 and 2 below show examples of the mixed NSST sets proposed in this embodiment.

[0206]

Table 1

[0207]

Table 2

[0208] Embodiment 7

[0209] In one embodiment of the present invention, in the process of determining the secondary conversion set, a method for determining the NSST set is proposed by considering the intra prediction mode and the block size.

[0210] The method proposed in this embodiment is related to Embodiment 6 to configure a conversion set adapted to the intra prediction mode so that kernels of various sizes can be configured and applied to blocks.

[0211] FIG. 19 is an embodiment to which the present invention is applied, and shows an example of a method for configuring a mixed NSST set for each intra prediction mode.

[0212] FIG. 19 is an example of a table according to a method of applying the method proposed in Embodiment 2 in relation to Embodiment 6. That is, as shown in FIG. 19, an index (“Mixed Type”) is defined to indicate whether to follow the existing NSST set configuration method for each intra prediction mode or the NSST set configuration method of another method.

[0213] More specifically, in the case of the intra prediction mode where the index ("Mixed Type") is defined as "1" in FIG. 19, regardless of the JEM NSST set configuration method, the NSST set is configured using the NSST set configuration method defined in the system. Here, the NSST set configuration method defined in the system means the mixed NSST set proposed in Embodiment 6.

[0214] As yet another embodiment, the table in FIG. 19 describes two types of transform set configuration methods (JEM-based NSST set configuration, the mixed type NSST set configuration method proposed in the embodiments of the present invention) based on mixed type information (flags) related to the intra prediction mode. However, there may be one or more mixed type NSST configuration methods, and here, the mixed type information can be expressed as N (N>2) various values.

[0215] As yet another embodiment, it is possible to determine whether to configure a transform set suitable for the current block in a mixed type considering both the intra prediction mode and the size of the transform block. For example, when the mode type corresponding to the intra prediction mode is 0, the NSST set of JEM is set accordingly; otherwise, (Mode Type = 1) various mixed type NSST sets can be determined according to the size of the transform block.

[0216] FIG. 20 shows an example of a method for selecting an NSST set (or kernel) considering the intra prediction mode and the size of the transform block in an embodiment to which the present invention is applied.

[0217] When the transform set is determined, the decoding device 200 can determine the NSST kernel used by using the NSST index information.

[0218] Embodiment 8

[0219] In one embodiment of the present invention, when constructing a transform set considering the intra prediction mode and the block size during the application of the secondary transform, a method for efficiently encoding the NSST index is provided by considering the change in the statistical distribution of the NSST index values transmitted after encoding. Embodiments of the present invention provide a method for selecting a kernel applied using a syntax indicating the kernel size.

[0220] Also, in an embodiment of the present invention, since the number of available NSST kernels differs for each transform set, for an efficient binary evolution method, a truncated unary binary evolution method as shown in Table 3 below is provided according to the maximum NSST index value available for each set.

[0221]

Table 3

[0222] Table 3 shows the binary evolution method of the NSST index value. Since the number of available NSST kernels differs for each transform set, the NSST index can be binary evolved by the maximum NSST index value.

[0223] Embodiment 9: Reduced Transform

[0224] Due to complexity issues in the transform (e.g., large block transform or non-separable transform), a reduced transform applicable to the core transform (e.g., DCT, DST, etc.) and the secondary transform (e.g., NSST) is provided.

[0225] The main idea of the reduced transform is to map an N-dimensional vector to an R-dimensional vector from another space, where R / N (R < N) is the reduction factor. The reduced transform is an R×N matrix as shown in Equation 3 below.

[0226] <Equation 3>

Mathematics

[0227] In Equation 1, the R rows of the transformation are the R bases of the new N-dimensional space. Therefore, the reason it is called a reduced transformation is that the number of elements of the vector output by the transformation is smaller than the number of elements of the input vector (R < N). The inverse transform matrix for the reduced transformation is the transpose of the forward transformation. The reduced transformations in the forward and reverse directions will be described with reference to FIGS. 13A and 13B.

[0228] FIGS. 21A and 21B are embodiments to which the present invention is applied and show reduced transformations in the forward and reverse directions.

[0229] The number of elements of the reduced transformation is RxN, which is smaller than the size of the complete matrix (N×N) by R / N, which means that the required memory is R / N of the complete matrix.

[0230] Also, the number of required multiplications is also R×N, which is less than the original N×N by R / N.

[0231] When X is a vector of N levels, R coefficients are obtained after applying the reduced transformation, which means that only R values need to be transmitted instead of the original N coefficients.

[0232] FIG. 22 shows an example of a flowchart of decoding using a reduced transformation according to an embodiment of the present invention.

[0233] The proposed reduced transformation (inverse transformation in the decoder) can be applied to the coefficients (inverse quantized coefficients), as shown in FIG. 21. A pre-determined reduction factor (R, or R / N) and a transformation kernel for performing the transformation may be required. Here, the transformation kernel can be determined based on available information such as block size (width, height), intra prediction mode, Cidx. When the current coding block is a luma block, CIdx is 0. Otherwise (Cb or Cr block), CIdx has a non-zero value, such as 1.

[0234] Hereinafter, the operators used in this document are defined as shown in Tables 4 and 5 below

[0235]

Table 4

[0236]

Table 5

[0237] FIG. 23 shows an example of a flowchart of the application of the conditionally reduced transformation according to an embodiment of the present invention. The operations in FIG. 23 are performed by the inverse quantization unit 140 and the inverse transformation unit 150 of the decoding apparatus 200.

[0238] In one embodiment, the reduced transformation can be used when certain conditions are met. For example, the reduced transformation can be applied to blocks larger than a certain size as follows.

[0239] - Width > TH && Height > HT (where TH is a pre-defined value (e.g., 4))

[0240] Or

[0241] - Width * Height > K && MIN(Width, Height) > TH (where K and TH are predefined values)

[0242] That is, when the width of the current block is greater than a predefined value (TH) and the height of the current block is greater than the predefined value (TH) as in the above condition, a reduced transformation can be applied. Alternatively, when the product of the width and height of the current block is greater than a predefined value (K) and the smaller value of the width and height of the current block is greater than the predefined value (TH), a reduced transformation can be applied.

[0243] The reduced transformation can be applied to a predefined group of blocks as follows.

[0244] - Width == TH && Height == TH

[0245] Or

[0246] - Width == Height

[0247] That is, when the width and height of the current block are each the same as a predefined value (TH), or when the width and height of the current block are the same (when the current block is a square block), a reduced transformation can be applied.

[0248] If the conditions for using the reduced transformation are not satisfied, a regular transformation is applied. The regular transformation can be a predefined and available transformation in a video coding system. Examples of regular transformations are as follows.

[0249] - DCT-2, DCT-4, DCT-5, DCT-7, DCT-8

[0250] Or

[0251] -DST-1, DST-4, DST-7

[0252] or

[0253] -Non-separable transform

[0254] or

[0255] -JEM-NSST(HyGT)

[0256] As shown in FIG. 23, the reduced conversion conditions depend on an index (Transform_idx) indicating which conversion (e.g., DCT-4, DST-1) is used or which kernel is applied (when multiple kernels are available). In particular, Transform_idx can be transmitted twice. One is the index indicating the horizontal conversion (Transform_idx_h), and the other is the index indicating the vertical conversion (Transform_idx_v).

[0257] More specifically, referring to FIG. 23, the decoding device 200 performs inverse quantization on the input bitstream (S2305). Then, the decoding device 200 determines whether to apply a conversion (S2310). The decoding device 200 determines whether to apply a conversion based on a flag indicating whether to skip the conversion.

[0258] When a conversion is applied, the decoding device 200 parses a conversion index (Transform_idx) indicating the applied conversion (S2315). Also, the decoding device 200 selects a conversion kernel (S2330). For example, the decoding device 200 selects a conversion kernel corresponding to the conversion index (Transform_idx). Also, the decoding device 200 selects a conversion kernel in consideration of the block size (width, height), the intra prediction mode, and CIdx (luma, chroma).

[0259] The decoding device 200 determines whether or not to satisfy the conditions for applying the reduced transformation (S2320). The conditions for applying the reduced transformation include the conditions as described above. When the reduced transformation is not applied, the decoding device 200 applies a normal inverse transformation (S2325). For example, the decoding device 200 determines an inverse transformation matrix from the transformation kernel selected in step S2330, and applies the determined inverse transformation matrix to the current block including the transformation coefficients.

[0260] When the reduced transformation is applied, the decoding device 200 applies a reduced inverse transformation (S2335). For example, the decoding device 200 determines a reduced inverse transformation matrix in consideration of a reduction factor from the transformation kernel selected in step S2330, and applies the reduced inverse transformation matrix to the current block including the transformation coefficients.

[0261] FIG. 24 shows an example of a decoding flowchart for a secondary inverse transformation to which a conditionally reduced transformation according to an embodiment of the present invention is applied. The operation of FIG. 24 is performed by the inverse transformation unit 230 of the decoding device 200.

[0262] In one embodiment, the reduced transformation can be applied to the secondary transformation as shown in FIG. 24. When the NSST index is purged, the reduced inverse transformation can be applied.

[0263] Referring to FIG. 24, the decoding device 200 performs inverse quantization (S2405). For the transformation coefficients generated by the inverse quantization, the decoding device 200 determines whether or not to apply NSST (S2410). That is, the decoding device 200 determines whether or not it is necessary to purge the NSST index (NSST_idx) depending on whether or not to apply NSST.

[0264] When NSST is applied, the decoding device 200 parses the NSST index (S2415) and determines whether the NSST index is greater than 0 (S2420). The NSST index is restored by a technique such as CABAC by the entropy decoding unit 210. When the NSST index is 0, the decoding device 200 omits the second inverse transform and applies the core inverse transform or the first inverse transform (S2445).

[0265] Also, when NSST is applied, the decoding device 200 selects a transform kernel for the second inverse transform (S2435). For example, the decoding device 200 selects a transform kernel corresponding to the NSST index (NSST_idx). Also, the decoding device 200 selects a transform kernel in consideration of the block size (width, height), the intra prediction mode, and CIdx (luma, chroma).

[0266] When the NSST index is greater than 0, the decoding device 200 determines whether the conditions for applying the reduced transform are satisfied (S2425). The conditions for applying the reduced transform include the conditions as described above. When the reduced transform is not applied, the decoding device 200 applies the normal second inverse transform (S2430). For example, the decoding device 200 determines a second inverse transform matrix from the transform kernel selected in step S2435 and applies the determined second inverse transform matrix to the current block including the transform coefficients.

[0267] When the reduced transform is applied, the decoding device 200 applies the reduced second inverse transform (S2440). For example, the decoding device 200 can determine a reduced inverse transform matrix in consideration of the reduction factor from the transform kernel selected in step S2335 and apply the reduced inverse transform matrix to the current block including the transform coefficients. Thereafter, the decoding device 200 applies the core inverse transform or the first inverse transform (S2445).

[0268] Embodiment 10: Reduced Transform as a Secondary Transform for Different Block Sizes

[0269] Figures 25A, 25B, 26A, and 26B show examples of reduced transformation and reduced inverse transformation according to embodiments of the present invention.

[0270] In one embodiment of the present invention, in a video codec for different block sizes such as 4×4, 8×8, 16×16, etc., the reduced transformation can be used as a secondary transformation and a secondary inverse transformation. As an example regarding the 8×8 block size and the reduction factor R = 16, the secondary transformation and the secondary inverse transformation can be set as shown in FIGS. 25A and 25B.

[0271] The pseudocode of the reduced transformation and the reduced inverse transformation is set as shown in FIG. 26.

[0272] [Table 6]

[0273] [Table 7]

[0274] Embodiment 11: Reduced Transform as a Secondary Transform with Non-Rectangular Shape

[0275] FIG. 27 shows an example of a region to which the reduced secondary transformation according to an embodiment of the present invention is applied.

[0276] As described above, due to the complexity argument in the secondary transformation, the secondary transformation can be applied to the 4×4 and 8×8 corners. The reduced transformation can also be applied to non-square shapes.

[0277] As shown in FIG. 27, RST can be applied only to a partial area (diagonal area) of a block. In FIG. 27, each square represents a 4×4 area, and RST is applied to 10 4×4 pixels (i.e., 160 pixels). When the reduction factor R = 16, the entire RST matrix is a 16×16 matrix, which can be an acceptable computational load.

[0278] Embodiment 12: Reduction Factor

[0279] FIG. 28 shows a reduction by a reduction factor according to an embodiment of the present invention.

[0280] Changing the reduction factor can change the memory and multiplication complexity. As described above, changing the reduction factor reduces the memory and multiplication complexity by a factor of R / N. For example, for 8×8 NSST, when R = 16, the memory and multiplication complexity are reduced by a factor of 1 / 4.

[0281] Embodiment 13: High Level Syntax

[0282] The following syntax elements are used to process RST in video coding. The semantics related to the reduced transform are present in the SPS (Sequence Parameter Set) or the slice header.

[0283] A Reduced_transform_enabled_flag value of 1 indicates that reduced transform is enabled and applied. A Reduced_transform_enabled_flag value of 0 indicates that reduced transform is not enabled. When Reduced_transform_enabled_flag is not present, it is inferred to be equal to 0 (Reduced_transform_enabled_flag equals to 1 specifies that reduced transform is enabled and applied. Reduced_transform_enabled_flag equal to 0 specifies that reduced transform is not enabled. When Reduced_transform_enabled_flag is not present, it is inferred to be equal to 0).

[0284] Reduced_transform_factor specifies the number of reduced dimensions to keep for reduced transform. When Reduced_transform_factor is not present, it is inferred to be equal to R (Reduced_transform_factor specifies that the number of reduced dimensions to keep for reduced transform. When Reduced_transform_factor is not present, it is inferred to be equal to R).

[0285] min_reduced_transform_size indicates the minimum transform size for applying the reduced transform. If min_reduced_transform_size does not exist, it is inferred to be 0 (min_reduced_transform_size specifies that the minimum transform size to apply reduced transform. When min_reduced_transform_size is not present, it is inferred to be equal to 0).

[0286] max_reduced_transform_size indicates the maximum transform size for applying the reduced transform. If max_reduced_transform_size does not exist, it is inferred to be 0.

[0287] reduced_transform_size indicates the number of reduced dimensions to maintain for the reduced transform. If reduced_transform_size does not exist, it is inferred to be 0 (reduced_transform_size specifies that the number of reduced dimensions to keep for reduced transform. When Reduced_transform_factor is not present, it is inferred to be equal to 0.).

[0288]

Table 8

[0289] Embodiment 14: Conditional Application of 4×4 RST for Worst Case Handling

[0290] The non-separable second-order transform (4×4 NSST) applicable to a 4×4 block is a 16×16 transform. The 4×4 NSST is applied secondarily to a block to which a primary transform such as DCT-2, DST-7, or DCT-8 has been applied. Assuming the size of the block to which the primary transform has been applied is N×M, when applying the 4×4 NSST to an N×M block, the following methods can be considered.

[0291] 1) The conditions for applying the 4×4 NSST to an N×M area are as follows in a) and b) below.

[0292] a) N >= 4

[0293] b) M >= 4

[0294] 2) The 4×4 NSST is not applied to the entire N×M area, but only to a partial area. For example, the 4×4 NSST can be applied only to the upper-left K×J area. The conditions for this case are as follows in a) and b) below.

[0295] a) K >= 4

[0296] b) J >= 4

[0297] 3) After dividing the area to which the second-order transform is applied into 4×4 blocks, the 4×4 NSST can be applied to each divided block.

[0298] The computational complexity of the 4×4 NSST is a very important factor to be considered in the encoder and decoder, so it will be analyzed in detail. In particular, the computational complexity of the 4×4 NSST is analyzed based on the number of multiplications. In the case of the forward NSST, the 16×16 secondary transformation is composed of 16 row-direction transformation basis vectors. When taking the inner product of a 16×1 vector with each transformation basis vector, the transformation coefficients for the corresponding transformation basis vector are obtained. The process of obtaining all the transformation coefficients for the 16 transformation basis vectors seems to be equivalent to multiplying a 16×16 non-separable transformation matrix by an input 16×1 vector. Therefore, the total number of multiplications required for the 4×4 forward NSST is 256.

[0299] In the decoder, when applying the inverse 16×16 non-separable transformation to the 16×1 transformation coefficients (when ignoring the effects such as quantization and integerization calculations), the coefficients of the original 4×4 primary transformation block can be restored. In other words, when multiplying the inverse 16×16 non-separable transformation matrix by the 16×1 transformation coefficient vector, data in the form of a 16×1 vector is obtained. When arranging the data according to the row-first or column-first order applied initially, the 4×4 block signal (primary transformation coefficients) can be restored. Therefore, the total number of multiplications required for the 4×4 inverse NSST is 256.

[0300] As described above, when the 4×4 NSST is applied, the number of multiplications required per sample is 16. This is the number obtained by dividing the total number of multiplications 256, which is obtained in the inner product process between the 16×1 vector and each transformation basis vector during the execution of the 4×4 NSST, by the total number of samples 16. The number of multiplications required identically for both the forward 4×4 NSST and the inverse 4×4 NSST is 16.

[0301] In the case of an 8×8 block, the number of multiplications per sample required when applying the 4×4 NSST is determined as follows according to the area where the 4×4 NSST is applied.

[0302] 1. When the 4×4 NSST is applied only to the upper left 4×4 region: 256 (the number of multiplications required in the 4×4 NSST process) / 64 (the total number of samples in the 8×8 block) = 4 multiplications / sample

[0303] 2. When the 4×4 NSST is applied to the upper left 4×4 region and the upper right 4×4 region: 512 (the number of multiplications required in two 4×4 NSST processes) / 64 (the number of samples in the 8×8 block) = 8 multiplications / sample

[0304] 3. When the 4×4 NSST is applied to all 4×4 regions of the 8×8 block: 1024 (the number of multiplications required in four 4×4 NSST processes) / 64 (the number of samples in the 8×8 block) = 16 multiplications / sample

[0305] As described above, when the block size is large, the range to which the 4×4 NSST is applied can be reduced to decrease the number of multiplications in the worst case required for each sample.

[0306] Therefore, when using the 4×4 NSST, the worst case occurs when the size of the TU is 4×4. In this case, the methods for reducing the worst case complexity are as follows.

[0307] Method 1. Do not apply the 4×4 NSST to small TUs (i.e., 4×4 TUs).

[0308] Method 2. In the case of 4×4 blocks (4×4 TUs), apply the 4×4 RST instead of the 4×4 NSST.

[0309] In the case of Method 1, it has been experimentally observed that not applying the 4×4 NSST causes a significant decrease in the coding performance. In the case of Method 2, it has been revealed that due to the statistical characteristics of the elements constituting the 16×1 conversion coefficient vector, by applying the inverse transformation to some of the conversion coefficients located in the front even without using all the conversion coefficients, a signal very close to the original signal can be restored and most of the coding performance can be maintained.

[0310] Specifically, in the case of 4×4 RST, assuming that the 16×16 non-separable transform in the reverse direction (or forward direction) is composed of 16 column basis vectors, only L column basis vectors are left to form a 16×L matrix. By leaving only the more important L transform coefficients out of the transform coefficients, when multiplying the 16×L matrix by an L×1 vector, a 16×1 vector with a small error from the original 16×1 vector data can be restored.

[0311] As a result, since only L coefficients intervene in data restoration, instead of obtaining a 16×1 transform coefficient vector to obtain the transform coefficients, an L×1 transform coefficient vector may be obtained. That is, by selecting L row direction transform vectors in the forward 16×16 non-separable transform matrix, an L×16 transform matrix is formed, and when multiplying the L×16 transform matrix by a 16×1 input vector, L transform coefficients are obtained.

[0312] The value of L has a range of 1 <= L < 16. Generally, L out of 16 transform basis vectors can be selected in any way. However, as described above, it may be advantageous from the viewpoint of coding efficiency to select the transform basis vectors with high energy importance of the signal from the aspects of encoding and decoding. The worst-case number of multiplications per sample in a 4×4 block due to the change of the L value is as shown in Table 9 below.

[0313]

Table 9

[0314] As described above, for the reduction of the worst-case multiplication complexity, 4×4 NSST and 4×4 RST can be used in combination as shown in Table 10 below (however, the following example explains the conditions for applying 4×4 NSST and 4×4 RST under the condition for applying 4×4 NSST (that is, when the width and height of the current block are all greater than or equal to 4)).

[0315] As described above, the 4×4 NSST for a 4×4 block is a square (16x16) transformation matrix that takes 16 data inputs and outputs 16 data, and the 4×4 RST means a non-square (8×16) transformation matrix that takes 16 data inputs and outputs R data smaller than 16 (e.g., 8) based on the encoder side. Based on the decoder side, the 4×4 RST means a non-square (16×8) transformation matrix that takes R data smaller than 16 (e.g., 8) as inputs and outputs 16 data.

[0316]

Table 10

[0317] Referring to Table 10, when the width and height of the current block are both 4, a 4×4 RST based on an 8×16 matrix is applied to the current block, and otherwise (when either one of the width or height of the current block is not 4), a 4×4 NSST can be applied to the 4×4 area in the upper left corner of the current block. More specifically, when the size of the current block is 4×4, a non-separable transformation with an input length of 16 and an output length of 8 can be applied. In the case of the inverse non-separable transformation, conversely, a non-separable transformation with an input length of 8 and an output length of 16 can be applied.

[0318] As described above, for the reduction of the multiplication complexity in the worst case, the 4×4 NSST and the 4×4 RST can be used in combination as shown in Table 11 below (however, the following example explains the conditions for applying the 4×4 NSST and the 4×4 RST under the condition for applying the 4×4 NSST, i.e., when the width and height of the current block are both greater than or equal to 4).

[0319]

Table 11

[0320] Referring to Table 11, when the width and height of the current block are both 4, 4×4 RST based on an 8×16 matrix is applied. When the product of the width and height of the current block is smaller than the threshold (TH), 4×4 NSST is applied to the upper left 4×4 area of the current block. When the width of the current block is greater than or equal to the height, 4×4 NSST is applied to the upper left 4×4 area and the 4×4 area located on the right side of the upper left 4×4 area. In the remaining cases (when it is smaller than the height of the current block), 4×4 NSST is applied to the upper left 4×4 area and the 4×4 area located below the upper left 4×4 area.

[0321] As a conclusion, for reducing the computational complexity of multiplication in the worst case, 4×4 RST (e.g., an 8×16 matrix) can be applied to 4×4 blocks instead of 4×4 NSST.

[0322] Embodiment 15: Conditional Application of 8×8 RST for Worst Case Handling

[0323] The non-separable second-order transform (8×8 NSST) applicable to an 8×8 block is a 64×64 transform. 8×8 NSST is applied secondarily to a block to which a first-order transform such as DCT-2, DST-7, or DCT-8 has been applied. Assuming the size of the block to which the first-order transform has been applied is N×M, when applying 8×8 NSST to the N×M block, the following methods are considered.

[0324] 1) The conditions for applying 8×8 NSST to the N×M area are as follows in c) and d).

[0325] c) N >= 8

[0326] d) M >= 8

[0327] 2) Instead of applying 8×8 NSST to the entire N×M area, it may be applied only to some areas. For example, 8×8 NSST is applied only to the upper left K×J area. The conditions for this case are as follows in c) and d).

[0328] c) K >= 8

[0329] d) J >= 8

[0330] 3) After dividing the area to which the secondary transformation is applied into 8×8 blocks, 8×8 NSST can be applied to each divided block.

[0331] The computational complexity of 8×8 NSST is a very important factor to be considered in the encoder and decoder, so it will be analyzed in detail. In particular, the computational complexity of 8×8 NSST is analyzed based on the number of multiplications. In the case of the forward NSST, the 64×64 non-separable secondary transformation is composed of 64 row-direction transformation basis vectors. When taking the inner product of a 64×1 vector with each transformation basis vector, the transformation coefficients for the corresponding transformation basis vector are obtained. The process of obtaining all the transformation coefficients for 64 transformation basis vectors seems to be multiplying a 64×64 non-separable transformation matrix by an input 64×1 vector. Therefore, the total number of multiplications required for 8×8 forward NSST is 4096.

[0332] In the decoder, when applying the inverse 64×64 non-separable transformation to the 64×1 transformation coefficients (when ignoring the effects such as quantization and integerization calculations), the coefficients of the original 8×8 primary transformation block can be restored. In other words, when multiplying the inverse 64×64 non-separable transformation matrix by the 64×1 transformation coefficient vector, data in the form of a 64×1 vector is obtained, and when arranging the data according to the row-first or column-first order applied initially, the 8×8 block signal (primary transformation coefficients) can be restored. Therefore, the total number of multiplications required for 8×8 inverse NSST is 4096.

[0333] As described above, when 8×8 NSST is applied, the number of multiplications required per sample unit is 64. This is the number obtained by dividing the total number of multiplications 4096, which is obtained in the process of taking the inner product of the 64×1 vector, which is the 8×8 NSST execution process, and each transformation basis vector, by the number of total samples 64. The number of multiplications required identically for both the forward 8×8 NSST and the reverse 8×8 NSST is 64.

[0334] In the case of a 16×16 block, the number of multiplications per sample required when applying 8×8 NSST is determined as follows depending on the area where 8×8 NSST is applied.

[0335] 1. When 8×8 NSST is applied only to the upper left 8×8 area: 4096 (the number of multiplications required in the 8×8 NSST process) / 256 (the number of total samples in the 16×16 block) = 16 multiplications / sample

[0336] 2. When 8×8 NSST is applied to the upper left 8×8 area and the upper right 8×8 area: 8192 (the number of multiplications required in two 8×8 NSST processes) / 256 (the number of total samples in the 16×16 block) = 32 multiplications / sample

[0337] 3. When 8×8 NSST is applied to all 8×8 areas of the 16×16 block: 16384 (the number of multiplications required in four 8×8 NSST processes) / 256 (the number of total samples in the 16×16 block) = 64 multiplications / sample

[0338] As described above, when the block size is large, the range where 8×8 NSST is applied can be reduced in order to reduce the number of multiplications in the worst case required per sample.

[0339] When 8×8 NSST is applied, since the 8×8 block is the smallest TU to which 8×8 NSST can be applied, from the perspective of the number of multiplications required per sample, the case where the TU size is 8×8 corresponds to the worst case. The methods for reducing the worst case complexity are as follows.

[0340] Method 1. Do not apply 8×8 NSST to small TUs (i.e., 8×8 TUs).

[0341] Method 2. For 8×8 blocks (8×8 TUs), apply 8×8 RST instead of 8×8 NSST.

[0342] In the case of Method 1, it has been experimentally observed that not applying 8×8 NSST causes a significant decrease in coding performance. In the case of Method 2, it has been revealed that due to the statistical characteristics of the elements constituting the 64×1 conversion coefficient vector, a signal quite close to the original signal can be restored by applying inverse conversion to some of the conversion coefficients located at the front side without using all the conversion coefficients, and most of the coding performance can be maintained.

[0343] Specifically, in the case of 8×8 RST, assuming that the inverse (or forward) 64×64 non-separable transform is composed of 16 column basis vectors, only L column basis vectors are left to form a 64×L matrix. By multiplying the 64×L matrix and the L×1 vector by leaving only the more important L conversion coefficients among the conversion coefficients, a 64×1 vector with a not large error from the original 64×1 vector data can be restored.

[0344] As a result, since only L coefficients intervene in data restoration, in order to obtain the conversion coefficients, instead of a 64×1 conversion coefficient vector, an L×1 conversion coefficient vector may be obtained. That is, an L×64 conversion matrix is constructed by selecting L row-direction conversion vectors in the forward 64×64 non-separable conversion matrix, and when the L×64 conversion matrix is multiplied by the 64×1 input vector, L conversion coefficients are obtained.

[0345] The L value has a range of 1 <= L < 64. Generally, L out of 64 conversion basis vectors are selected in an arbitrary manner. However, as described above, it may be advantageous from the viewpoint of coding efficiency to select the conversion basis vectors with high energy importance of the signal from the aspects of coding and decoding. The worst-case number of multiplications per sample in an 8×8 block due to the conversion of the L value is as shown in Table 12 below.

[0346] [Table 12]

[0347] As described above, for the reduction of the worst-case multiplication complexity, 8×8 RSTs with different L values can be used in combination as shown in Table 13 below (however, the following example explains the conditions for applying 8×8 RST under the conditions for applying 8×8 NSST, that is, when the width and height of the current block are all greater than or equal to 8).

[0348] [Table 13]

[0349] Referring to Table 13, when the width and height of the current block are both 8, an 8×8 RST based on an 8×64 matrix is applied to the current block. Otherwise (when either the width or height of the current block is not 8), an 8×8 RST based on a 16×64 matrix can be applied to the current block. More specifically, when the size of the current block is 8×8, a non-separable transform with an input length of 64 and an output length of 8 is applied. Otherwise, a non-separable transform with an input length of 64 and an output length of 16 is applied. In the case of the inverse non-separable transform, when the current block is 8×8, a non-separable transform with an input length of 8 and an output length of 64 is applied. Otherwise, a non-separable transform with an input length of 16 and an output length of 64 is applied.

[0350] Table 14 is an example regarding the application of various 8×8 RSTs under the condition for applying an 8×8 NSST (that is, when the width and height of the current block are greater than 8).

[0351]

Table 14

[0352] Referring to Table 14, when the width and height of the current block are both 8, an 8×8 RST based on an 8×64 matrix is applied. When the product of the width and height of the current block is smaller than the threshold (TH), an 8×8 RST based on a 16×64 matrix is applied to the 8×8 area at the upper left side of the current block. When the width of the current block is greater than or equal to the height, an 8 RST based on a 32×64 matrix is applied to the 4×4 area located in the 8×8 area at the upper left side of the current block. In the remaining cases (when the product of the width and height of the current block is greater than or equal to the threshold and the width of the current block is smaller than the height), an 8×8 RST based on a 32×64 matrix is applied to the 8×8 area at the upper left side of the current block.

[0353] FIG. 29 shows an example of a flowchart of decoding to which the conversion according to an embodiment of the present invention is applied. The operations in FIG. 29 are performed by the inverse transform unit 230 of the decoding device 200.

[0354] In step S2905, the decoding device 200 determines the input length and output length of the non-separable transform based on the height and width of the current block. Here, when the height and width of the current block are both 4, the input length of the non-separable transform is determined to be 8 and the output length is determined to be 16. That is, the inverse transform of 4×4 RST based on an 8×16 matrix (inverse 4×4 RST in the reverse direction based on a 16×8 matrix) is applied. When it does not correspond to the case where the height and width of the current block are both 4, the input length and output length of the non-separable transform are each determined to be 16.

[0355] In step S2910, the decoding device 200 determines a non-separable transform matrix corresponding to the input length and output length of the non-separable transform. For example, when the input length of the non-separable transform is 8 and the output length is 16 (when the size of the current block is 4×4), a 16×8 matrix derived from the transform kernel is determined as the non-separable transform block. When the input length of the non-separable transform is 16 and the output length is 16 (for example, when the current block is smaller than 8×8 and not 4×4), a 16×16 transform kernel can be determined as the non-separable transform.

[0356] According to an embodiment of the present invention, the decoding device 200 determines a non-separable transform set index (for example, NSST index) based on the intra prediction mode of the current block, determines a non-separable transform kernel corresponding to the non-separable transform index within the non-separable transform set included in the non-separable transform set index, and can determine a non-separable transform matrix from the non-separable transform kernel based on the input length and output length determined in step S2905.

[0357] In step S2915, the decoding device 200 applies the non-separable transform matrix determined for the current block to the current block. For example, when the input length of the non-separable transform is 8 and the output length is 16, an 8×16 matrix derived from the transform kernel is applied to the current block. When the input length of the non-separable transform is 16 and the output length is 16, a 16×16 matrix derived from the transform kernel can be applied to the coefficients of the 4×4 area in the upper left corner of the current block.

[0358] Also, for the case where the height and width of the current block are not 4 respectively, when the product of the width and height of the current block is smaller than the threshold, the decoding device 200 applies the non-separable transform matrix to the upper left 4×4 area of the current block. When the width of the current block is greater than or equal to the height, the non-separable transform matrix is applied to the upper left 4×4 area of the current block and the 4×4 area located on the right side of the upper left 4×4 area. When the product of the width and height of the current block is greater than or equal to the threshold and the width of the current block is smaller than the height, the non-separable transform matrix is applied to the upper left 4×4 area of the current block and the 4×4 area located below the upper left 4×4 area.

[0359] FIG. 30 shows an example of a block diagram of a device for processing a video signal in an embodiment to which the present invention is applied. The image processing device 3000 in FIG. 30 may correspond to the encoding device 100 in FIG. 1 or the decoding device 200 in FIG. 2.

[0360] The image processing device 3000 for processing an image signal includes a memory 3020 for storing the image signal and a processor 3010 for processing the image signal while being coupled to the memory.

[0361] The processor 3010 according to an embodiment of the present invention is composed of at least one processing circuit for processing an image signal, and can process the image signal by executing instruction words for encoding or decoding the image signal. That is, the processor 3010 encodes the original image data or decodes the encoded image signal by executing the above-described encoding or decoding method.

[0362] FIG. 31 shows an example of an image coding system in an embodiment to which the present invention is applied.

[0363] The image coding system includes a source device and a receiving device. The source device transmits encoded video / image information or data to the receiving device via a digital storage medium or a network in the form of a file or a stream.

[0364] The source device includes a video source, an encoding device, and a transmitter. The receiving device includes a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.

[0365] The video source acquires video / images through processes such as video / image capture, synthesis, or generation. The video source includes a video / image capture device and / or a video / image generation device. The video / image capture device includes, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device includes, for example, computers, tablets, and smartphones, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer, etc., and in this case, the process of generating related data can replace the process of video / image capture.

[0366] The encoding device encodes the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) is output in bitstream format.

[0367] The transmitting unit transmits the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit includes elements for generating a media file in a predetermined file format and elements for transmission via a broadcast / communication network. The receiver extracts the bitstream and transmits it to the decoding device.

[0368] The decoding device decodes the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.

[0369] The renderer renders the decoded video / image. The rendered video / image is displayed via the display unit.

[0370] FIG. 32 is a structural diagram of a content streaming system according to an embodiment to which the present invention is applied.

[0371] The content streaming system to which the present invention is applied includes an encoding server, a streaming server, a web server, a media storage device (repository), a user device, and a multimedia input device.

[0372] The encoding server serves to compress the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server may be omitted.

[0373] The bitstream is generated by an encoding method or a bitstream generation method to which the present invention is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0374] The streaming server transmits multimedia data to the user device based on a user request via the web server, and the web server serves as a medium for informing the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. Here, the content streaming system may include a separate control server, and in this case, the control server plays a role of controlling commands / responses between each device in the content streaming system.

[0375] The streaming server receives content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0376] Examples of user devices can include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (Head Mounted Displays)), digital TVs, desktop computers, digital signage, and the like.

[0377] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be processed distributively.

[0378] Also, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable storage medium. The multimedia data having the data structure according to the present invention can also be stored in a recording medium readable by a computer. The above computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which computer-readable data is stored. The above computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy (registered trademark) disk, and optical data storage device. Also, the above computer-readable recording medium includes a medium realized in the form of a carrier wave (for example, transmission through the Internet). Also, the bitstream generated by the encoding method can be stored in a computer-readable recording medium or transferred via a wired or wireless communication network.

[0379] Also, the embodiments of the present invention can be realized as a computer program product by program code, and the above program code can be executed on a computer according to the embodiments of the present invention. The above program code can be stored on a carrier readable by a computer.

[0380] As described above, the embodiments described in the present invention can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each figure can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.

[0381] In addition, the decoders and encoders to which the present invention is applied can be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video intercom devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, videophones, video devices, and medical video devices, etc., and are used for processing video signals or data signals. For example, an over-the-top (OTT) video device can include a game machine, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0382] In addition, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable storage medium. The multimedia data having the data structure according to the present invention can also be stored in a computer-readable storage medium. The above-mentioned computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which computer-readable data is stored. The above-mentioned computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy (registered trademark) disk, and optical data storage device. Further, the above-mentioned computer-readable recording medium includes a medium realized in the form of a carrier wave (for example, transmission through the Internet). Also, the bitstream generated by the encoding method can be stored in a computer-readable recording medium or transferred via a wired or wireless communication network.

[0383] In addition, the embodiments of the present invention can be realized as a computer program product by program code, and the above program code can be executed on a computer according to the embodiments of the present invention. The above program code can be stored on a carrier readable by a computer.

[0384] The embodiments described above are those in which the components and features of the present invention are combined in a predetermined form. Each component or feature should be considered as optional unless otherwise explicitly mentioned. Each component or feature can be implemented in a form that is not combined with other components or features. Also, it is possible to combine some components and / or features to constitute an embodiment of the present invention. The order of operations described in the embodiments of the present invention can be changed. Some components and features of any embodiment can be included in other embodiments, or can be replaced with corresponding components or features of other embodiments. It is obvious that embodiments can be constituted by combining claims without an explicit citation relationship in the claims, or can be included as new claims by amendment after filing.

[0385] The embodiments according to the present invention can be realized by various means, for example, hardware, firmware, software, or a combination thereof. In the case of realization by hardware, an embodiment of the present invention can be realized by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and the like.

[0386] In the case of realization by firmware or software, an embodiment of the present invention can be realized in the form of modules, procedures, functions, etc. that execute the functions or operations described above. The software code can be stored in a memory and driven by a processor. The memory can be located inside or outside the processor and can transmit and receive data with the processor by various means already known.

[0387] It is obvious to those skilled in the art that the present invention can be embodied in other specific forms without departing from the essential features of the present invention. Therefore, the above detailed description should not be construed restrictively in all aspects and should be regarded as exemplary. The scope of the present invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present invention are included in the scope of the present invention.

Industrial Applicability

[0388] As described above, the preferred embodiments of the present invention described above are disclosed for illustrative purposes, and those skilled in the art can make various improvements, modifications, substitutions, or additions to other embodiments within the technical idea and technical scope of the present invention disclosed in the appended patent claims below.

Claims

1. performing an inverse transform on transform coefficients of a current block to generate residual samples of the current block; generating reconstructed samples of the current block based on the residual samples of the current block; The step of performing the inverse transformation comprises: determining a non-separable transformation matrix based on the non-separable transformation indexes of the current block; applying the non-separable transform matrix to the transform coefficients of the current block; The method wherein the input length and output length of the inverse transform are determined as 8 and 16 based on the size of the current block being 4x4.

2. The method of claim 1 , wherein applying the non-separable transform matrix comprises applying the non-separable transform matrix to a number of the transform coefficients of the current block that corresponds to the input length.

3. The method of claim 1 , wherein the input length and the output length of the inverse transform are determined as 16 and 16 based on the size of the current block being 4×8 or 8×4.

4. 4. The method of claim 3, wherein applying the non-separable transformation matrix comprises applying the non-separable transformation matrix to an upper left 4x4 region of the current block based on each of a height and a width of the current block not being equal to 4 and a product of the width and the height being less than a threshold.

5. generating residual samples for the current block; generating transform coefficients for the current block from the residual samples of the current block based on a transform for the current block; encoding the transform coefficients of the current block to generate a bitstream; The step of generating transform coefficients comprises: determining a non-separable transformation matrix for the current block; generating the transform coefficients based on the non-separable transform matrix; a non-separable transform index associated with the non-separable transform matrix is ​​encoded in the bitstream; A method wherein the input length and output length of the transform are determined as 16 and 8 based on the size of the current block being 4x4.

6. The method of claim 5 , wherein the input length and the output length of the transform are determined as 16 and 16 based on the size of the current block being 4×8 or 8×4.

7. obtaining a bitstream relating to an image, said bitstream comprising: generating residual samples for the current block; generating transform coefficients for the current block from the residual samples of the current block based on a transform for the current block; encoding the transform coefficients of the current block to generate the bitstream; transmitting data including the bitstream relating to the image; The step of generating transform coefficients comprises: determining a non-separable transformation matrix for the current block; generating the transform coefficients based on the non-separable transform matrix; a non-separable transform index associated with the non-separable transform matrix is ​​encoded in the bitstream; A method wherein the input length and output length of the transform are determined as 16 and 8 based on the size of the current block being 4x4.