Method and apparatus for processing video signals
By applying a non-separable transform matrix based on block size for video signals, the method enhances coding efficiency and reduces complexity in processing high-resolution video content, addressing the challenges of next-generation video coding.
Patent Information
- Application Number
- JP2024098614
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-05
- Filing Date
- 2024-06-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-09-05
AI Technical Summary
Existing video coding technologies face challenges in efficiently processing next-generation video content with high spatial resolution, high frame rate, and high dimensionality, requiring improved coding tools for spatial domain to frequency domain conversion and prediction techniques with higher accuracy to reduce memory storage and processing demands.
The method involves applying a non-separable transform matrix to video signals based on the size of the current block, determining the input and output lengths, and using a non-separable transform set index for intra prediction modes, which includes a non-separable transform kernel and matrix, to enhance coding efficiency and reduce complexity.
This approach provides a video coding method with high coding efficiency and low complexity, effectively handling next-generation video content by optimizing memory storage and processing requirements.
Smart Images

Figure 0007708935000024 
Figure 0007708935000025 
Figure 0007708935000026
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for processing video signals, and more particularly to a method and apparatus for encoding or decoding video signals by performing a conversion.
Background Art
[0002] Compression encoding (encoding) refers to a series of signal processing techniques for transferring digitized information via a communication line or storing it in a form suitable for a storage medium. Media such as video, video, and audio can be the target of compression encoding, and in particular, the technique of performing compression encoding on video is called video compression.
[0003] Next-generation video content will have characteristics such as high spatial resolution, high frame rate, and high dimensionality of scene representation. Processing such content will result in a huge increase in terms of memory storage, memory access rate, and processing power.
[0004] Therefore, it is necessary to design coding tools for more efficiently processing next-generation video content. In particular, video codec standards after the HEVC (High Efficiency Video Coding) standard require efficient conversion techniques for converting video signals in the spatial domain to the frequency domain, along with prediction techniques having higher accuracy.
Summary of the Invention
Problems to be Solved by the Invention
[0005] Embodiments of the present invention aim to provide an image signal processing method and apparatus that apply a transformation having high coding efficiency and low complexity.
[0006] The technical problems to be solved by the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned should be clearly understood by those with ordinary knowledge in the technical field to which the present invention pertains from the following description.
Means for Solving the Problems
[0007] A method for decoding an image signal according to an embodiment of the present invention includes steps of determining an input length and an output length of a non-separable transform based on a height and a width of a current block, determining a non-separable transform matrix corresponding to the input length and the output length of the non-separable transform, and applying the non-separable transform matrix to coefficients for a number corresponding to the input length in the current block, where the height and the width of the block are greater than or equal to 8, and when the height and the width of the current block are each 8, the input length of the non-separable transform is determined to be 8.
[0008] Also, when the height and the width of the current block do not correspond to the case where they are each 8, the input length of the non-separable transform is determined to be 16.
[0009] Also, the output length is determined to be 48 or 64.
[0010] Also, the step of applying the non-separable transform matrix to the current block includes a step of applying the non-separable transform matrix to the upper left 4×4 region of the current block when the height and the width do not each correspond to 8 and the product of the width and the height is smaller than a critical value.
[0011] The step of determining the non-separable transform matrix includes: determining a non-separable transform set index based on the intra prediction mode of the current block; determining a non-separable transform kernel corresponding to the non-separable transform index within the non-separable transform set included in the non-separable transform set index; and determining the non-separable transform matrix from the non-separable transform kernel based on the input length and the output length.
[0012] An image signal processing apparatus according to another embodiment of the present invention includes a memory for storing the image signal and a processor coupled to the memory. The processor determines an input length and an output length of a non-separable transform based on a height and a width of a current block, determines a non-separable transform matrix corresponding to the input length and the output length of the non-separable transform, and is set to apply the non-separable transform matrix to coefficients corresponding to the number of the input length in the current block. The height and the width of the current block are greater than or equal to 8. When the height and the width of the current block are both 8, the input length of the non-separable transform is determined to be 8, and the output length is determined to be a value greater than the input length and less than or equal to 64.
Advantages of the Invention
[0013] According to an embodiment of the present invention, by applying a transform based on the size of a current block, a video coding method and apparatus having high coding efficiency and low complexity can be provided.
[0014] The effects obtained by the present invention are not limited to the effects mentioned above, and other effects not mentioned should be clearly understood by those of ordinary skill in the technical field to which the present invention belongs from the following description.
Brief Description of the Drawings
[0015] The accompanying drawings, which are included in and constitute a part of this detailed description, illustrate embodiments of the present invention and, together with the detailed description, explain the technical features of the present invention.
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21A
Figure 21B
Figure 22
Figure 23
Figure 24
Figure 25A
Figure 25B
Figure 26A
Figure 26B
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Embodiments for Carrying Out the Invention
[0017] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. The detailed description disclosed below together with the accompanying drawings is intended to explain exemplary embodiments of the present invention and is not intended to show the only embodiments in which the present invention can be implemented. The following detailed description includes specific details to provide a complete understanding of the present invention. However, those skilled in the art will understand that the present invention can be implemented without such specific details.
[0018] In some cases, to avoid obscuring the concept of the present invention, known structures and devices may be omitted or shown in the form of block diagrams centered on the core functions of each structure and device.
[0019] In some cases, to avoid obscuring the concept of the present invention, known structures and devices may be omitted or shown in the form of block diagrams centered on the core functions of each structure and device.
[0020] The specific terms used in the following description are provided to assist in understanding the present invention, and the use of such specific terms can be changed to other forms without departing from the technical idea of the present invention. For example, in the case of signals, data, samples, pictures, frames, blocks, etc., they may be appropriately substituted and interpreted in each coding process.
[0021] Hereinafter, in this specification, "processing unit" means a unit in which encoding / decoding processing such as prediction, conversion, and / or quantization is performed. Also, the processing unit can be interpreted to include a unit of a luma component and a unit of a chroma component. For example, the processing unit can correspond to a block, a coding unit (CU), a prediction unit (PU), or a transform unit (TU).
[0022] Also, the processing unit can be interpreted as a unit of a luma component or a unit of a chroma component. For example, the processing unit can correspond to a CTB, a CB, a PU, or a TB of the luma component. Or, the processing unit can correspond to a CTB, a CB, a PU, or a TB of the chroma component. Also, without being limited thereto, the processing unit may be interpreted to include a unit of a luma component and a unit of a chroma component.
[0023] Also, the processing unit is not necessarily limited to a square block and may be configured in the form of a polygon having three or more vertices.
[0024] Hereinafter, in this specification, pixels, picture elements, or coefficients (transformation coefficients or transformation coefficients after primary transformation) are collectively referred to as samples. And using samples means using pixel values, picture element values, or coefficients (transformation coefficients or transformation coefficients after primary transformation), etc.
[0025] Hereinafter, regarding an encoding / decoding method for a still image or a moving image, a design and an application method of a reduced secondary transform (RST) considering the computational complexity in the worst case will be described.
[0026] Embodiments of the present invention provide an image and video compression method and apparatus. The compressed data has the form of a bitstream, and the bitstream can be stored in various forms of storage, or can be streamed via a network and transmitted to a terminal having a decoder. In the terminal, when a display device is attached, the image decoded by the display device may be displayed, or the bitstream data may simply be stored. The methods and apparatuses proposed in the embodiments of the present invention can be applied to both the encoder and the decoder, and can be applied to all apparatuses that generate or receive a bitstream, regardless of whether the output is via a display device in the terminal or not.
[0027] The image compression apparatus is composed of a prediction unit, a transform and quantization unit, and an entropy coding unit. The schematic block diagrams of the encoding apparatus and the decoding apparatus are as shown in FIGS. 1 and 2. Among them, in the transform and quantization unit, the prediction signal is subtracted from the original signal to obtain a residual signal, which is then transformed into a frequency domain signal by a transform such as DCT (discrete cosine transform)-2, and then quantization is applied to greatly reduce the number of non-zero signals, enabling image compression.
[0028] FIG. 1 shows a schematic block diagram of an encoding apparatus in an embodiment to which the present invention is applied, where video / image signal encoding is performed.
[0029] The image segmentation unit 110 divides an input image (or picture, frame) input to the encoding device 100 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit is recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a Quad-tree binary-tree (QTBT) structure. For example, one coding unit is divided into a plurality of coding units with a deeper depth based on a quad-tree structure and / or a binary-tree structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure may be applied later. Or, the binary-tree structure may be applied first. The coding procedure according to the present invention is performed based on the final coding unit that is no longer divided. In this case, the largest coding unit may be immediately used as the final coding unit based on coding efficiency according to image characteristics, or, if necessary, the coding unit may be recursively divided into coding units with a deeper depth so that a coding unit of an optimal size is used as the final coding unit. Here, the coding procedure includes procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit are respectively divided or partitioned from the aforementioned final coding unit. The prediction unit is a unit of sample prediction, and the transformation unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0030] The unit may, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block represents a set of samples or transform coefficients consisting of M columns and N rows. Samples generally represent pixels or pixel values, and may represent only the pixel / pixel values of the luma component, or only the pixel / pixel values of the chroma component. Samples can be used as terms corresponding to pixels (or pels) of one picture (or image).
[0031] The encoding device 100 subtracts the predicted signal (predicted block, predicted sample array) output from the inter prediction unit 180 or the intra prediction unit 185 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 120. In this case, as shown in the figure, the unit that subtracts the predicted signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 100 may be called the subtraction unit 115. The prediction unit performs prediction on the block to be processed (hereinafter referred to as the current block), and generates a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit generates various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmits it to the entropy encoding unit 190. The information related to prediction is encoded in the entropy encoding unit 190 and output in the form of a bitstream.
[0032] The intra prediction unit 185 predicts the current block by referring to samples within the current picture. The samples to be referred to are located either in the neighborhood of the current block or at a distance depending on the prediction mode. In intra prediction, the prediction mode includes a plurality of non-directional modes and a plurality of directional modes. The non-directional modes include, for example, the DC mode and the Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is merely an example, and a greater or smaller number of directional prediction modes may be used depending on the setting. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0033] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information includes a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral blocks include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be referred to by names such as collocated reference blocks and collocated CUs (colCUs), and the reference picture including the temporal neighboring blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on the peripheral blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the peripheral blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the peripheral block is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0034] The prediction signal generated by the inter prediction unit 180 or the intra prediction unit 185 is used to generate a restored signal or to generate a residual signal.
[0035] The transform unit 120 applies a transform technique to the residual signal to generate transform coefficients. For example, the transform technique includes at least one of DCT, DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means a transform obtained from a graph when representing relationship information between pixels as a graph. CNT means a transform obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the transform process may be applied to a pixel block having the same size of a square or may be applied to a block of a variable size that is not square.
[0036] The quantization unit 130 quantizes the transform coefficients and transmits them to the entropy encoding unit 190. The entropy encoding unit 190 encodes the quantized signal (information regarding the quantized transform coefficients) and outputs it as a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. The quantization unit 130 can also reorder the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding). The entropy encoding unit 190 can encode, together or separately, in addition to the quantized transform coefficients, information necessary for video / image restoration (for example, values of syntax elements). The encoded information (for example, video / image information) is transmitted or stored in units of NAL (network abstraction layer) units in the form of a bitstream. The bitstream is transmitted via a network or stored in a digital storage medium. Here, the network includes a broadcast network and / or a communication network, etc., and the digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 may be configured as internal / external elements of the encoding device 100, or the transmission unit may be a component of the entropy encoding unit 190.
[0037] The quantized transform coefficients output from the quantization unit 130 can be used to generate a prediction signal. For example, the quantized transform coefficients can restore the residual signal by applying inverse quantization and inverse transformation by the inverse quantization unit 140 and the inverse transformation unit 150 within the loop. The addition unit 155 generates a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 may be referred to as a restoration unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next processing target block within the current picture, or may be used for inter prediction of the next picture after passing through filtering as described later.
[0038] The filtering unit 160 can apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 applies various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and transmits the modified reconstructed picture to the decoded picture buffer 170. Various filtering methods include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like. The filtering unit 160 generates various information related to filtering as described later in the description of each filtering method and transmits it to the entropy encoding unit 190. The information related to filtering is encoded in the entropy encoding unit 190 and output in the form of a bitstream.
[0039] The corrected decoded picture transmitted to the decoded picture buffer 170 is used as a reference picture in the inter prediction unit 180. When inter prediction is applied in this way, the encoding device 100 can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve the encoding efficiency.
[0040] The decoded picture buffer 170 can store the corrected restored picture for use as a reference picture in the inter prediction unit 180.
[0041] FIG. 2 shows a schematic block diagram of a decoding device in an embodiment to which the present invention is applied, in which an image signal is decoded.
[0042] As shown in FIG. 2, the decoding device 200 includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a decoded picture buffer (DPB) 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a prediction unit. That is, the prediction unit includes the inter prediction unit 180 and the intra prediction unit 185. The inverse quantization unit 220 and the inverse transform unit 230 may be collectively referred to as a residual processing unit. That is, the residual processing unit includes the inverse quantization unit 220 and the inverse transform unit 230. The above-described entropy decoding unit 210, inverse quantization unit 220, inverse transform unit 230, addition unit 235, filtering unit 240, inter prediction unit 260, and intra prediction unit 265 are configured by one hardware component (for example, a decoder or a processor) according to an embodiment. Also, the decoded picture buffer 250 is configured by one hardware component (for example, a memory or a digital storage medium) according to an embodiment.
[0043] When a bitstream including video / image information is input, the decoding device 200 can restore an image corresponding to the process in which the video / image information was processed in the encoding device 100 of FIG. 2. For example, the decoding device 200 performs decoding using the processing unit applied in the encoding device 100. Therefore, the processing unit for decoding is, for example, an encoding unit, and the encoding unit is divided from an encoding tree unit or a maximum encoding unit by a quad tree structure and / or a binary tree structure. Then, the restored image signal decoded and output by the decoding device 200 is reproduced by a reproducing device.
[0044] The decoding device 200 receives the signal output from the encoding device 100 of FIG. 2 in the form of a bitstream, and the received signal is decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 parses the bitstream to derive information (e.g., video / image information) necessary for image restoration (or, picture restoration). For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient related to the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding and the block to be decoded, or the information of the symbol / bin decoded in the previous stage, predicts the occurrence probability of the bin according to the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. Here, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin. Among the information decoded in the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual value obtained by performing entropy decoding in the entropy decoding unit 210, that is, the quantized transform coefficient and related parameter information, are input to the inverse quantization unit 220. Also, among the information decoded in the entropy decoding unit 210, the information related to filtering is provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device 100 may be further configured as an internal / external element of the decoding device 200, or the receiving unit may be a component of the entropy decoding unit 210.
[0045] In the inverse quantization unit 220, the quantized transform coefficients are inverse quantized to output transform coefficients. The inverse quantization unit 220 rearranges the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device 100. The inverse quantization unit 220 performs inverse quantization on the quantized transform coefficients using quantization parameters (for example, quantization step size information) to obtain transform coefficients.
[0046] The inverse transform unit 230 obtains a residual signal (residual block, residual sample array) by inverse transforming the transform coefficients.
[0047] The prediction unit performs prediction on the current block and generates a predicted block including predicted samples for the current block. The prediction unit determines whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode.
[0048] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block or at a distance depending on the prediction mode. In intra prediction, the prediction mode includes all of a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 265 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0049] The inter prediction unit 260 derives a predicted block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information is predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 constructs a motion information candidate list based on information regarding the prediction of neighboring blocks, and derives the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction is performed based on various prediction modes, and the information regarding the prediction includes information indicating the mode of inter prediction for the current block.
[0050] The adder 235 generates a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 260 or the intra prediction unit 265. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.
[0051] The adder 235 may also be referred to as a restoration unit or a restored block generation unit. The generated restored signal may be used for intra prediction of the next processing target block within the current picture, or may be used for inter prediction of the next picture after passing through filtering as described later.
[0052] The filtering unit 240 can improve the subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 applies various filtering methods to the restored picture to generate a modified restored picture and transmits the modified restored picture to the decoded picture buffer 250. The various filtering methods include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter (ALF), and bilateral filter.
[0053] The modified decoded picture transmitted to the decoded picture buffer 250 can be used as a reference picture by the inter prediction unit 260.
[0054] In this document, the embodiments described in the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the encoding device 100 are applied to the filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of the decoding device 200 identically or correspondingly.
[0055] FIG. 3 is an embodiment to which the present invention can be applied. FIG. 3A is a diagram for explaining a block division structure by QT (quadtree: QT), FIG. 3B is a diagram for explaining a block division structure by BT (binary tree: BT), FIG. 3C is a diagram for explaining a block division structure by TT (ternary tree: TT), and FIG. 3D is a diagram for explaining a block division structure by AT (asymmetric tree: AT).
[0056] In video coding, one block can be divided based on QT. Also, one subblock divided by QT may be recursively further divided using QT. A leaf block that cannot be further divided by QT is divided by at least one of the BT, TT, or AT methods. BT can have two forms of division: horizontal BT (2N×N, 2N×N) and vertical BT (N×2N, N×2N). TT can have two forms of division: horizontal TT (2N×1 / 2N, 2N×N, 2N×1 / 2N) and vertical TT (1 / 2N×2N, N×2N, 1 / 2N×2N). AT can have four forms of division: horizontal-up AT (2N×1 / 2N, 2N×3 / 2N), horizontal-down AT (2N×3 / 2N, 2N×1 / 2N), vertical-left AT (1 / 2N×2N, 3 / 2N×2N), and vertical-right AT (3 / 2N×2N, 1 / 2N×2N). Each of the BT, TT, and AT may be recursively further divided using BT, TT, and AT.
[0057] Figure 3A shows an example of QT division. Block A is divided into four subblocks (A0, A1, A2, A3) by QT. Subblock A1 is again divided into four subblocks (B0, B1, B2, B3) by QT.
[0058] Figure 3B shows an example of BT division. Block B3, which cannot be further divided by QT, is divided into vertical BT (C0, C1) or horizontal BT (D0, D1). Each subblock, like block C0, is recursively further divided in the form of horizontal BT (E0, E1) or vertical BT (F0, F1).
[0059] Figure 3C shows an example of TT division. The block B3 that cannot be further divided by QT is divided into vertical TT (C0, C1, C2) or horizontal TT (D0, D1, D2). Each sub-block, like block C1, is recursively further divided, such as in the form of horizontal TT (E0, E1, E2) or vertical TT (F0, F1, F2).
[0060] Figure 3D shows an example of AT division. The block B3 that cannot be further divided by QT is divided into vertical AT (C0, C1) or horizontal AT (D0, D1). Each sub-block, like block C1, can be recursively further divided, such as in the form of horizontal AT (E0, E1) or vertical TT (F0, F1).
[0061] On the other hand, BT, TT, and AT divisions may all be used together. For example, the sub-blocks divided by BT can be divided by TT or AT. Also, the sub-blocks divided by TT can be divided by BT or AT. The sub-blocks divided by AT can be divided by BT or TT. For example, after horizontal BT division, each sub-block can be divided vertically by BT, or after vertical BT division, each sub-block can also be divided horizontally by BT. In this case, although the division order is different, the finally divided shapes are the same.
[0062] Also, when a block is divided, the order of exploring the blocks can be defined in various ways. Generally, the exploration is carried out from left to right and from the top end to the bottom end. Exploring a block means the order of determining the possibility of additional block division for each divided sub-block, or when the block cannot be further divided, it means the encoding order of each sub-block, or the exploration order when referring to the information of other adjacent blocks in the sub-block.
[0063] Conversions can be performed separately for each processing unit (or transformation block) divided by a divided structure such as FIGS. 3A to 3D. In particular, the conversion matrix can be applied by dividing in the row direction and the column direction. According to an embodiment of the present invention, different conversion types can be used according to the length of the processing unit (or transformation block) in the row direction or the column direction.
[0064] FIGS. 4 and 5 are embodiments to which the present invention is applied. FIG. 4 shows a schematic block diagram of the conversion and quantization units 120 / 130 and the inverse quantization and inverse conversion units 140 / 150 in the encoding device 100 of FIG. 1, and FIG. 5 shows a schematic block diagram of the inverse quantization and inverse conversion units 220 / 230 in the decoding device 200.
[0065] As shown in FIG. 4, the conversion and quantization units 120 / 130 include a primary transform unit 121, a secondary transform unit 122, and a quantization unit 130. The inverse quantization and inverse conversion units 140 / 150 include an inverse quantization unit 140, an inverse secondary transform unit 151, and an inverse primary transform unit 152.
[0066] As shown in FIG. 5, the inverse quantization and inverse conversion units 220 / 230 include an inverse quantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.
[0067] In the present invention, when performing conversion, the conversion is performed through a plurality of stages. For example, as shown in FIG. 4, two stages of primary transform and secondary transform can be applied, and more conversion stages can also be used depending on the algorithm. Here, the primary transform may be referred to as a core transform.
[0068] The primary conversion unit 121 applies primary conversion to the residual signal, where the primary conversion can already be defined as a table in the encoder and / or decoder.
[0069] The secondary conversion unit 122 applies secondary conversion to the primarily converted signal, where the secondary conversion can already be defined as a table in the encoder and / or decoder.
[0070] In one embodiment, a non-separable secondary transform (NSST) can be conditionally applied as the secondary conversion. For example, NSST is applied only when it is a prediction block within the screen, and can have a set of transforms applicable for each prediction mode group.
[0071] Here, the prediction mode group is set based on the symmetry with respect to the prediction direction. For example, prediction mode 52 and prediction mode 16 are symmetric with respect to prediction mode 34 (diagonal direction), so they form one group and the same transform set can be applied. Here, when applying the transform for prediction mode 52, it is applied after transposing the input data, because the transform set for prediction mode 16 is the same.
[0072] On the other hand, in the case of the Planar mode and the DC mode, there is no symmetry with respect to direction, so they have their own transform sets, and the transform set can be composed of two transforms. For the remaining directional modes, the transform set can be composed of three transforms for each.
[0073] The quantization unit 130 quantizes the secondarily converted signal.
[0074] The inverse quantization and inverse conversion units 140 / 150 perform the above-described process in reverse, and redundant descriptions are omitted.
[0075] FIG. 5 shows a schematic block diagram of the inverse quantization and inverse transform units 220 / 230 in the decoding apparatus 200.
[0076] As shown in FIG. 5, the inverse quantization and inverse transform unit 220 / 230 includes an inverse quantization unit 220, an inverse secondary transform unit 231, and an inverse primary transform unit 232.
[0077] The inverse quantization unit 220 obtains transform coefficients from the entropy-decoded signal using quantization step size information.
[0078] In the inverse secondary transform unit 231, an inverse secondary transform is performed on the transform coefficients. Here, the inverse secondary transform represents the inverse transform of the secondary transform described with reference to FIG. 4.
[0079] The inverse primary transform unit 232 performs an inverse primary transform on the inversely secondarily transformed signal (or block) to obtain a residual signal. Here, the inverse primary transform represents the inverse transform of the primary transform described with reference to FIG. 4.
[0080] FIG. 6 is a flowchart showing an embodiment to which the present invention is applied, in which a video signal is encoded by primary transform and secondary transform. Each operation shown in FIG. 6 is performed by the transform unit 120 of the encoding apparatus 100.
[0081] The encoding apparatus 100 determines (or selects) a forward secondary transform based on at least one of the prediction mode, block shape, and / or block size of the current block (S610).
[0082] The encoding device 100 can determine an optimal forward secondary transformation by rate-distortion optimization. The optimal forward secondary transformation corresponds to one of a plurality of transformation combinations, and the plurality of transformation combinations are defined by transformation indexes. For example, for rate-distortion optimization, the encoding device 100 can compare the results of performing all of forward secondary transformation, quantization, residual coding, etc. for each candidate.
[0083] The encoding device 100 signals (S620) a secondary transformation index corresponding to the optimal forward secondary transformation. Here, the secondary transformation index can be applied to other embodiments described in this specification.
[0084] On the other hand, the encoding device 100 performs a forward primary transformation on the current block (residual block) (S630).
[0085] The encoding device 100 performs a forward secondary transformation on the current block using the optimal forward secondary transformation (S640). On the other hand, the forward secondary transformation can be the RST described below. RST means a transformation in which N pieces of residual data (N×1 residual vector) are input and R pieces of transform coefficient data (R×1 transform coefficient vector) are output (R < N).
[0086] As one embodiment, RST can be applied to a specific region of the current block. For example, when the current block is N×N, the specific region can mean the upper left N / 2×N / 2 region. However, the present invention is not limited to this, and is set to be different depending on at least one of the prediction mode, block shape, or block size. For example, when the current block is N×N, the specific region can mean the upper left M×M region (M≦N).
[0087] On the other hand, the encoding device 100 generates a transform coefficient block by performing quantization on the current block (S650).
[0088] The encoding device 100 can perform entropy encoding on the transform coefficient block to generate a bitstream.
[0089] FIG. 7 shows a flowchart for decoding a video signal by inverse secondary transformation and inverse primary transformation, which is an embodiment to which the present invention is applied. Each operation shown in FIG. 7 is performed by the inverse transformation unit 230 of the decoding device 200.
[0090] The decoding device 200 acquires a secondary transformation index from the bitstream (S710).
[0091] The decoding device 200 induces a secondary transformation corresponding to the secondary transformation index (S720).
[0092] However, the steps S710 and S720 are one embodiment, and the present invention is not limited thereto. For example, the decoding device 200 can induce a secondary transformation based on at least one of the prediction mode, block shape, and / or block size of the current block without acquiring the secondary transformation index.
[0093] On the other hand, the decoder 200 entropy-decodes the bitstream to obtain a transform coefficient block, and performs inverse quantization on the transform coefficient block (S730).
[0094] The decoder 200 performs an inverse-direction secondary transformation on the inverse-quantized transform coefficient block (S740). For example, the inverse-direction secondary transformation can be an inverse-direction RST. The inverse-direction RST is the transpose matrix of the RST described in FIG. 6, and means a transformation in which R transform coefficient data (Rx1 transform coefficient vector) are input and N residual data (Nx1 residual vector) are output.
[0095] As one embodiment, the reduced secondary transform can be applied to a specific region of the current block. For example, when the current block is N×N, the specific region may mean the upper left N / 2×N / 2 region. However, the present invention is not limited thereto and is set to vary depending on at least one of the prediction mode, block shape, or block size. For example, when the current block is N×N, the specific region may mean the upper left M×M region (M≦N) or M×L (M≦N, L≦N).
[0096] Then, the decoder 200 performs an inverse first transform on the result of the inverse-direction secondary transform (S750).
[0097] The decoder 200 generates a residual block in step S750, and generates a restored block by adding the residual block and the prediction block.
[0098] FIG. 8 shows an example of a transform configuration group to which AMT (adaptive multiple transform) according to an embodiment of the present invention is applied.
[0099] According to FIG. 8, the transform configuration group is determined based on the prediction mode, and the number of groups can be six in total (G0 to G5). And G0 to G4 correspond to the case where intra prediction is applied, and G5 indicates a transform combination (or transform set, transform combination set) applied to a residual block generated by inter prediction.
[0100] One transform combination is composed of a horizontal transform (or row transform) applied to the rows of the corresponding two-dimensional block and a vertical transform (or column transform) applied to the columns.
[0101] Here, each of all the conversion setting groups includes four conversion combination candidates. The four conversion combination candidates are selected or determined by conversion combination indices from 0 to 3, and the conversion combination index is transmitted from the encoding device 100 to the decoding device 200 by an encoding procedure.
[0102] As one embodiment, the residual data (or residual signal) obtained by intra prediction has different statistical characteristics according to the intra prediction mode. Therefore, other conversions other than the general cosine conversion can be applied for each intra prediction mode as shown in FIG. 8. In this document, the conversion type may be expressed as, for example, DCT-Type 2, DCT-II, DCT-2.
[0103] As shown in FIG. 8, the conversion set configurations for the case where 35 intra prediction modes are used and the case where 67 intra prediction modes are used are respectively illustrated. A plurality of conversion combinations can be applied for each conversion setting group divided in the intra prediction mode column. For example, a plurality of conversion combinations (row direction conversion, column direction conversion) are composed of four combinations. More specifically, in group 0, DST-7 and DCT-5 can be applied to all in the row (horizontal) direction and the column (vertical) direction, so four combinations are possible.
[0104] Since a combination of a total of four conversion kernels can be applied to each intra prediction mode, a conversion combination index for selecting one of them is transmitted for each transform unit. In this document, the conversion combination index is referred to as an AMT index and may be expressed as amt_idx.
[0105] In addition to the conversion kernel shown in FIG. 8, due to the characteristics of the residual signal, DCT-2 may be optimal for both the row direction and the column direction. Therefore, adaptive conversion can be performed by defining an AMT flag for each coding unit. Here, when the AMT flag is 0, DCT-2 is applied to both the row direction and the column direction. When the AMT flag is 1, one of the four combinations can be selected or determined by the AMT index.
[0106] As one embodiment, when the AMT flag is 0 and the number of conversion coefficients for one conversion unit is less than 3, instead of applying the conversion kernel of FIG. 8, DST-7 is applied to both the row direction and the column direction.
[0107] As one embodiment, by first parsing the value of the conversion coefficient and applying DST-7 without parsing the AMT index when the number of conversion coefficients is less than 3, the amount of additional information transmitted can be reduced.
[0108] As one embodiment, AMT can be applied only when the width and height of the conversion unit are both 32 or less.
[0109] As one embodiment, FIG. 8 may be preset by off-line training.
[0110] As one embodiment, the AMT index can be defined by one index that can simultaneously indicate a combination of horizontal conversion and vertical conversion. Alternatively, the AMT index can be defined separately by a horizontal conversion index and a vertical conversion index.
[0111] The technique of applying a selected transform from among a plurality of transform kernels (e.g., DCT-2, DST-7, DCT-8) like the aforementioned AMT may be referred to as MTS (multiple transform selection) or EMT (enhanced multiple transform), and the AMT index may be referred to as the MT index.
[0112] FIG. 9 shows an example of an encoding flowchart to which AMT according to an embodiment of the present invention is applied. The operations shown in FIG. 9 are performed by the transform unit 120 of the encoding device 100.
[0113] This document basically describes embodiments that apply transforms separately for the horizontal and vertical directions, but the transform combination can also be composed of non-separable transforms.
[0114] Also, it can be composed of a mixture of separable and non-separable transforms. In this case, when a non-separable transform is used, transform selection by row / column or selection by horizontal / vertical direction becomes unnecessary, and the transform combination of FIG. 8 is used only when a separable transform is selected.
[0115] Moreover, the method proposed in this document can be applied regardless of whether it is a primary transform or a secondary transform. That is, there is no restriction that it must be applied to only one of the two, and it can be applied to both. Here, the primary transform means a transform for first transforming the residual block, and the secondary transform can mean a transform for applying a transform to the block generated as a result of the primary transform.
[0116] First, the encoding device 100 determines a transform setting group corresponding to the current block (S910). Here, the transform setting group can also be configured in a combination as shown in FIG. 8.
[0117] The encoding device 100 performs conversion on the combination of candidate conversions available within the conversion setting group (S920).
[0118] As a result of the conversion execution, the encoding device 100 determines or selects the conversion combination with the smallest RD (rate distortion) cost (S930).
[0119] The encoding device 100 encodes the conversion combination index corresponding to the selected conversion combination (S940).
[0120] FIG. 10 shows an example of a decode flowchart to which AMT according to an embodiment of the present invention is applied. The operations shown in FIG. 10 are performed by the inverse conversion unit 230 of the decoding device 200.
[0121] First, the decoding device 200 determines the conversion setting group for the current block (S1010). The decoding device 200 parses (or acquires) the conversion combination index from the video signal, where the conversion combination index corresponds to any one of a plurality of conversion combinations within the conversion setting group (S1020). For example, the conversion setting group includes DCT-2, DST-7, or DCT-8.
[0122] The decoding device 200 derives the conversion combination corresponding to the conversion combination index. Here, the conversion combination is composed of a horizontal conversion and a vertical conversion, and includes at least one of DCT-2, DST-7, or DCT-8. Also, the conversion combination may use the conversion combinations described in FIG. 8.
[0123] The decoding device 200 performs inverse transformation on the current block based on the induced transformation combination (S1040). When the transformation combination consists of a row (horizontal) transformation and a column (vertical) transformation, the row (horizontal) transformation can be applied first, and then the column (vertical) transformation can be applied. However, the present invention is not limited to this. Conversely, if it is applied in the reverse order or consists of a non-separable transformation, the non-separable transformation can be immediately applied.
[0124] In one embodiment, when the vertical transformation or the horizontal transformation is DST-7 or DCT-8, the inverse transformation of DST-7 or the inverse transformation of DCT-8 is applied for each column first, and then for each row. Also, different transformations are applied for each row and / or each column for the vertical transformation or the horizontal transformation.
[0125] In one embodiment, the transformation combination index can be obtained based on an AMT flag indicating whether AMT is performed. That is, the transformation combination index can be obtained only when AMT is performed by the AMT flag. Also, the decoding device 200 checks whether the number of non-zero coefficients is greater than a threshold value. Here, the transformation combination index can be parsed only when the number of non-zero coefficients is greater than the threshold value.
[0126] In one embodiment, the AMT flag or the AMT index is defined at at least one level of a sequence, a picture, a slice, a block, a coding unit, a transform unit, or a prediction unit.
[0127] On the other hand, as another embodiment, the process of determining the transformation setting group and the process of parsing the transformation combination index can be performed simultaneously. Or, the step S1010 can be omitted as it has already been set in the encoding device 100 and / or the decoding device 200.
[0128] FIG. 11 shows an example of a flowchart for encoding an AMT flag and an AMT index according to an embodiment of the present invention. The operations in FIG. 11 are performed by the conversion unit 120 of the encoding device 100.
[0129] The encoding device 100 determines whether AMT is applied to the current block (S1110).
[0130] If AMT is applied, the encoding device 100 encodes with AMT flag = 1 (S1120).
[0131] Then, the encoding device 100 determines the AMT index based on at least one of the prediction mode, horizontal conversion, and vertical conversion of the current block (S1130). Here, the AMT index indicates an index indicating any one of a plurality of conversion combinations for each intra prediction mode, and the AMT index is transmitted for each conversion unit.
[0132] When the AMT index is determined, the encoding device 100 encodes the AMT index (S1140).
[0133] On the other hand, if AMT is not applied, the encoding device 100 encodes with AMT flag = 0 (S1150).
[0134] FIG. 12 shows an example of a decoding flowchart for performing conversion based on the AMT flag and the AMT index.
[0135] The decoding device 200 parses the AMT flag from the bitstream (S1210). Here, the AMT flag indicates whether AMT is applied to the current block.
[0136] The decoding device 200 checks whether AMT is applied to the current block based on the AMT flag (S1220). For example, it checks whether the AMT flag is 1.
[0137] If the AMT flag is 1, the decoding device 200 parses the AMT index (S1230). Here, the AMT index means an index indicating any one of a plurality of conversion combinations for each intra prediction mode, and the AMT index can be transmitted for each conversion unit. Alternatively, the AMT index means an index indicating any one of the conversion combinations defined in a preset conversion combination table. Here, the preset conversion combination table may mean FIG. 8, but the present invention is not limited thereto.
[0138] The decoding device 200 induces or determines horizontal conversion and vertical conversion based on at least one of the AMT index or the prediction mode (S1240).
[0139] Alternatively, the decoding device 200 induces a conversion combination corresponding to the AMT index. For example, the decoding device 200 induces or determines horizontal conversion and vertical conversion corresponding to the AMT index.
[0140] On the other hand, if the AMT flag is 0, the decoding device 200 applies the preset inverse vertical conversion for each column (S1250). For example, the inverse vertical conversion may be an inverse DCT-2 conversion.
[0141] Then, the decoding device 200 applies the preset inverse horizontal conversion for each row (S1260). For example, the inverse horizontal conversion may be an inverse DCT-2 conversion. That is, when the AMT flag is 0, the conversion kernel already set in the encoding device 100 or the decoding device 200 is used. For example, instead of being defined in a conversion combination table such as FIG. 8, a frequently used conversion kernel may be used.
[0142] NSST (non-separable secondary transform)
[0143] The secondary transformation means applying the transformation kernel once again with the result obtained by applying the primary transformation as the input. The primary transformation includes DCT-2, DST-7 in HEVC, and the aforementioned AMT, etc. The non-separable transform does not sequentially apply the N×N transformation kernel to the row direction and the column direction, but rather regards the N×N two-dimensional residual block as an N 2 ×1 vector, and then applies the N 2 ×N 2 transformation kernel only once.
[0144] That is, NSST refers to a non-separable square matrix applied to a vector composed of the coefficients of the transformation block. Also, in the embodiments of this document, NSST is mainly described as an example of the non-separable transform applied to the upper left region (low-frequency region) determined by the block size. However, the embodiments of the present invention are not limited to the term NSST, and any type of non-separable transform can be applied to the embodiments of the present invention. For example, the non-separable transform applied to the upper left region (low-frequency region) determined by the block size is called LFNST (low frequency non-separable transform). In this document, the M×N transform (or transform matrix) means a matrix composed of M rows and N columns.
[0145] In NSST, after dividing the two-dimensional block data obtained by applying the primary transformation into M×M blocks, for each M×M block, M 2 ×M 2Apply the non-separable transform. The value of M can be 4 or 8. Instead of applying the NSST to all regions of the 2D block obtained by the primary transform, it is also possible to apply it only to some regions. For example, the NSST can be applied only to the top-left 8×8 block. Also, the 64×64 non-separable transform can be applied only to the top-left 8×8 region when both the width and height of the 2D block obtained by the primary transform are 8 or more, and for the remaining cases, the 16×16 non-separable transform can be applied separately to 4×4 blocks.
[0146] M 2 ×M 2 The non-separable transform can also be applied in the form of a matrix product, but for reducing the computational amount and memory requirement, it can be approximated by a combination of a Givens rotation layer and a permutation layer. FIG. 13 shows one Givens rotation. It can be explained by one angle of one Givens rotation as shown in FIG. 13.
[0147] FIGS. 13 and 14 are embodiments to which the present invention is applied. FIG. 13 shows a diagram for explaining a Givens rotation, and FIG. 14 shows the configuration of one round in a 4×4 NSST composed of a Givens rotation layer and a permutation.
[0148] Both the 8×8 NSST and the 4×4 NSST can be composed of a hierarchical combination of Givens rotations. The matrix corresponding to one Givens rotation is as shown in Equation 1, and when expressing the matrix product as a diagram, it is as shown in FIG. 13.
[0149]
Equation
[0150] In FIG. 13, t output by the Givens rotation m and t nIt can be calculated as shown in Equation 2.
[0151]
Number
[0152] As shown in FIG. 13, since one Givens rotation rotates two data, 32 or 8 Givens rotations are respectively required for processing 64 data (in the case of 8×8 NSST) or 16 data (in the case of 4×4 NSST). Therefore, a bundle of 32 or 8 Givens rotations can form a Givens rotation layer. As shown in FIG. 15, the output data for one Givens rotation layer is transmitted to the input data for the next Givens rotation layer by replacement (shuffling). The pattern to be replaced as shown in FIG. 15 is regularly defined. In the case of 4×4 NSST, four Givens rotation layers and corresponding replacements form one round. 4×4 NSST is performed by two rounds, and 8×8 NSST is performed by four rounds. Different rounds use the same replacement pattern, but the applied Givens rotation angles are different. Therefore, it is necessary to store the angle data for all Givens rotations constituting each transformation.
[0153] As a final step, one final replacement is further performed on the data output through the Givens rotation layer, and the information regarding the replacement is stored separately for each transformation. This replacement is performed at the end of the forward NSST, and the inverse NSST first applies the inverse replacement.
[0154] The inverse NSST performs the Givens rotation layers and replacements applied in the forward NSST in reverse order, and rotates by taking a minus (-) value for the angle of each Givens rotation.
[0155] FIG. 15 shows an example of the configuration of a set of non-separable transforms for each intra prediction mode according to an embodiment of the present invention.
[0156] Intra prediction modes to which the same NSST or set of NSSTs is applied can form a group. Figure 15 classifies 67 intra prediction modes into 35 groups. For example, both mode 20 and mode 48 belong to group 20 (hereinafter, mode group).
[0157] A plurality of NSSTs other than one NSST can be configured as a set for each mode group. Each set includes the case where no NSST is applied. For example, when three different NSSTs can be applied to one mode group, it can be configured to select one of four cases including the case where no NSST is applied. Here, an index is transmitted in TU units to distinguish one of the four cases. The number of NSSTs can be configured to be different for each mode group. For example, mode groups 0 and 1 are signaled to select one of three cases each including the case where no NSST is applied.
[0158] Embodiment 1: RST applicable to a 4×4 block
[0159] A non-separable transform applicable to a single 4×4 block is a 16×16 transform. That is, when the data elements constituting the 4×4 block are arranged in a column in row-first or column-first order, they become a 16×1 vector, and the non-separable transform can be applied to the 16×1 vector. The forward 16×16 transform is composed of 16 row-direction transform basis vectors. When taking the inner product of the 16×1 vector and each transform basis vector, the transform coefficient for the transform basis vector is obtained. The process of obtaining all the transform coefficients for the 16 transform basis vectors seems to be multiplying a 16×16 non-separable transform matrix by the input 16×1 vector. The transform coefficients obtained by the matrix product have the form of a 16×1 vector, but the statistical characteristics may differ for different transform coefficients. For example, if the 16×1 transform coefficient vector is composed of elements from the 0th to the 15th, the variance of the 0th element may be larger than the variance of the 15th element. That is, the earlier the element is located, the larger the variance value and the larger the energy value it can have.
[0160] When applying the inverse 16×16 non-separable transform to the 16×1 transform coefficients, (when ignoring the effects of quantization and integerization calculations, etc.), the original 4×4 block signal can be restored. If the forward 16×16 non-separable transform is an orthonormal transform, the inverse 16×16 transform can be obtained by taking the transpose of the matrix for the forward 16×16 transform. Simply put, multiplying the inverse 16×16 non-separable transform matrix by the 16×1 transform coefficient vector gives data in the form of a 16×1 vector, and arranging it in the row-first or column-first order applied initially can restore the 4×4 block signal.
[0161] As described above, the elements forming the 16×1 conversion coefficient vector may have different statistical characteristics. As in the above example, if the conversion coefficients arranged in the front (close to the 0th element) have larger energy, even if the inverse conversion is applied to some of the conversion coefficients that appear first without using all the conversion coefficients, a signal quite close to the original signal can be restored. For example, assuming that the inverse 16×16 non-separable conversion is composed of 16 column basis vectors, an L×16 matrix is formed by leaving only L column basis vectors, and after leaving only the more important L conversion coefficients from among the conversion coefficients (L×1 vector, which can appear first as in the above example), multiplying the L×16 matrix by the L×1 vector can restore a 16×1 vector with a small error from the original input 16×1 vector data. As a result, since only L coefficients participate in data restoration, when obtaining the conversion coefficients, it is only necessary to obtain an L×1 conversion coefficient vector instead of a 16×1 conversion coefficient vector. That is, in the forward 16×16 non-separable conversion matrix, after selecting L of the row-direction conversion vectors to form an L×16 conversion and multiplying by a 16×1 input vector, L important conversion coefficients can be obtained.
[0162] Embodiment 2: Application area setting of 4×4 RST and arrangement of transform coefficients
[0163] 4×4RST can be applied as a secondary conversion, and at this time, it can be applied secondarily to the block to which a primary conversion such as DCT-type2 has been applied. When the size of the block to which the primary conversion has been applied is N×N, it is usually larger than 4×4. Therefore, when applying 4×4RST to an N×N block, the following two methods can be considered.
[0164] 1) Instead of applying 4×4RST to the entire N×N area, it can be applied only to a part of the area. For example, it can be applied only to the top-left M×M area (M <= N).
[0165] 2) After dividing the area to which the secondary conversion is applied into 4×4 blocks, 4×4RST can be applied to each divided block.
[0166] The above methods 1) and 2) can be applied in combination. For example, after dividing only the upper left M×M region into 4×4 blocks, 4×4 RST can be applied.
[0167] As a specific embodiment, secondary conversion is applied only to the upper left 8×8 region. When the N×N block is larger than or equal to 8×8, 8×8 RST is applied. When the N×N block is smaller than 8×8, after dividing it into 4×4 blocks as in (4×4, 8×4, 4×8) of the above 2), 4×4 RST can be applied to each of them.
[0168] After applying 4×4 RST, if L transformation coefficients (1 <= L < 16) are generated, there is a degree of freedom in how to arrange the L transformation coefficients. However, when reading and processing the transformation coefficients in the residual coding part, since there is a defined order, the coding performance may vary depending on how the L transformation coefficients are arranged in a 2-dimensional block. Residual coding in the HEVC (high efficiency video coding) standard starts coding from the position farthest from the DC position. This is to improve the coding performance by taking advantage of the fact that the value of the coefficient after quantization is 0 or close to 0 as the distance from the DC position increases. Therefore, arranging the L transformation coefficients so that they have high energy and the more important coefficients are coded later in the order of residual coding may be advantageous in terms of coding performance.
[0169] FIG. 16 shows three forward scan orders for the transformation coefficients or transformation coefficient blocks applied in the HEVC standard, where (a) is a diagonal scan, (b) is a horizontal scan, and (c) is a vertical scan.
[0170] FIG. 16 illustrates three forward scan orders for the transform coefficients or transform coefficient blocks (4×4 blocks, Coefficient Group (CG)) applied in the HEVC standard, and the residual coding is performed in the reverse order of the scan order of (a), (b), or (c) (i.e., coded in the order from 16 to 1). Since the three scan orders shown in (a), (b), and (c) are selected according to the intra-prediction mode, the scan order can be determined for the L transform coefficients in the same manner according to the intra-prediction mode.
[0171] The L value has a range of 1 <= L < 16. Generally, L out of 16 transform basis vectors can be selected in any way. However, from the viewpoints of encoding and decoding, it may be advantageous from the viewpoint of encoding efficiency to select the transform basis vectors that are more important in terms of signal energy as in the examples presented above.
[0172] FIGS. 17 and 18 are embodiments to which the present invention is applied. FIG. 17 shows the positions of the transform coefficients when the forward diagonal scan is applied when applying 4×4 RST to a 4×8 block, and FIG. 18 shows an example when the valid transform coefficients of two 4×4 blocks are merged into one block.
[0173] When the upper left 4×8 block is divided into 4×4 blocks and 4×4 RST is applied according to the diagonal scan order of (a), when the L value is 8 (i.e., only 8 out of 16 transform coefficients are left), the transform coefficients are positioned as shown in FIG. 17, but only half of each 4×4 block can have transform coefficients, and the positions indicated by X are filled with a default value of 0. Therefore, assuming that L transform coefficients are arranged for each 4×4 block according to the scan order presented in (a) and the remaining (16 - L) positions of each 4×4 block are filled with 0, the residual coding (e.g., residual coding in HEVC) can be applied.
[0174] Also, as shown in FIG. 18, the L transform coefficients arranged in two 4×4 blocks can be configured into one block. In particular, when the L value is 8, the transform coefficients of the two 4×4 blocks completely fill one 4×4 block, so there are no remaining transform coefficients in other blocks. Therefore, since residual coding is not required for the 4×4 block that has become empty of transform coefficients, in the case of HEVC, a flag (coded_sub_block_flag) indicating whether or not to apply residual coding to the block is coded as 0. The combination methods for the positions of the transform coefficients of the two 4×4 blocks are diverse. For example, they can be combined in any order, but the following methods may also be applied.
[0175] 1) Alternately combine the transform coefficients of the two 4×4 blocks in scan order. That is, in FIGS. 8A, 8B, and 8C, when the transform coefficient for the upper block is JPEG0007708935000003.jpg556 and the transform coefficient of the lower block is JPEG0007708935000004.jpg552, They can be combined alternately one by one as in JPEG0007708935000005.jpg568. Also, JPEG0007708935000006.jpg55 and JPEG0007708935000007.jpg55 can be reordered JPEG0007708935000008.jpg1190.
[0176] 2) The transform coefficients for the first 4×4 block can be arranged first, and then the transform coefficients for the second 4×4 block can be arranged. That is, They can be concatenated and arranged as in JPEG0007708935000009.jpg667. Of course, The order can also be changed as in JPEG0007708935000010.jpg669.
[0177] Embodiment 3: Method for coding NSST (non-separable secondary transform) index for 4×4 RST
[0178] When 4×4 RST is applied as shown in FIG. 17, values of 0 are filled from the (L + 1)-th position to the 16-th position in accordance with the transform coefficient scan order for each 4×4 block. Therefore, if there is a non-zero value among the positions from the (L + 1)-th position to the 16-th position in even one of the two 4×4 blocks, it is derived that the case is one where 4×4 RST is not applied. When 4×4 RST has a structure in which a selected transform from among a set of transforms prepared like JEM (joint experiment model) NSST is applied, an index (hereinafter referred to as the NSST index) for what transform to apply is signaled.
[0179] In a certain decoder, the NSST index can be known by bitstream parsing, and the bitstream parsing can be performed after residual coding. In this case, if there is a non-zero transform coefficient between the (L + 1)-th position and the 16-th position by residual decoding, since it is certain that 4×4 RST is not applied in the decoder, the decoder does not parse the NSST index. Therefore, the signaling cost is reduced by selectively parsing the NSST index only when necessary.
[0180] As shown in Fig. 17, when 4×4 RST is applied to a plurality of 4×4 blocks within a specific region (at this time, the same 4×4 RST may be applied, or different 4×4 RSTs may be applied respectively), the (same or different) 4×4 RSTs applied to all 4×4 blocks are specified by one NSST index. Since the 4×4 RST and the applicability of the 4×4 RST for all 4×4 blocks are determined by one NSST index, as a result of investigating during the residual decoding process whether there are non-zero conversion coefficients at positions from the (L + 1)-th to the 16-th positions for all 4×4 blocks, if there are non-zero conversion coefficients at positions not allowed in the 4×4 block (from the (L + 1)-th position to the 16-th position), the encoding device 100 can be set not to code the NSST index.
[0181] The encoding device 100 can also separately signal each NSST index for the luminance block and the chrominance block. In the case of the chrominance block, it can also separately signal a different NSST index for the Cb component and the Cr component, or can also use one common NSST index. When one NSST index is used, the signaling of the NSST index is also performed only once. When one NSST index is shared for the Cb component and the Cr component, the 4×4 RST indicated by the same NSST index is applied. In this case, the 4×4 RSTs for the Cb component and the Cr component themselves may be the same, or although the NSST index is the same, individual 4×4 RSTs may be set for the Cb component and the Cr component. When the NSST index shared for the Cb component and the Cr component is used, it is checked whether there are non-zero conversion coefficients at positions from the (L + 1)-th position to the 16-th position for all 4×4 blocks for the above-mentioned conditional signaling. If non-zero conversion coefficients are found at positions from the (L + 1)-th position to the 16-th position, the signaling for the NSST index may be omitted.
[0182] As shown in FIG. 18, even when the transform coefficients for two 4×4 blocks are merged into one 4×4 block, when 4×4 RST is applied, after checking whether a non-zero transform coefficient appears at a position where there is no valid transform coefficient, the encoding device 100 can determine whether signaling for the NSST index is possible. In particular, since the L value is 8 as shown in FIG. 18, when there is no valid transform coefficient in one 4×4 block when 4×4 RST is applied (the block indicated by X in FIG. 18(b)), the flag (coded_sub_block_flag) for whether the residual coding of the block is possible is checked, and if it is 1, the NSST index can be set not to be signaled. As described above, in the following description, NSST will be mainly described as an example of non-separable transform, but other known terms (for example, LFNST) may be used for non-separable transform. For example, the NSST set and the NSST index may be used in place of the LFNST set and the LFNST index. Also, the RST described in this document can be used in place of LFNST as an example of non-separable transform that uses a non-square transform matrix having a reduced input length and / or a reduced output length in a square non-separable transform matrix applied to at least a partial region of the transform block (the remaining region excluding the lower right 4×4 region in the upper left 4×4, 8×8 region or 8×8 block).
[0183] Embodiment 4: Optimization method for the case where coding for 4×4 index is performed before residual coding
[0184] When coding for the NSST index is performed before residual coding, since the applicability of 4×4RST is determined in advance, residual coding for positions filled with conversion coefficients of 0 may be omitted. Here, the applicability of 4×4RST can be configured to be determined by the NSST index value (for example, if the NSST index is 0, 4×4RST is not applied), or the applicability of 4×4RST can be signaled by a separate syntax element (for example, the NSST flag). For example, assuming that the separate syntax element is the NSST flag, the decoding device 200 determines the applicability of 4×4RST by parsing the NSST flag first, and then, if the NSST flag value is 1, for positions where no valid conversion coefficients can exist as described above, residual coding (decoding) can be omitted.
[0185] In the case of HEVC, during the execution of residual coding, coding is performed at the position of the last non-zero coefficient in the TU for the first time. If coding for the NSST index is performed after coding for the position of the last non-zero coefficient, and when assuming the applicability of 4×4RST at the position of the last non-zero coefficient, if it is a position where no non-zero coefficients can exist, the decoding device 200 can be set not to apply 4×4RST without coding the NSST index. For example, in the case of the position indicated as X in FIG. 17, when 4×4RST is applied, since no valid conversion coefficients are located (the value of 0 can be filled), if the last non-zero coefficient is located in the area indicated as X, the decoding device 200 can omit coding for the NSST index. If the last non-zero coefficient is not located in the area indicated as X, the decoding device 200 can perform coding for the NSST index.
[0186] When the NSST index is conditionally coded after coding for the positions of non-zero coefficients, if it is determined whether 4×4RST can be applied, hereinafter, the remaining part where residual coding is performed can be processed in the following two ways.
[0187] 1) When 4×4RST is not applied, general residual coding is performed. That is, coding is performed under the assumption that non-zero transform coefficients may exist at any position from the position of the last non-zero coefficient to DC.
[0188] 2) When 4×4RST is applied, since the transform coefficient does not exist for a specific position or a specific 4×4 block (for example, the X position in FIG. 17) (filled with 0 by default), residual coding may be omitted for the position or block. For example, when reaching the position indicated as X while scanning in accordance with the scan order in FIG. 17, coding for the flag (sig_coeff_flag) regarding whether a non-zero coefficient exists at the position in the HEVC standard can be omitted. When the transform coefficients of two blocks are merged into one block as shown in FIG. 18, coding for the flag (for example, code_sub_block_flag in the HEVC standard) indicating the availability of residual coding for the 4×4 block filled with 0 can be omitted, and the corresponding value can be induced to 0, and the corresponding 4×4 block can be filled with all 0 values without separate coding.
[0189] When coding the NSST index after coding for the position of the last non-zero coefficient, if the x-position (Px) and y-position (Py) of the last non-zero coefficient are each smaller than Tx and Ty, the coding of the NSST index can be omitted and the 4×4RST can be set not to be applied. For example, when Tx = 1, Ty = 1 and the last non-zero coefficient exists at the DC position, the NSST index coding is omitted. The method for determining whether the NSST index coding is possible for comparison with such a threshold value can be applied differently to the luminance component and the chrominance component. For example, different Tx and Ty may be applied to the luminance component and the chrominance component respectively, or a threshold value may be applied to the luminance component and no threshold value may be applied to the chrominance component. Conversely, a threshold value may be applied to the chrominance component and no threshold value may be applied to the luminance component.
[0190] The two methods described above (when the last non-zero coefficient is located in an area where there is no valid transform coefficient, the NSST index coding is omitted; when the X coordinate and Y coordinate for the last non-zero coefficient are each smaller than the threshold value, the NSST index coding is omitted) may be applied simultaneously. For example, after first checking the threshold value for the position coordinates of the last non-zero coefficient, it is possible to check whether the last non-zero coefficient is located in an area where there is no valid transform coefficient, and the order of the two methods can be changed.
[0191] The method presented in Embodiment 4) can also be applied to 8×8RST. That is, when the last non-zero coefficient is located in the upper left 4×4 non-region within the upper left 8×8 region, the coding for the NSST index can be omitted; otherwise, the coding for the NSST index can be performed. Also, when all the values of the X and Y coordinates for the position of the last non-zero coefficient are less than a certain threshold value, the coding for the NSST index can be omitted. The two methods can also be applied simultaneously.
[0192] Embodiment 5: Applying different NSST index coding and residual coding methods to luminance component and chrominance components respectively when applying RST
[0193] In Embodiment 3 and Embodiment 4, the methods described can be applied differently to the luminance component and the chrominance components. That is, the NSST index coding and the residual coding method can be applied differently to the luminance component and the chrominance components. For example, the method described in Embodiment 4 can be applied to the luminance component, and the method described in Embodiment 3 can be applied to the chrominance components. Also, conditional NSST index coding proposed in Embodiment 3 or Embodiment 4 can be applied to the luminance component, and it is also possible that conditional NSST index coding is not applied to the luminance component, and vice versa (conditional NSST index coding is applied to the chrominance components and not applied to the luminance component).
[0194] Embodiment 6
[0195] One embodiment of the present invention provides a mixed NSST transform set (MNTS) for applying various NSST conditions in the process of applying NSST and a method for constructing the corresponding MNTS.
[0196] According to JEM, depending on the size of the pre-selected sub-block, the 4×4 NSST set only includes 4×4 kernels, and the 8×8 NSST set only includes 8×8 kernels. Embodiments of the present invention additionally propose a method for constructing a mixed NSST set as follows.
[0197] - The size of the NSST kernel available in the NSST set is not fixed, and an NSST kernel having one or more variable sizes can be included in the NSST set (for example, both a 4×4 NSST kernel and an 8×8 NSST kernel are included in one NSST set).
[0198] - The number of NSST kernels available within the NSST set can be variable rather than fixed (e.g., the first set includes 3 kernels and the second set includes 4 kernels).
[0199] - The order of the NSST kernels may not be fixed and can be defined to be different for different NSST sets (e.g., in the first set, NSST kernels 1, 2, 3 are mapped to NSST indices 1, 2, 3 respectively, while in the second set, NSST kernels 3, 2, 1 are mapped to NSST indices 1, 2, 3 respectively).
[0200] More specifically, an example of a method for configuring a mixed NSST transform set is as follows.
[0201] - The priority of the NSST kernels available in the NSST transform set can be determined by the size of the NSST kernels (e.g., 4×4 NSST and 8×8 NSST).
[0202] For example, when the block is large, the 8×8 NSST kernel may be more important than the 4×4 NSST kernel, so an NSST index with a lower value is assigned to the 8×8 NSST kernel.
[0203] - The priority of the NSST kernels available in the NSST transform set can be determined by the order of the NSST kernels.
[0204] For example, a given 4×4 NSST first kernel may be preferred over a 4×4 NSST second kernel.
[0205] Since the NSST index is encoded and transmitted, by assigning a higher priority (smaller index) to the more frequently occurring NSST kernels, the NSST index can be signaled with fewer bits.
[0206] Tables 1 and 2 below show examples of the mixed NSST sets proposed in this embodiment.
[0207]
Table 1
[0208]
Table 2
[0209] Embodiment 7
[0210] In one embodiment of the present invention, in the process of determining the secondary conversion set, a method for determining the NSST set is proposed by considering the intra prediction mode and the block size.
[0211] The method proposed in this embodiment is related to Embodiment 6 to form a conversion set adapted to the intra prediction mode, so as to form kernels of various sizes and apply them to blocks.
[0212] FIG. 19 is an embodiment to which the present invention is applied, and shows an example of a method for constructing a mixed NSST set for each intra prediction mode.
[0213] FIG. 19 is an example of a table by a method of applying the method proposed in Embodiment 2 in relation to Embodiment 6. That is, as shown in FIG. 19, an index ("Mixed Type") indicating whether to follow the existing NSST set construction method or the NSST set construction method of another method for each intra prediction mode is defined.
[0214] More specifically, in the case of the intra prediction mode where the index ("Mixed Type") is defined as "1" in FIG. 19, regardless of the NSST set construction method of JEM, the NSST set is constructed using the NSST set construction method defined in the system. Here, the NSST set construction method defined in the system means the mixed NSST set proposed in Embodiment 6.
[0215] As another embodiment, the table in FIG. 19 describes two types of transform set construction methods (JEM-based NSST set construction, the mixed type NSST set construction method proposed in the embodiments of the present invention) based on mixed type information (flags) related to the intra prediction mode. However, there may be one or more mixed type NSST construction methods, and here, the mixed type information can be represented by N (N>2) various values.
[0216] As another embodiment, it is possible to determine whether to configure a transform set suitable for the current block in a mixed type by considering both the intra prediction mode and the size of the transform block. For example, if the mode type corresponding to the intra prediction mode is 0, the NSST set of JEM is set accordingly; otherwise, (Mode Type = 1) various mixed type NSST sets can be determined according to the size of the transform block.
[0217] FIG. 20 shows an example of a method for selecting an NSST set (or kernel) considering the intra prediction mode and the size of the transform block in an embodiment to which the present invention is applied.
[0218] When the transform set is determined, the decoding device 200 can determine the NSST kernel used by using the NSST index information.
[0219] Embodiment 8
[0220] In one embodiment of the present invention, when constructing a transform set by considering the intra prediction mode and the block size during the application of the secondary transform, a method for efficiently encoding the NSST index is provided by considering the change in the statistical distribution of the NSST index values transmitted after encoding. Embodiments of the present invention provide a method for selecting a kernel to be applied using a syntax indicating the kernel size.
[0221] In addition, in embodiments of the present invention, since the number of available NSST kernels differs for each transform set, a method of truncated unary binary evolution as shown in Table 3 below is provided according to the maximum NSST index value available for each set for an efficient binary evolution method.
[0222]
Table 3
[0223] Table 3 shows the binary evolution method of the NSST index value. Since the number of available NSST kernels differs for each transform set, the NSST index can be binary evolved by the maximum NSST index value.
[0224] Embodiment 9: Reduced Transform
[0225] Due to complexity issues in the transform (e.g., large block transform or non-separable transform), a reduced transform applicable to the core transform (e.g., DCT, DST, etc.) and the secondary transform (e.g., NSST) is provided.
[0226] The main idea of the reduced transform is to map an N-dimensional vector from another space to an R-dimensional vector, where R / N (R < N) is the reduction factor. The reduced transform is an R×N matrix as shown in Equation 3 below.
[0227]
Number
[0228] In Equation 1, the R rows of the transformation are the R bases of the new N-dimensional space. Therefore, the reason it is called a reduced transformation is that the number of elements of the vector output by the transformation is smaller than the number of elements of the input vector (R < N). The inverse transform matrix for the reduced transformation is the transpose of the forward transformation. The forward and inverse reduced transformations will be described with reference to FIGS. 13A and 13B.
[0229] FIGS. 21A and 21B show embodiments to which the present invention is applied and show forward and inverse reduced transformations.
[0230] The number of elements of the reduced transformation is RxN, which is smaller than the size of the complete matrix (N×N) by R / N, which means that the required memory is R / N of the complete matrix.
[0231] Also, the number of required multiplications is also R×N, which is smaller than the original N×N by R / N.
[0232] If X is an N-level vector, then after applying the reduced transformation, R coefficients are obtained, which means that only R values need to be transmitted instead of the original N coefficients.
[0233] FIG. 22 shows an example of a decode flowchart using the reduced transformation according to an embodiment of the present invention.
[0234] The proposed reduced transformation (inverse transformation in the decoder) can be applied to the coefficients (inverse quantized coefficients), as shown in FIG. 21. A predetermined reduction factor (R, or R / N) and a transformation kernel for performing the transformation may be required. Here, the transformation kernel can be determined based on information that can be used such as the block size (width, height), intra prediction mode, and Cidx. If the current coding block is a luma block, Cidx is 0. Otherwise (Cb or Cr block), Cidx is a non-zero value such as 1.
[0235] Hereinafter, the operators used in this document are defined as shown in Tables 4 and 5 below
[0236] [Table 4]
[0237] [Table 5]
[0238] FIG. 23 shows an example of a flowchart for the application of conditional reduced transformation according to an embodiment of the present invention. The operations in FIG. 23 are performed by the inverse quantization unit 140 and the inverse transformation unit 150 of the decoding apparatus 200.
[0239] In one embodiment, the reduced transformation can be used when certain conditions are satisfied. For example, the reduced transformation can be applied to blocks larger than a certain size as follows.
[0240] - Width > TH && Height > HT (where TH is a predefined value (e.g., 4))
[0241] Or
[0242] - Width * Height > K && MIN(width, height) > TH (K and TH are predefined values)
[0243] That is, when the width of the current block is greater than a predefined value (TH) and the height of the current block is greater than the predefined value (TH) as described in the above conditions, a reduced transformation can be applied. Or, when the product of the width and height of the current block is greater than a predefined value (K) and the smaller value of the width and height of the current block is greater than the predefined value (TH), a reduced transformation can be applied.
[0244] The reduced transformation can be applied to a group of predefined blocks as follows.
[0245] - Width == TH && Height == TH
[0246] Or
[0247] - Width == Height
[0248] That is, when the width and height of the current block are respectively the same as the predefined value (TH), or when the width and height of the current block are the same (when the current block is a square block), a reduced transformation can be applied.
[0249] If the conditions for using the reduced transformation are not satisfied, a regular transformation is applied. The regular transformation can be a transformation that is predefined and available in a video coding system. Examples of the regular transformation are as follows.
[0250] - DCT-2, DCT-4, DCT-5, DCT-7, DCT-8
[0251] Or
[0252] - DST-1, DST-4, DST-7
[0253] Or
[0254] - Non-separable transformation
[0255] or
[0256] -JEM-NSST(HyGT)
[0257] As shown in FIG. 23, the reduced conversion conditions depend on an index (Transform_idx) indicating which conversion (e.g., DCT-4, DST-1) is used or which kernel is applied (when multiple kernels are available). In particular, Transform_idx can be transmitted twice. One is the index indicating the horizontal conversion (Transform_idx_h), and the other is the index indicating the vertical conversion (Transform_idx_v).
[0258] More specifically, referring to FIG. 23, the decoding device 200 performs inverse quantization on the input bitstream (S2305). Thereafter, the decoding device 200 determines whether to apply a conversion (S2310). The decoding device 200 determines whether to apply a conversion based on a flag indicating whether to skip the conversion.
[0259] When a conversion is applied, the decoding device 200 parses a conversion index (Transform_idx) indicating the applied conversion (S2315). Also, the decoding device 200 selects a conversion kernel (S2330). For example, the decoding device 200 selects a conversion kernel corresponding to the conversion index (Transform_idx). Also, the decoding device 200 selects a conversion kernel in consideration of the block size (width, height), intra prediction mode, and CIdx (luma, chroma).
[0260] The decoding device 200 determines whether or not to satisfy the conditions for applying the reduced transformation (S2320). The conditions for applying the reduced transformation include the conditions as described above. When the reduced transformation is not applied, the decoding device 200 applies a normal inverse transformation (S2325). For example, the decoding device 200 determines an inverse transformation matrix from the transformation kernel selected in step S2330, and applies the determined inverse transformation matrix to the current block including the transformation coefficients.
[0261] When the reduced transformation is applied, the decoding device 200 applies a reduced inverse transformation (S2335). For example, the decoding device 200 determines a reduced inverse transformation matrix in consideration of a reduction factor from the transformation kernel selected in step S2330, and applies the reduced inverse transformation matrix to the current block including the transformation coefficients.
[0262] FIG. 24 shows an example of a decoding flowchart for a secondary inverse transformation to which a conditionally reduced transformation according to an embodiment of the present invention is applied. The operation of FIG. 24 is performed by the inverse transformation unit 230 of the decoding device 200.
[0263] In one embodiment, the reduced transformation can be applied to the secondary transformation as shown in FIG. 24. When the NSST index is parsed, the reduced inverse transformation can be applied.
[0264] Referring to FIG. 24, the decoding device 200 performs inverse quantization (S2405). For the transformation coefficients generated by the inverse quantization, the decoding device 200 determines whether or not to apply NSST (S2410). That is, the decoding device 200 determines whether or not it is necessary to parse the NSST index (NSST_idx) depending on whether or not to apply NSST.
[0265] When NSST is applied, the decoding device 200 parses the NSST index (S2415) and determines whether the NSST index is greater than 0 (S2420). The NSST index is restored by an entropy decoding unit 210 using a technique such as CABAC. When the NSST index is 0, the decoding device 200 omits the secondary inverse transformation and applies the core inverse transformation or the primary inverse transformation (S2445).
[0266] Also, when NSST is applied, the decoding device 200 selects a transformation kernel for the secondary inverse transformation (S2435). For example, the decoding device 200 selects a transformation kernel corresponding to the NSST index (NSST_idx). Also, the decoding device 200 selects a transformation kernel in consideration of the block size (width, height), the intra prediction mode, and CIdx (luma, chroma).
[0267] When the NSST index is greater than 0, the decoding device 200 determines whether the conditions for applying the reduced transformation are satisfied (S2425). The conditions for applying the reduced transformation include the conditions as described above. When the reduced transformation is not applied, the decoding device 200 applies the normal secondary inverse transformation (S2430). For example, the decoding device 200 determines a secondary inverse transformation matrix from the transformation kernel selected in step S2435 and applies the determined secondary inverse transformation matrix to the current block including the transformation coefficients.
[0268] When the reduced transformation is applied, the decoding device 200 applies the reduced secondary inverse transformation (S2440). For example, the decoding device 200 can determine a reduced inverse transformation matrix in consideration of the reduction factor from the transformation kernel selected in step S2335 and apply the reduced inverse transformation matrix to the current block including the transformation coefficients. Thereafter, the decoding device 200 applies the core inverse transformation or the primary inverse transformation (S2445).
[0269] Embodiment 10: Reduced Transform as a Secondary Transform with Different Block Size
[0270] Figures 25A, 25B, 26A, and 26B show examples of reduced transformation and reduced inverse transformation according to embodiments of the present invention.
[0271] In one embodiment of the present invention, reduced transformation in a video codec for different block sizes such as 4×4, 8×8, 16×16, etc. can be used as secondary transformation and secondary inverse transformation. As an example for an 8×8 block size and a reduction factor R = 16, the secondary transformation and the secondary inverse transformation can be set as shown in FIGS. 25A and 25B.
[0272] The pseudocode of the reduced transformation and the reduced inverse transformation is set as shown in FIG. 26.
[0273] Embodiment 11: Reduced Transform as a Secondary Transform with Non-Rectangular Shape
[0274] FIG. 27 shows an example of a region to which the reduced secondary transformation according to an embodiment of the present invention is applied.
[0275] As described above, due to complexity issues in the secondary transformation, the secondary transformation can be applied to the 4×4 and 8×8 corners. The reduced transformation can also be applied to non-rectangular shapes.
[0276] As described above, due to complexity issues in the secondary transformation, the secondary transformation can be applied to the 4×4 and 8×8 corners. The reduced transformation can also be applied to non-rectangular shapes.
[0277] In another example, when RST is applied to an 8×8 block, the non-separable transformation (RST) can be applied only to the three 4×4 blocks (a total of 48 transform coefficients) on the top-left, top-right, and bottom-left sides, excluding the bottom-right 4×4 block.
[0278] Embodiment 12: Reduction Factor
[0279] FIG. 28 shows a reduced transformation by a reduction factor according to an embodiment of the present invention.
[0280] Changing the reduction factor can change the memory and multiplication complexity. As described above, changing the reduction factor reduces the memory and multiplication complexity by a factor of (factor) R / N. For example, for an 8×8 NSST, when R = 16, the memory and multiplication complexity are reduced by a factor of 1 / 4.
[0281] Embodiment 13: High Level Syntax
[0282] Syntax elements such as those in Table 6 below can be used to process RST in video coding. The semantics related to the reduced transformation may be present in the SPS (sequence parameter set) or the slice header.
[0283] A Reduced_transform_enabled_flag equal to 1 indicates that the reduced transformation is possible and applied. A Reduced_transform_enabled_flag equal to 0 indicates that the reduced transformation is not possible. When the Reduced_transform_enabled_flag is not present, it is inferred to be equal to 0 (Reduced_transform_enabled_flag equals to 1 specifies that reduced transform is enabled and applied. Reduced_transform_enabled_flag equal to 0 specifies that reduced transform is not enabled. When Reduced_transform_enabled_flag is not present, it is inferred to be equal to 0).
[0284] Reduced_transform_factor indicates the number of reduced dimensions to maintain for the reduced transform. If Reduced_transform_factor does not exist, it is inferred to be the same as R (Reduced_transform_factor specifies that the number of reduced dimensions to keep for reduced transform. When Reduced_transform_factor is not present, it is inferred to be equal to R).
[0285] min_reduced_transform_size indicates the minimum transform size to apply the reduced transform. If min_reduced_transform_size does not exist, it is inferred to be 0 (min_reduced_transform_size specifies that the minimum transform size to apply reduced transform. When min_reduced_transform_size is not present, it is inferred to be equal to 0).
[0286] max_reduced_transform_size indicates the maximum transform size to apply the reduced transform. If max_reduced_transform_size does not exist, it is inferred to be 0.
[0287] reduced_transform_size indicates the number of reduced dimensions to maintain for the reduced transform. If reduced_transform_size does not exist, it is inferred to be 0. (reduced_transform_size specifies that the number of reduced dimensions to keep for reduced transform. When Reduced_transform_factor is not present, it is inferred to be equal to 0.)
[0288] [Table 6]
[0289] Embodiment 14: Conditional application of 4×4 RST for worst case handling
[0290] The non-separable second-order transform (4×4 NSST) applicable to a 4×4 block is a 16×16 transform. 4×4 NSST is applied secondarily to a block to which a first-order transform such as DCT-2, DST-7, or DCT-8 has been applied. Assuming the size of the block to which the first-order transform has been applied is N×M, when applying 4×4 NSST to an N×M block, the following methods can be considered.
[0291] 1) The conditions for applying 4×4 NSST to an N×M area are as follows in a) and b).
[0292] a) N >= 4
[0293] b) M >= 4
[0294] 2) Instead of applying 4×4 NSST to the entire N×M area, it is applied only to some areas. For example, 4×4 NSST can be applied only to the upper-left K×J area. The conditions for this case are as follows in a) and b).
[0295] a) K >= 4
[0296] b) J >= 4
[0297] 3) After dividing the area to which the second conversion is applied into 4×4 blocks, 4×4 NSST can be applied to each divided block.
[0298] Since the computational complexity of 4×4 NSST is a very important consideration factor for both the encoder and the decoder, it will be analyzed in detail. In particular, the computational complexity of 4×4 NSST is analyzed based on the number of multiplications. In the case of the forward NSST, the 16×16 second conversion is composed of 16 row-direction conversion basis vectors. When taking the inner product of a 16×1 vector with each conversion basis vector, the conversion coefficient for the corresponding conversion basis vector is obtained. The process of obtaining all the conversion coefficients for the 16 conversion basis vectors seems to be multiplying a 16×16 non-separable conversion matrix by an input 16×1 vector. Therefore, the total number of multiplications required for 4×4 forward NSST is 256.
[0299] In the decoder, when applying the inverse 16×16 non-separable conversion to the 16×1 conversion coefficients (when ignoring the effects such as quantization and integerization calculations), the coefficients of the original 4×4 first conversion block can be restored. In other words, when multiplying the inverse 16×16 non-separable conversion matrix by the 16×1 conversion coefficient vector, data in the form of a 16×1 vector is obtained. When arranging the data according to the row-first or column-first order applied initially, the 4×4 block signal (first conversion coefficient) can be restored. Therefore, the total number of multiplications required for 4×4 inverse NSST is 256.
[0300] As described above, when 4×4 NSST is applied, the number of multiplications required per sample unit is 16. This number is obtained by dividing the total number of multiplications 256, which is obtained in the inner product process between the 16×1 vector, which is the 4×4 NSST execution process, and each conversion basis vector, by the number of all samples 16. The number of multiplications required identically for both the forward 4×4 NSST and the inverse 4×4 NSST is 16.
[0301] In the case of an 8×8 block, the number of multiplications per sample required when applying 4×4 NSST is determined as follows depending on the area to which 4×4 NSST is applied.
[0302] 1. When 4×4 NSST is applied only to the upper left 4×4 area: 256 (the number of multiplications required in the 4×4 NSST process) / 64 (the total number of samples in the 8×8 block) = 4 multiplications / sample
[0303] 2. When 4×4 NSST is applied to the upper left 4×4 area and the upper right 4×4 area: 512 (the number of multiplications required in two 4×4 NSST processes) / 64 (the total number of samples in the 8×8 block) = 8 multiplications / sample
[0304] 3. When 4×4 NSST is applied to all 4×4 areas of the 8×8 block: 1024 (the number of multiplications required in four 4×4 NSST processes) / 64 (the total number of samples in the 8×8 block) = 16 multiplications / sample
[0305] As described above, when the block size is large, the range to which 4×4 NSST is applied can be reduced in order to decrease the number of multiplications in the worst case required for each sample.
[0306] Therefore, when using 4×4 NSST, the worst case occurs when the size of the TU is 4×4. The methods for reducing the worst case complexity are as follows.
[0307] Method 1. Do not apply 4×4 NSST to small TUs (i.e., 4×4 TUs).
[0308] Method 2. For 4×4 blocks (4×4 TUs), apply 4×4 RST instead of 4×4 NSST.
[0309] In the case of Method 1, it was experimentally observed that not applying the 4×4 NSST caused a significant degradation in the encoding performance. In the case of Method 2, it was revealed that by applying the inverse transformation to some of the conversion coefficients located at the front side without using all the conversion coefficients depending on the statistical characteristics of the elements constituting the 16×1 conversion coefficient vector, a signal very close to the original signal can be restored, and most of the encoding performance can be maintained.
[0310] Specifically, in the case of the 4×4 RST, assuming that the 16×16 non-separable transformation in the reverse direction (or forward direction) is composed of 16 column basis vectors, only L column basis vectors are left to form a 16×L matrix. By leaving only the more important L conversion coefficients among the conversion coefficients and multiplying the 16×L matrix by the L×1 vector, a 16×1 vector with a small error from the original 16×1 vector data can be restored.
[0311] As a result, since only L coefficients are involved in data restoration, instead of obtaining a 16×1 conversion coefficient vector to obtain the conversion coefficients, an L×1 conversion coefficient vector should be obtained. That is, by selecting L row-direction conversion vectors in the forward 16×16 non-separable transformation matrix, an L×16 transformation matrix is formed, and when the L×16 transformation matrix is multiplied by the 16×1 input vector, L conversion coefficients are obtained.
[0312] The value of L has a range of 1 <= L < 16. Generally, L can be selected from any of the 16 conversion basis vectors in any way. However, as described above, it may be advantageous from the perspective of encoding efficiency to select the conversion basis vectors with high energy importance of the signal from the aspects of encoding and decoding. The worst-case number of multiplications per sample in the 4×4 block due to the change in the L value is as shown in Table 7 below.
[0313]
Table 7
[0314] As described above, for reducing the multiplication complexity in the worst case, 4×4 NSST and 4×4 RST can be used in combination as shown in Table 8 below (however, the following example explains the conditions for applying 4×4 NSST and 4×4 RST under the condition for applying 4×4 NSST (i.e., when the width and height of the current block are both greater than or equal to 4)).
[0315] As described above, the 4×4 NSST for a 4×4 block is a square (16x16) transformation matrix that takes 16 data as input and outputs 16 data, and the 4×4 RST means a non-square (8×16) transformation matrix that takes 16 data as input and outputs R data smaller than 16 (for example, 8) based on the encoder side. Based on the decoder side, the 4×4 RST means a non-square (16×8) transformation matrix that takes R data smaller than 16 (for example, 8) as input and outputs 16 data.
[0316]
Table 8
[0317] Referring to Table 8, when the width and height of the current block are both 4, 4×4 RST based on an 8×16 matrix is applied to the current block, and otherwise (when either one of the width or height of the current block is not 4), 4×4 NSST can be applied to the 4×4 area in the upper left corner of the current block. More specifically, when the size of the current block is 4×4, a non-separable transformation with an input length of 16 and an output length of 8 can be applied. In the case of the inverse non-separable transformation, conversely, a non-separable transformation with an input length of 8 and an output length of 16 can be applied.
[0318] As described above, for reducing the multiplication complexity in the worst case, 4×4 NSST and 4×4 RST can be combined and used as shown in Table 9 below (however, the following example explains the conditions for applying 4×4 NSST and 4×4 RST under the condition for applying 4×4 NSST (i.e., when the width and height of the current block are both greater than or equal to 4)).
[0319]
Table 9
[0320] Referring to Table 9, when the width and height of the current block are both 4, a 4×4 RST based on an 8×16 matrix is applied. If the product of the width and height of the current block is less than the critical value (TH), a 4×4 NSST is applied to the upper left 4×4 area of the current block. If the width of the current block is greater than or equal to its height, a 4×4 NSST is applied to the upper left 4×4 area and the 4×4 area located on the right side of the upper left 4×4 area of the current block. In the remaining cases (when it is less than the height of the current block), a 4×4 NSST is applied to the upper left 4×4 area and the 4×4 area located below the upper left 4×4 area of the current block.
[0321] In conclusion, for reducing the computational complexity of multiplication in the worst case, a 4×4 RST (e.g., an 8×16 matrix) can be applied to 4×4 blocks instead of a 4×4 NSST.
[0322] Embodiment 15: Conditional application of 8×8 RST for worst case handling
[0323] The non-separable secondary transform (8×8 NSST) applicable to an 8×8 block is a 64×64 transform. The 8×8 NSST is secondarily applied to blocks to which a primary transform such as DCT-2, DST-7, or DCT-8 has been applied. Assuming the size of the block to which the primary transform has been applied is N×M, when applying the 8×8 NSST to an N×M block, the following methods are considered.
[0324] 1) The conditions for applying the 8×8 NSST to an N×M area are as follows in c) and d).
[0325] c) N >= 8
[0326] d) M >= 8
[0327] 2) Instead of applying the 8×8 NSST to the entire N×M region, it may be applied only to some regions. For example, the 8×8 NSST is applied only to the upper left K×J region. The conditions for this case are as follows in c) and d).
[0328] c) K >= 8
[0329] d) J >= 8
[0330] 3) After dividing the region to which the secondary transformation is applied into 8×8 blocks, the 8×8 NSST can be applied to each divided block.
[0331] The computational complexity of the 8×8 NSST is a very important consideration factor for both the encoder and the decoder, so it will be analyzed in detail. In particular, the computational complexity of the 8×8 NSST is analyzed based on the number of multiplications. In the case of the forward NSST, the 64×64 non-separable secondary transformation is composed of 64 row-direction transformation basis vectors. When taking the inner product of a 64×1 vector with each transformation basis vector, the transformation coefficient for the corresponding transformation basis vector is obtained. The process of obtaining all the transformation coefficients for the 64 transformation basis vectors seems to be multiplying a 64×64 non-separable transformation matrix by the input 64×1 vector. Therefore, the total number of multiplications required for the 8×8 forward NSST is 4096.
[0332] In the decoder, when applying the inverse 64×64 non-separable transformation to the 64×1 transformation coefficients (when ignoring the effects such as quantization and integerization calculations), the coefficients of the original 8×8 primary transformation block can be restored. In other words, when multiplying the inverse 64×64 non-separable transformation matrix by the 64×1 transformation coefficient vector, data in the form of a 64×1 vector is obtained. When arranging the data according to the row-first or column-first order applied initially, the 8×8 block signal (primary transformation coefficients) can be restored. Therefore, the total number of multiplications required for the 8×8 inverse NSST is 4096.
[0333] As described above, when the 8×8 NSST is applied, the number of multiplications required per sample unit is 64. This is the number obtained by dividing the total number of multiplications 4096, which is obtained in the inner product process between the 64×1 vector, which is the 8×8 NSST execution process, and each transformation basis vector, by the number of all samples 64. The number of multiplications required identically for both the forward 8×8 NSST and the reverse 8×8 NSST is 64.
[0334] If it is the case of a 16×16 block, when the 8×8 NSST is applied, the number of multiplications per sample required is determined as follows depending on the area where the 8×8 NSST is applied.
[0335] 1. When the 8×8 NSST is applied only to the upper left 8×8 area: 4096 (the number of multiplications required in the 8×8 NSST process) / 256 (the number of all samples in the 16×16 block) = 16 multiplications / sample
[0336] 2. When the 8×8 NSST is applied to the upper left 8×8 area and the upper right 8×8 area: 8192 (the number of multiplications required in two 8×8 NSST processes) / 256 (the number of all samples in the 16×16 block) = 32 multiplications / sample
[0337] 3. When the 8×8 NSST is applied to all 8×8 areas of the 16×16 block: 16384 (the number of multiplications required in four 8×8 NSST processes) / 256 (the number of all samples in the 16×16 block) = 64 multiplications / sample
[0338] As described above, when the block size is large, the range where the 8×8 NSST is applied can be reduced in order to reduce the number of multiplications in the worst case required per sample.
[0339] When 8×8 NSST is applied, since the 8×8 block is the smallest TU to which 8×8 NSST can be applied, the case where the TU size is 8×8 corresponds to the worst case from the perspective of the number of multiplications required per sample. In this case, the methods for reducing the worst case complexity are as follows.
[0340] Method 1. Do not apply 8×8 NSST to small TUs (i.e., 8×8 TUs).
[0341] Method 2. In the case of an 8×8 block (8×8 TU), apply 8×8 RST instead of 8×8 NSST.
[0342] In the case of Method 1, it has been experimentally observed that not applying 8×8 NSST causes a significant degradation in coding performance. In the case of Method 2, it has been clarified that, due to the statistical characteristics of the elements constituting the 64×1 conversion coefficient vector, a signal that is quite close to the original signal can be restored by applying inverse conversion to some of the conversion coefficients located at the front side without using all the conversion coefficients, and most of the coding performance can be maintained.
[0343] Specifically, in the case of 8×8 RST, assuming that the inverse (or forward) 64×64 non-separable transform is composed of 16 column basis vectors, only L column basis vectors are left to form a 64×L matrix. By leaving only the more important L conversion coefficients among the conversion coefficients and multiplying the 64×L matrix by an L×1 vector, a 64×1 vector with a small error from the original 64×1 vector data can be restored.
[0344] Also, as described in Embodiment 11, RST can be applied not to all 64 conversion coefficients included in the 8×8 block, but to some regions (for example, the remaining region excluding the lower right 4×4 region in the 8×8 block).
[0345] As a result, since only L coefficients intervene in data restoration, an L×1 conversion coefficient vector may be obtained instead of a 64×1 conversion coefficient vector to acquire the conversion coefficients. That is, an L×64 conversion matrix is constructed by selecting L row-direction conversion vectors in the forward 64×64 non-separable conversion matrix, and when the L×64 conversion matrix is multiplied by a 64×1 input vector, L conversion coefficients are obtained.
[0346] The L value has a range of 1 <= L < 64. Generally, L out of 64 conversion basis vectors can be selected in any manner. However, as described above, it may be advantageous from the perspective of encoding efficiency to select the conversion basis vectors with high signal energy importance from the aspects of encoding and decoding. The worst-case number of multiplications per sample in an 8×8 block due to changes in the L value is as shown in Table 10 below.
[0347]
Table 10
[0348] As described above, different L values of 8×8 RSTs can be used in combination as shown in Table 11 below for reducing the worst-case multiplication complexity (however, the following example explains the conditions for applying 8×8 RST under the conditions for applying 8×8 NSST, that is, when the width and height of the current block are all greater than or equal to 8).
[0349]
Table 11
[0350] Referring to Table 11, when the width and height of the current block are both 8, an 8×8 RST based on an 8×64 matrix is applied to the current block; otherwise (when either the width or height of the current block is not 8), an 8×8 RST based on a 16×64 matrix can be applied to the current block. More specifically, when the size of the current block is 8×8, a non-separable transform with an input length of 64 and an output length of 8 is applied; otherwise, a non-separable transform with an input length of 64 and an output length of 16 is applied. In the case of the inverse non-separable transform, when the current block is 8×8, a non-separable transform with an input length of 8 and an output length of 64 is applied; otherwise, a non-separable transform with an input length of 16 and an output length of 64 is applied.
[0351] Also, as described in Embodiment 11, since the RST can be applied only to a partial region instead of the entire 8×8 block, for example, when the RST is applied to the remaining region excluding the lower right 4×4 region of the 8×8 block, an 8×8 RST based on an 8×48 or 16×18 matrix can be applied. That is, when the width and height of the current block each correspond to 8, an 8×8 RST based on an 8×48 matrix is applied; otherwise (when the width or height of the current block is not 8), an 8×8 RST based on a 16×48 matrix is applied.
[0352] In the case of the forward non-separable transform, when the current block is 8×8, a non-separable transform with an input length of 48 and an output length of 8 is applied; otherwise, a non-separable transform with an input length of 48 and an output length of 16 is applied.
[0353] In the case of the inverse non-separable transform, when the current block is 8×8, a non-separable transform with an input length of 8 and an output length of 48 is applied; otherwise, a non-separable transform with an input length of 16 and an output length of 48 is applied.
[0354] As a conclusion, when RST is applied to blocks larger than 8×8, based on the encoder side, if the height and width of the block respectively correspond to 8, a non-separable transformation matrix (8×48 or 8×64 matrix) with an input length less than or equal to 64 (e.g., 48 or 64) and an output length less than 64 (e.g., 8) is applied; if the height or width of the block does not correspond to 8, a non-separable transformation matrix (16×48 or 16×64 matrix) with an input length less than or equal to 64 (e.g., 48 or 64) and an output length less than 64 (e.g., 16) is applied.
[0355] Also, when RST is applied to blocks larger than 8×8 based on the decoder side, if the height and width of the block respectively correspond to 8, a non-separable transformation matrix (48×8 or 64×8 matrix) with an input length less than 64 (e.g., 8) and an output length less than or equal to 64 (e.g., 48 or 64) is applied; if the height or width of the block does not correspond to 8, a non-separable transformation matrix (48×16 or 64×16 matrix) with an input length less than 64 (e.g., 16) and an output length less than or equal to 64 (e.g., 48 or 64) is applied.
[0356] Table 12 shows examples regarding the application of various 8×8 RSTs under the condition for applying 8×8 NSST (i.e., when the width and height of the current block are greater than or equal to 8).
[0357]
Table 12
[0358] Referring to Table 12, when the width and height of the current block are each 8, an 8×8 RST based on an 8×64 matrix (or an 8×48 matrix) is applied. When the product of the width and height of the current block is smaller than the threshold value (TH), an 8×8 RST based on a 16×64 matrix (or a 16×48 matrix) is applied to the 8×8 area in the upper left corner of the current block. In the remaining cases (when the width or height of the current block is not 8 and the product of the width and height of the current block is greater than or equal to the threshold value), an 8×8 RST based on a 32×64 matrix (or a 32×48 matrix) is applied to the 8×8 area in the upper left corner of the current block.
[0359] FIG. 29 is an embodiment to which the present invention is applied, and shows an example of an encoding flowchart for performing conversion.
[0360] The encoding device 100 performs primary conversion on the residual block (S2910). The primary conversion may be referred to as core conversion. As an embodiment, the encoding device 100 performs primary conversion using the aforementioned MTS. Further, the encoding device 100 transmits an MTS index indicating a specific MTS among the MTS candidates to the decoding device 200. Here, the MTS candidates are configured based on the intra prediction mode of the current block.
[0361] The encoding device 100 determines whether to apply secondary conversion (S2920). As an example, the encoding device 100 determines whether to apply secondary conversion based on the primary-converted residual conversion coefficients. For example, the secondary conversion can be NSST or RST.
[0362] The encoding device 100 determines the secondary conversion (S2930). At this time, the encoding device 100 determines the secondary conversion based on the NSST (or RST) conversion set specified by the intra prediction mode.
[0363] Also, as an example, the encoding device 100 determines the area to which the secondary conversion is applied based on the size of the current block prior to step S2930.
[0364] The encoding device 100 performs the secondary conversion using the secondary conversion determined in step S2930 (S2940).
[0365] FIG. 30 shows an example of a decoding flowchart for performing conversion in an embodiment to which the present invention is applied.
[0366] The decoding device 200 determines whether to apply the secondary inverse conversion (S3010). For example, the secondary inverse conversion can be NSST or RST. As an example, the decoding device 200 determines whether to apply the secondary inverse conversion based on the secondary conversion flag received from the encoding device 100.
[0367] The decoding device 200 determines the secondary inverse conversion (S3020). Here, the decoding device 200 can determine the secondary inverse conversion to be applied to the current block based on the NSST (or RST) conversion set specified by the intra prediction mode described above.
[0368] Also, as an example, the decoding device 200 determines the area to which the secondary inverse conversion is applied based on the size of the current block prior to step S3020.
[0369] The decoding device 200 performs the secondary inverse conversion on the inverse quantized residual block using the secondary inverse conversion determined in step S3020 (S3030).
[0370] The decoding device 200 performs the primary inverse conversion on the secondary inverse converted residual block (S3040). The primary inverse conversion may be referred to as the core inverse conversion. As an embodiment, the decoding device 200 performs the primary inverse conversion using the MTS described above. Also, as an example, the decoding device 200 can determine whether MTS is applied to the current block prior to step S3040. In this case, the decoding flowchart of FIG. 30 may further include a step of determining whether MTS is applied.
[0371] As an example, when MTS is applied to the current block (i.e., cu_mts_flag = 1), the decoding device 200 constructs MTS candidates based on the intra prediction mode of the current block. In this case, the decoding flowchart of FIG. 30 may further include steps for constructing MTS candidates. Then, the decoding device 200 can determine the primary inverse transform applied to the current block using the mts_idx indicating a specific MTS among the constructed MTS candidates.
[0372] FIG. 31 shows an example of a detailed block diagram of the conversion unit 120 in the encoding device 100 according to an embodiment to which the present invention is applied.
[0373] The encoding device 100 to which the embodiment of the present invention is applied includes a primary conversion unit 3110, a secondary conversion applicability determination unit 3120, a secondary conversion determination unit 3130, and a secondary conversion unit 3140.
[0374] The primary conversion unit 3110 can perform a primary conversion on the residual block. The primary conversion may be referred to as a core conversion. As an embodiment, the primary conversion unit 3110 performs the primary conversion using the aforementioned MTS. Also, the primary conversion unit 3110 transmits an MTS index indicating a specific MTS among the MTS candidates to the decoding device 200. Here, the MTS candidates are constructed based on the intra prediction mode of the current block.
[0375] The secondary conversion applicability determination unit 3120 can determine whether to apply the secondary conversion. As an example, the secondary conversion applicability determination unit 3120 can determine whether to apply the secondary conversion based on the conversion coefficients of the primarily converted residual block. For example, the secondary conversion can be NSST or RST.
[0376] The secondary conversion determination unit 3130 determines the secondary conversion. At this time, the secondary conversion determination unit 3130 determines the secondary conversion based on the NSST (or, RST) conversion set specified by the intra prediction mode as described above.
[0377] Also, as an example, the secondary conversion determination unit 3130 can determine the area to which the secondary conversion is applied based on the size of the current block.
[0378] The secondary conversion unit 3140 can perform the secondary conversion using the determined secondary conversion.
[0379] FIG. 32 shows an example of a detailed block diagram of the inverse conversion unit 230 in the decoding apparatus 200, which is an embodiment to which the present invention is applied.
[0380] The decoding apparatus 200 to which the present invention is applied includes a secondary inverse conversion applicability determination unit 3210, a secondary inverse conversion determination unit 3220, a secondary inverse conversion unit 3230, and a primary inverse conversion unit 3240.
[0381] The secondary inverse conversion applicability determination unit 3210 determines whether to apply the secondary inverse conversion. For example, the secondary inverse conversion can be NSST or RST. As an example, the secondary inverse conversion applicability determination unit 3210 determines whether to apply the secondary inverse conversion based on the secondary conversion flag received from the encoding apparatus 100. As another example, the secondary inverse conversion applicability determination unit 3210 can also determine whether to apply the secondary inverse conversion based on the conversion coefficients of the residual block.
[0382] The secondary inverse conversion determination unit 3220 determines the secondary inverse conversion. At this time, the secondary inverse conversion determination unit 3220 determines the secondary inverse conversion to be applied to the current block based on the NSST (or RST) conversion set specified by the intra prediction mode.
[0383] Also, as an example, the secondary inverse conversion determination unit 3220 can determine the area to which the secondary inverse conversion is applied based on the size of the current block.
[0384] Also, as an example, the secondary inverse conversion unit 3230 can perform the secondary inverse conversion on the inverse quantized residual block using the determined secondary inverse conversion.
[0385] The primary inverse transform unit 3240 performs a primary inverse transform on the secondarily inverse-transformed residual block. As an embodiment, the primary inverse transform unit 3240 performs the primary transform using the aforementioned MTS. Also, as an example, the primary inverse transform unit 3240 can determine whether the MTS is applied to the current block.
[0386] As an example, when the MTS is applied to the current block (i.e., cu_mts_flag = 1), the primary inverse transform unit 3240 constructs MTS candidates based on the intra prediction mode of the current block. Then, the primary inverse transform unit 3240 determines the primary transform applied to the current block using the mts_idx indicating a specific MTS among the constructed MTS candidates.
[0387] FIG. 33 shows an example of a decode flowchart to which the transform according to the embodiment of the present invention is applied. The operation in FIG. 33 is performed by the inverse transform unit 230 of the decoding device 100.
[0388] In step S3305, the decoding device 200 determines the input length and output length of the non-separable transform based on the height and width of the current block. Here, when the height and width of the current block are both 8, the input length of the non-separable transform is 8, and the output length is determined to be a value greater than the input length and less than or equal to 64 (e.g., 48 or 64). For example, when the non-separable transform is applied to the entire transform coefficients of an 8×8 block on the encoder side, the output length is determined to be 64, and when the non-separable transform is applied to a part of the transform coefficients of an 8×8 block (e.g., the part excluding the lower-right 4×4 region of the 8×8 block), the output length is determined to be 48.
[0389] In step S3310, the decoding device 200 determines a non-separable transform matrix corresponding to the input length and output length of the non-separable transform. For example, when the input length of the non-separable transform is 8 and the output length is 48 or 64 (when the size of the current block is 4×4), a 48×8 or 64×8 matrix derived from the transform kernel is determined as the non-separable transform. When the input length of the non-separable transform is 16 and the output length is 48 or 64 (for example, when the current block is smaller than 8×8 and not 4×4), a 48×16 or 64×16 transform kernel can be determined as non-separable.
[0390] According to an embodiment of the present invention, the decoding device 200 determines a non-separable transform set index (for example, an NSST index) based on the intra prediction mode of the current block, determines a non-separable transform kernel corresponding to the non-separable transform index within the non-separable transform set included in the non-separable transform set index, and can determine a non-separable transform matrix from the non-separable transform kernel based on the input length and output length determined in step S3305.
[0391] In step S3315, the decoding device 200 applies the non-separable transform matrix determined for the current block to the coefficients corresponding to the input length (8 or 16) determined for the current block. For example, when the input length of the non-separable transform is 8 and the output length is 48 or 64, the 48×8 or 64×8 matrix derived from the transform kernel is applied to the 8 coefficients included in the current block. When the input length of the non-separable transform is 16 and the output length is 48 or 64, the 48×16 or 64×16 matrix derived from the transform kernel can be applied to the 16 coefficients in the 4×4 area at the upper left of the current block. Here, the coefficients to which the non-separable transform is applied are the coefficients up to the position corresponding to the input length (for example, 8 or 16) in the scan order determined from the DC position of the current block (for example, (a), (b), or (c) in FIG. 16).
[0392] On the other hand, when it does not correspond to the case where the current block height and width are both 8, if the product of the width and height of the current block is less than the threshold value, the decoding device 200 inputs the 16 coefficients in the upper left 4×4 region in the current block and outputs a non-separable conversion matrix (48×16 or 64×16 matrix) that outputs the converted coefficients for the output length (e.g., 48 or 64). If the product of the width and height of the current block is greater than or equal to the threshold value, the decoding device 200 inputs 32 coefficients in the current block and outputs a non-separable conversion matrix (48×32 or 64×32 matrix) that outputs the converted coefficients for the output length (e.g., 48 or 64).
[0393] When the output length is 64, 64 converted data (converted coefficients) to which non-separable conversion is applied to an 8×8 block are arranged by applying the non-separable conversion matrix. When the output length is 48, 48 converted data (converted coefficients) to which non-separable conversion is applied to the remaining region excluding the lower right 4×4 region in the 8×8 block are arranged by applying the non-separable conversion matrix.
[0394] FIG. 34 shows an example of a block diagram of an apparatus for processing a video signal, which is an embodiment to which the present invention is applied. The image processing apparatus 3400 in FIG. 34 may correspond to the encoding apparatus 100 in FIG. 1 or the decoding apparatus 200 in FIG. 2.
[0395] The image processing apparatus 3400 for processing an image signal includes a memory 3420 for storing the image signal and a processor 3410 coupled to the memory and processing the image signal.
[0396] The processor 3410 according to an embodiment of the present invention is composed of at least one processing circuit for processing an image signal and can process the image signal by executing instruction words for encoding or decoding the image signal. That is, the processor 3410 can encode the original image data or decode the encoded image signal by executing the above-described encoding or decoding method.
[0397] FIG. 35 shows an example of an image coding system, which is an embodiment to which the present invention is applied.
[0398] The image coding system includes a source device and a receiving device. The source device transmits encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.
[0399] The source device includes a video source, an encoding device, and a transmitter. The receiving device includes a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0400] The video source acquires video / images through processes such as video / image capture, synthesis, or generation. The video source includes a video / image capture device and / or a video / image generation device. The video / image capture device includes, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device includes, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer, etc., and in this case, the video / image capture process can be replaced by the process of generating related data.
[0401] The encoding device encodes the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) is output in the form of a bitstream.
[0402] The transmitting unit transmits the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device in a file or streaming form via a digital storage medium or a network. The digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit includes elements for generating a media file in a predetermined file format and elements for transmission via a broadcast / communication network. The receiver extracts the bitstream and transmits it to the decoding device.
[0403] The decoding device decodes the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0404] The renderer renders the decoded video / image. The rendered video / image is displayed via the display unit.
[0405] FIG. 36 is a structural diagram of a content streaming system according to an embodiment to which the present invention is applied.
[0406] The content streaming system to which the present invention is applied includes an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0407] The encoding server serves to compress the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server may be omitted.
[0408] The bitstream is generated by an encoding method or a bitstream generation method to which the present invention is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0409] The streaming server transmits multimedia data to the user device based on a user request via the web server, and the web server serves as a medium for informing the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. Here, the content streaming system may include a separate control server, and in this case, the control server plays a role of controlling commands / responses between each device in the content streaming system.
[0410] The streaming server receives content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0411] Examples of user devices can include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head mounted displays)), digital TVs, desktop computers, digital signage, and the like.
[0412] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed distributively.
[0413] In addition, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable storage medium. Multimedia data having the data structure according to the present invention can also be stored in a recording medium readable by a computer. The recording medium readable by the computer includes all kinds of storage devices and distributed storage devices in which computer-readable data is stored. The recording medium readable by the computer can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Further, the recording medium readable by the computer includes a medium realized in the form of a carrier wave (for example, transmission through the Internet). Also, a bit stream generated by an encoding method can be stored in a recording medium readable by a computer or transferred via a wired / wireless communication network.
[0414] In addition, the embodiments of the present invention can be realized as a computer program product by program code, and the program code can be executed on a computer according to the embodiments of the present invention. The program code can be stored on a carrier readable by a computer.
[0415] As described above, the embodiments described in the present invention can be realized and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each figure can be realized and executed on a computer, a processor, a microprocessor, a controller, or a chip.
[0416] In addition, the decoder and encoder to which the present invention is applied can be included in real-time communication devices such as multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conferencing devices, video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, videophones, video devices, and medical video devices, and are used to process video signals or data signals. For example, an over-the-top (OTT) video device can include game machines, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), and the like.
[0417] In addition, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable storage medium. Multimedia data having a data structure according to the present invention can also be stored in a computer-readable storage medium. The computer-readable recording medium includes any type of storage device and distributed storage device in which computer-readable data is stored. The computer-readable recording medium can include, for example, Blu-ray Discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, the computer-readable recording medium includes a medium realized in the form of a carrier wave (for example, transmission through the Internet). Further, the bitstream generated by the encoding method can be stored in a computer-readable recording medium or transferred via a wired / wireless communication network.
[0418] In addition, embodiments of the present invention can be realized as a computer program product by program code, and the program code can be executed on a computer according to embodiments of the present invention. The program code can be stored on a carrier readable by a computer.
[0419] The embodiments described above are those in which the components and features of the present invention are combined in a predetermined form. Each component or feature should be considered optional unless otherwise explicitly mentioned. Each component or feature can be implemented in a form not combined with other components or features. Also, it is possible to configure embodiments of the present invention by combining some components and / or features. The order of operations described in the embodiments of the present invention can be changed. Some configurations and features of any embodiment can be included in other embodiments or can be replaced with corresponding configurations or features of other embodiments. It is obvious that embodiments can be configured by combining claims without an explicit citation relationship in the claims, or can be included as new claims by amendment after filing.
[0420] Embodiments according to the present invention can be realized by various means, such as hardware, firmware, software, or combinations thereof. In the case of realization by hardware, an embodiment of the present invention can be realized by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and the like.
[0421] In the case of implementation by firmware or software, an embodiment of the present invention can be implemented in the form of modules, procedures, functions, etc. that execute the functions or operations described above. The software code can be stored in a memory and driven by a processor. The memory can be located inside or outside the processor and can transmit and receive data to and from the processor by various means already known.
[0422] It is obvious to those skilled in the art that the present invention can be embodied in other specific forms without departing from the essential features of the present invention. Therefore, the above detailed description should not be construed restrictively in all aspects and should be regarded as exemplary. The scope of the present invention should be determined by a reasonable interpretation of the appended claims, and all changes within the equivalent scope of the present invention are included in the scope of the present invention.
Industrial Applicability
[0423] As described above, the preferred embodiments of the present invention described above are disclosed for illustrative purposes, and those skilled in the art can improve, change, substitute, or add various other embodiments within the technical idea and technical scope of the present invention disclosed in the appended patent claims below.
Claims
In a method for a device to decode an image signal, determining an input length and an output length of a non-separable transform based on a height and a width of a current block; determining a non-separable transform matrix related to the input length and the output length of the non-separable transform; applying the non-separable transform matrix to coefficients of the current block, wherein a number of the coefficients is related to the input length of the non-separable transform; performing an inverse primary transform on the coefficients to which the non-separable transform has been applied, the method comprising: the input length and the output length of the non-separable transform are determined individually; based on the height and the width of the current block, a size of the non-separable transform matrix is determined as one of four predetermined sizes; based on both the height and the width of the current block being greater than 8, the input length of the non-separable transform is determined to be 16, and the output length of the non-separable transform is greater than the input length of the non-separable transform; based on both the height and the width of the current block being equal to 8, the input length of the non-separable transform is determined to be 8, and the output length of the non-separable transform is greater than the input length of the non-separable transform. In a method for a device to encode an image signal, performing a primary transform on a current block; determining an input length and an output length of a non-separable transform based on a height and a width of the current block; determining a non-separable transform matrix for the current block, wherein a size of the non-separable transform matrix is determined based on the input length and the output length of the non-separable transform; applying the non-separable transform matrix to coefficients of the primarily transformed current block, wherein a number of the coefficients is related to the input length of the non-separable transform; encoding non-separable transform index information for the non-separable transform matrix for the current block, the method comprising: the input length and the output length of the non-separable transform are determined individually; based on the height and the width of the current block, the size of the non-separable transform matrix is determined as one of four predetermined sizes; based on both the height and the width of the current block being greater than 8, the output length of the non-separable transform is determined to be 16, and the input length of the non-separable transform is greater than the output length of the non-separable transform; A method in which, based on both the height and the width of the current block being equal to 8, the output length of the non-separable transform is determined to be 8, and the input length of the non-separable transform is greater than the output length of the non-separable transform. A method of transmitting a bitstream generated by a method for encoding an image signal by an apparatus, wherein the method for encoding the image signal comprises: performing a primary transform on a current block; determining an input length and an output length of a non-separable transform based on the height and width of the current block; determining a non-separable transform matrix for the current block, wherein the size of the non-separable transform matrix is determined based on the input length and the output length of the non-separable transform; applying the non-separable transform matrix to the coefficients of the primarily transformed current block, wherein the number of the coefficients is related to the input length of the non-separable transform; encoding non-separable transform index information for the non-separable transform matrix for the current block into the bitstream; the input length and the output length of the non-separable transform are determined individually; based on the height and width of the current block, the size of the non-separable transform matrix is determined as one of four predetermined sizes; A method in which, based on both the height and the width of the current block being greater than 8, the output length of the non-separable transform is determined to be 16, and the input length of the non-separable transform is greater than the output length of the non-separable transform. A method in which, based on both the height and the width of the current block being equal to 8, the output length of the non-separable transform is determined to be 8, and the input length of the non-separable transform is greater than the output length of the non-separable transform.
Citation Information
Patent Citations
Reduced size inverse transform for decoding and encoding
US20170034530A1
Method and apparatus for video coding
US20200304782A1
Non-separable secondary transform for video coding
WO2017058614A1