Video encoding / decoding method and apparatus based on non-separable linear transform, and recording medium for storing bitstreams

The implementation of non-separable linear transforms in video encoding/decoding methods enhances efficiency for high-resolution video, addressing the need for cost-effective transmission and storage.

JP2025533227APending Publication Date: 2025-10-03LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025521030
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-10-12
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality video has led to a need for more efficient video compression techniques to reduce transmission and storage costs.

Method used

A video encoding/decoding method and apparatus utilizing non-separable linear transforms, including transform information for current blocks, and the application of non-separable and separable transformation kernels based on block characteristics, to enhance encoding/decoding efficiency.

Benefits of technology

Improves encoding/decoding efficiency and enables effective transmission and storage of high-resolution, high-quality video by utilizing non-separable linear transforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533227000001_ABST
    Figure 2025533227000001_ABST
Patent Text Reader

Abstract

A video encoding / decoding method and apparatus according to the present disclosure may include obtaining transform information including information about a non-separable linear transform, obtaining transform coefficients for a current block, and generating a residual block by inverse transforming the transform coefficients based on the transform information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a video encoding / decoding method, an apparatus, and a recording medium for storing a bitstream, and more particularly to a video encoding / decoding method, an apparatus, and a recording medium for storing a bitstream generated by the video encoding method / apparatus of the present disclosure. [Background technology]

[0002] In recent years, demand for high-resolution, high-quality video, such as HD (High Definition) video and UHD (Ultra High Definition) video, has been increasing in various fields. As video data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases compared to existing video data. The increase in the amount of information or bits to be transmitted leads to an increase in transmission costs and storage costs.

[0003] Therefore, there is a demand for a highly efficient video compression technique for effectively transmitting, storing, and reproducing high-resolution, high-quality video information. Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide a video encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide a video encoding / decoding method and apparatus that performs intra prediction mode.

[0006] Another object of the present disclosure is to provide a video encoding / decoding method and apparatus that performs inter prediction mode.

[0007] Another object of the present disclosure is to provide a video encoding / decoding method and apparatus based on a non-separable linear transform.

[0008] Another object of the present disclosure is to provide a non-transitory computer-readable recording medium that stores a bitstream generated by the video encoding method or apparatus according to the present disclosure.

[0009] Another object of the present disclosure is to provide a non-transitory computer-readable recording medium that stores a bitstream that is received and decoded by a video decoding device according to the present disclosure and is used to restore a video.

[0010] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the video encoding method or apparatus according to the present disclosure.

[0011] The technical problems to be solved by the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the following description. [Means for solving the problem]

[0012] According to one embodiment of the present disclosure, a video decoding method performed in a video decoding device may include (may comprise; may configure; may construct; may set; may encompass; may contain; may include) a step of obtaining transform information including information about a non-separable linear transform, a step of obtaining transform coefficients for a current block, and a step of generating a residual block by inverse transforming the transform coefficients based on the transform information.

[0013] According to one embodiment of the present disclosure, the transformation information includes a non-separable primary transformation activation flag indicating whether a non-separable primary transformation is possible for the current block, and indicates that a separable primary transformation is possible for the current block when the value of the non-separable primary transformation activation flag is a first value, and indicates that either a non-separable primary transformation or a separable primary transformation is possible for the current block when the value of the non-separable primary transformation activation flag is a second value.

[0014] According to one embodiment of the present disclosure, based on at least one of the shape of the current block or the size of the current block, the transformation information may include a non-separable linear transformation application flag indicating whether a non-separable linear transformation is applied to the current block.

[0015] According to one embodiment of the present disclosure, the transformation information may include a kernel index indicating one of a plurality of transformation kernels, and the plurality of transformation kernels may include at least one non-separable transformation kernel and at least one separable transformation kernel.

[0016] According to one embodiment of the present disclosure, the transformation information may include a non-separable linear transform application flag indicating whether a non-separable linear transform is applied to the current block, and a non-separable transform kernel index indicating one of a plurality of non-separable transform kernels may be obtained based on the non-separable linear transform application flag indicating that a non-separable linear transform is applied to the current block, and a separable transform kernel index indicating one of a plurality of separable transform kernels may be obtained based on the non-separable linear transform application flag indicating that a non-separable linear transform is not applied to the current block.

[0017] According to one embodiment of the present disclosure, the number of the non-separable transform kernels or the separated transform kernels may be obtained based on at least one of whether intra prediction is applied to the current block, the intra prediction mode, the block size, the surrounding sample values, whether a secondary transform is used, or the quantization parameter.

[0018] According to an embodiment of the present disclosure, based on the fact that a non-separable linear transform is applied to the current block, the secondary transform for the current block may be restricted to one that is not a non-separable transform.

[0019] According to one embodiment of the present disclosure, information regarding the non-separable linear transform may be parsed before information regarding the non-separable secondary transform for the current block, and parsing of the information regarding the non-separable secondary transform may be skipped based on the fact that a non-separable linear transform is applied to the current block.

[0020] According to an embodiment of the present disclosure, based on the fact that a non-separable secondary transform is applied to the current block, the linear transform for the current block may be restricted to one that is not a non-separable transform.

[0021] According to one embodiment of the present disclosure, information regarding a non-separable secondary transform for the current block may be parsed before information regarding the non-separable primary transform, and parsing of the information regarding the non-separable secondary transform may be skipped based on the fact that a non-separable secondary transform is applied to the current block.

[0022] According to one embodiment of the present disclosure, the information about the non-separable linear transform may be coded based on characteristics of residual coefficients.

[0023] According to one embodiment of the present disclosure, a video encoding method performed in a video encoding device may include the steps of obtaining a residual block for a current block, encoding transform information including information about a non-separable linear transform, and generating transform coefficients by transforming the residual block based on the transform information.

[0024] According to an embodiment of the present disclosure, a computer-readable recording medium may be provided that stores a bitstream generated by a video encoding method.

[0025] According to one embodiment of the present disclosure, in a method for transmitting a bitstream generated by a video encoding method, the video encoding method may include a step of obtaining a residual block for a current block, a step of encoding transform information including information about a non-separable linear transform, and a step of generating transform coefficients by transforming the residual block based on the transform information. [Effects of the Invention]

[0026] According to the present disclosure, it is possible to provide a video encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0027] Furthermore, according to the present disclosure, it is possible to provide a video encoding / decoding method and apparatus that performs intra prediction mode.

[0028] Furthermore, according to the present disclosure, it is possible to provide a video encoding / decoding method and apparatus that performs inter prediction mode.

[0029] Furthermore, the present disclosure can provide a video encoding / decoding method and apparatus based on a non-separable linear transform.

[0030] Furthermore, according to the present disclosure, it is possible to provide a non-transitory computer-readable recording medium that stores a bitstream generated by the video encoding method or apparatus according to the present disclosure.

[0031] In addition, according to the present disclosure, it is possible to provide a non-transitory computer-readable recording medium that stores a bitstream that is received and decoded by a video decoding device according to the present disclosure and used to restore a video.

[0032] Furthermore, according to the present disclosure, it is possible to provide a method for transmitting a bitstream generated by a video encoding method or apparatus according to the present disclosure.

[0033] The effects obtained by the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief explanation of the drawings]

[0034] [Figure 1] 1 is a schematic diagram illustrating a video coding system to which embodiments of the present disclosure can be applied; [Figure 2] 1 is a schematic diagram illustrating a video encoding device to which an embodiment of the present disclosure can be applied. [Figure 3] FIG. 1 is a schematic diagram illustrating a video decoding device to which an embodiment of the present disclosure can be applied. [Figure 4] 1 is a flowchart of a video encoding method according to the present disclosure. [Figure 5] 1 is a flowchart of a video decoding method according to the present disclosure. [Figure 6] FIG. 1 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION

[0035] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement the present disclosure, but the present disclosure may be embodied in various other forms and is not limited to the embodiments described herein.

[0036] In describing the embodiments of the present disclosure, if it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.

[0037] In this disclosure, when one component is "coupled," "coupled," or "connected" to another component, this may include not only a direct connection, but also an indirect connection where there is another component between them. Furthermore, when one component is described as "including" or "having" another component, this does not exclude the other component, but means that the other component may also be included, unless otherwise specified.

[0038] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0039] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component and do not necessarily mean that the components are separate. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not specifically stated, such integrated or distributed embodiments are also included within the scope of this disclosure.

[0040] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.

[0041] This disclosure relates to video encoding and decoding, and terms used in this disclosure may have ordinary meanings commonly used in the field of technology to which this disclosure pertains unless they are newly defined in this disclosure.

[0042] In this disclosure, "video" can refer to a collection of images over time.

[0043] In this disclosure, a "picture" generally refers to a unit representing one video image in a specific time period, a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. In addition, a slice / tile may include one or more coding tree units (CTUs).

[0044] In this disclosure, a "pixel" or a "pel" may refer to the smallest unit constituting one picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component, or may represent only a pixel / pixel value of a chroma component.

[0045] In this disclosure, a "unit" may refer to a basic unit of video processing. A unit may include at least one of a specific region of a picture and information related to that region. A unit may also be referred to as a "sample array," "block," or "area," depending on the situation. In general, an MxN block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0046] In this disclosure, a "current block" may refer to one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." When prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." When filtering is performed, a "current block" may refer to a "block to be filtered."

[0047] In this disclosure, unless explicitly stated as a chroma block, the term "current block" can refer to a block including both a luma component block and a chroma component block, or to the "luma block of the current block." The luma component block of the current block may be explicitly expressed as a "luma block" or a "current luma block," including the explicit statement that it is a luma component block. Also, the chroma component block of the current block may be explicitly expressed as a "chroma block" or a "current chroma block," including the explicit statement that it is a chroma component block.

[0048] In the present disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" may mean "at least one of A, B, and / or C."

[0049] In this disclosure, "or" may be interpreted as "and / or." For example, "A or B" can mean 1) "A" only, 2) "B" only, or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."

[0050] In this disclosure, "at least one of A, B, and C" can mean "just A," "just B," "just C," or "any combination of A, B, and C." Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C."

[0051] Parentheses used in the present disclosure may mean "for example." For example, when "prediction (intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction." In other words, "prediction" in the present disclosure is not limited to "intra prediction," and "intra prediction" may be suggested as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction."

[0052] Video Coding System Overview FIG. 1 is a schematic diagram illustrating a video coding system to which embodiments of the present disclosure can be applied.

[0053] A video coding system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0054] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / video encoding unit, and the decoding unit 22 may be referred to as a video / video decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be a separate device or an external component.

[0055] The video source generation unit 11 may acquire video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated by a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.

[0056] The encoder 12 may encode input video / image data. The encoder 12 may perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency. The encoder 12 may output encoded data (encoded video / image information) in the form of a bitstream.

[0057] The transmitter 13 may acquire encoded video / image information or data output in the form of a bitstream and transmit it to the receiver 21 of the decoding device 20 or another external object in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter 13 may include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The transmitter 13 may be provided as a transmitting device separate from the encoder 12. In this case, the transmitting device may include at least one processor for acquiring encoded video / image information or data output in the form of a bitstream and a transmitter for transmitting the same in the form of a file or streaming. The receiver 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoder 22.

[0058] The decoding unit 22 can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.

[0059] The rendering unit 23 can render the decoded video / image, and the rendered video / image can be displayed on a display unit.

[0060] Overview of video encoding equipment FIG. 2 is a schematic diagram illustrating a video encoding device to which an embodiment of the present disclosure can be applied.

[0061] 2, the video encoding device 100 may include a video division unit 110, a subtraction unit 115, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transform unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transform unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0062] Depending on the embodiment, all or at least some of the components constituting the video encoding device 100 may be implemented as a single hardware component (e.g., an encoder or a processor). Also, the memory 170 may include a decoded picture buffer (DPB) and may be implemented as a digital storage medium.

[0063] The video division unit 110 may divide an input video (or picture or frame) input to the video encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) using a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be divided into multiple coding units at deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary tree structure. To divide the coding units, a quad-tree structure may be applied first, and then a binary tree structure and / or a ternary tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is not further divided. The largest coding unit may be directly used as the final coding unit, or a lower-depth coding unit obtained by dividing the largest coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may each be divided or partitioned from the final coding unit. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0064] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a current block (current block) to be processed and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit may generate various information related to the prediction of the current block and transmit it to the entropy encoding unit 190. The prediction information may be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0065] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the intra prediction mode and / or intra prediction method. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the accuracy of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0066] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be referred to as collocated reference blocks, collocated control units (colCUs), etc. The reference picture including the temporal neighboring blocks may be referred to as a collocated picture (colPic). For example, the inter predictor 180 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 180 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0067] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the predictor may apply intra prediction or inter prediction to predict the current block, or may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) to predict the current block. Intra block copy may be used, for example, for coding content images / videos such as games, such as screen content coding (SCC). IBC is a method of predicting a current block using an already reconstructed reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but may be performed similarly to inter prediction in that a reference block is derived within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure.

[0068] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input video signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.

[0069] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, the GBT refers to a transform obtained from a graph when relationship information between pixels is expressed as a graph. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to blocks of variable size other than a square.

[0070] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy encoding unit 190. The entropy encoding unit 190 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0071] The entropy encoding unit 190 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit 190 may encode information required for video / image restoration (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / video information) may be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) units. The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The signaling information, transmitted information, and / or syntax elements referred to in this disclosure may be encoded according to the encoding procedures described above and included in the bitstream.

[0072] The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 may be provided as an internal / external element of the video encoding device 100, or the transmitter may be provided as a component of the entropy encoding unit 190.

[0073] The quantized transform coefficients output from the quantization unit 130 may be used to generate a residual signal. For example, the quantized transform coefficients may be subjected to inverse quantization and inverse transformation in the inverse quantization unit 140 and the inverse transform unit 150, respectively, to reconstruct a residual signal (residual block or residual sample).

[0074] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to a prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.

[0075] Meanwhile, luma mapping with chroma scaling (LMCS) may be applied during the picture encoding and / or reconstruction process.

[0076] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering and transmit it to the entropy encoding unit 190, as will be described later in the description of each filtering method. The filtering information may be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0077] The modified reconstructed picture transmitted to the memory 170 may be used as a reference picture in the inter prediction unit 180. This allows the video encoding device 100 to avoid prediction mismatch between the video encoding device 100 and the video decoding device when inter prediction is applied, and also improves encoding efficiency.

[0078] The DPB in the memory 170 may store a modified reconstructed picture to be used as a reference picture in the inter prediction unit 180. The memory 170 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter prediction unit 180 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra prediction unit 185.

[0079] Overview of the video decoder FIG. 3 is a schematic diagram illustrating a video decoding device to which an embodiment of the present disclosure can be applied.

[0080] 3, the video decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.

[0081] Depending on the embodiment, all or at least some of the components constituting the video decoding device 200 may be implemented as a single hardware component (e.g., a decoder or a processor). Also, the memory 170 may include a DPB and may be implemented as a digital storage medium.

[0082] The video decoding apparatus 200, which receives a bitstream including video / image information, may reconstruct an image by performing a process corresponding to the process performed by the video encoding apparatus 100 of FIG. 2. For example, the video decoding apparatus 200 may perform decoding using a processing unit applied in the video encoding apparatus 100. Therefore, the decoding processing unit may be, for example, a coding unit. The coding unit may be a coding tree unit or may be obtained by dividing a maximum coding unit. The reconstructed video signal decoded and output by the video decoding apparatus 200 may be reproduced by a playback device (not shown).

[0083] The video decoding apparatus 200 may receive a signal output from the video encoding apparatus 100 of FIG. 2 in the form of a bitstream. The received signal may be decoded by the entropy decoding unit 210. For example, the entropy decoding unit 210 may parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The video decoding apparatus 200 may further use the information on the parameter sets and / or the general constraint information to decode the video. Signaling information, received information, and / or syntax elements referred to in this disclosure may be obtained from the bitstream by being decoded by the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded, decoding information on neighboring blocks and the block to be decoded, or information on symbols / bins decoded in a previous step, predicts the occurrence probability of bins according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element.In this case, after determining a context model, the CABAC entropy decoding method may update the context model for the next symbol / bin using information about the decoded symbol / bin. Prediction information from the information decoded by the entropy decoding unit 210 may be provided to a prediction unit (the inter prediction unit 260 and the intra prediction unit 265), and residual values ​​entropy decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. In addition, filtering information from the information decoded by the entropy decoding unit 210 may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the video encoding device 100 may be further provided as an internal / external element of the video decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0084] Meanwhile, the video decoding apparatus 200 according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The video decoding apparatus 200 may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.

[0085] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the video encoding device 100. The inverse quantization unit 220 may inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0086] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0087] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction method).

[0088] As mentioned in the description of the prediction unit of the video encoding device 100, the prediction unit can generate a prediction signal based on various prediction methods (techniques) described below.

[0089] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265.

[0090] The inter prediction unit 260 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 260 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes (methods), and the prediction information may include information indicating the inter prediction mode (method) for the current block.

[0091] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The description of the adder 155 may also apply to the adder 235. The adder 235 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, or may be used for inter prediction of the next picture after undergoing filtering, as described below.

[0092] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in a DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0093] The (modified) reconstructed picture stored in the DPB of the memory 250 may be used as a reference picture in the inter prediction unit 260. The memory 250 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter prediction unit 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 250 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra prediction unit 265.

[0094] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the video encoding device 100 may also be applied identically or correspondingly to the filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of the video decoding device 200, respectively.

[0095] Conversion / Reverse Conversion As described above, the video encoding apparatus 100 may derive residual blocks (residual samples) based on blocks (prediction samples) predicted by intra / inter / IBC prediction, etc., and may derive quantized transform coefficients by applying transform and quantization to the derived residual samples. Information on the quantized transform coefficients (residual information) may be included in a residual coding syntax and output in the form of a bitstream after encoding.

[0096] The video decoding apparatus 200 may obtain information (residual information) about the (quantized) transform coefficients from the bitstream and decode the information to derive quantized transform coefficients. The video decoding apparatus 200 may derive residual samples by performing inverse quantization / inverse transform based on the quantized transform coefficients. As described above, at least one of the quantization / inverse quantization and / or transform / inverse transform may be omitted. If the quantization / inverse quantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. If the transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for consistency of expression. Whether the transform / inverse transform is omitted may be signaled based on transform_skip_flag.

[0097] Furthermore, in this disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about transform coefficients, and the information about the transform coefficients may be signaled using a residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients may be derived by inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on inverse transform (transform) of the scaled transform coefficients. This may be equally applied / expressed in other parts of this disclosure.

[0098] The transform / inverse transform may be performed based on a transform kernel. For example, according to the present disclosure, a multiple transform selection (MTS) scheme may be applied. In this case, a subset of a set of transform kernels may be selected and applied to the current block. The transform kernel may be referred to by various terms, such as a transform matrix or a transform type. For example, the set of transform kernels may represent a combination of a vertical transform kernel (vertical transform kernel) and a horizontal transform kernel (horizontal transform kernel).

[0099] For example, MTS index information (or mts_idx syntax element) may be generated / encoded by an encoding device to indicate one of the transform kernel sets and signaled to a decoding device. For example, the transform kernel set according to the value of the MTS index information may be derived as shown in Table 1 below.

[0100] [Table 1]

[0101] The transform kernel set may be determined based on, for example, cu_sbt_horizontal_flag and cu_sbt_pos_flag. cu_sbt_horizontal_flag may indicate that the current block is divided horizontally into two transform blocks when it has a value of 1, and that the current block is divided vertically into two transform blocks when it has a value of 0. cu_sbt_pos_flag may indicate that tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr for the first transform block of the current block are not present in the bitstream when it has a value of 1, and that tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr for the second transform block of the current block are not present in the bitstream when it has a value of 0. Table 2 below shows trTypeHor and trTypeVer according to cu_sbt_horizontal_flag and cu_sbt_pos_flag.

[0102] [Table 2]

[0103] In the above table, trTypeHor may indicate a horizontal transform kernel and trTypeVer may indicate a vertical transform kernel, where a trTypeHor / trTypeVer value of 0 may indicate DCT2, a trTypeHor / trTypeVer value of 1 may indicate DST7, and a trTypeHor / trTypeVer value of 2 may indicate DCT8, although this is an example and other values ​​may be mapped to other DCTs / DSTs by convention.

[0104] Table 3 below shows exemplary basis functions for the above-mentioned DCT2, DCT8, and DST7.

[0105] [Table 3]

[0106] In the present disclosure, the MTS-based transform is applied as a primary transform, and a secondary transform may be further applied. The secondary transform may be applied only to coefficients in the upper left wxh region of the coefficient block to which the primary transform is applied, and may be referred to as a reduced secondary transform (RST). For example, w and / or h may be 4 or 8. In the transform, the primary transform and the secondary transform may be sequentially applied to the residual block, and in the inverse transform, the inverse secondary transform and the inverse primary transform may be sequentially applied to the transform coefficients. The secondary transform (RST transform) may be referred to as a low frequency coefficients transform (LFCT) or a low frequency non-separable transform (LFNST). The inverse secondary transform may be referred to as an inverse LFCT or an inverse LFNST.

[0107] Hereinafter, video encoding / decoding methods according to various embodiments of the present disclosure will be described in detail.

[0108] Example 1

[0109] According to the present disclosure, a method of first-order transformation from the spatial (pixel sample value) domain to the frequency domain using a predefined transform may include a separable transform and a non-separable transform. Alternatively, a transform skip that does not apply a transform to the current block may also be considered as a transform in a broad sense.

[0110] Since image pixels are values ​​(pixel values ​​or sample values) existing in a two-dimensional space (positions having horizontal and vertical coordinates), when a transformation process is performed to convert these values ​​into another domain (i.e., the frequency domain) using a specific basis vector, the complexity required can vary greatly depending on the transformation method. In the case of a separable transformation, transformations in the horizontal and vertical directions may be performed independently. Therefore, the similarities and characteristics of pixel values ​​(sample values) existing in the horizontal or vertical direction may be expressed by the transformation. In this case, because a separable transformation performs transformations in the horizontal and vertical directions independently, the size of the required basis vectors is smaller than that of a non-separable transformation, and the required computational complexity is relatively low.

[0111] Meanwhile, in the case of a non-separable transform, a basis vector of a size corresponding to the total number of pixels in the two-dimensional space may be used to grasp the overall characteristics of pixel values ​​in the two-dimensional space. As a result, the complexity required for a non-separable transform may be greater than that required for a separable transform. However, a non-separable transform can very well represent the similarity between pixels in the two-dimensional space. Therefore, a non-separable transform can provide improved coding efficiency.

[0112] By appropriately selecting and using both the separable transform and the non-separable transform in consideration of the characteristics of the separable transform and the non-separable transform, it is possible to improve the compression performance of the video encoding device 100 and / or the video decoding device 200. Furthermore, by appropriately selecting and using both the separable transform and the non-separable transform, the video encoding device 100 and / or the video decoding device 200 can relatively efficiently maintain the required computational complexity.

[0113] According to an embodiment of the present disclosure, the present disclosure may provide a method of always applying only a separable transform as a primary transform method. Alternatively, the present disclosure may provide a method of always applying only a non-separable transform as a primary transform method. Alternatively, the present disclosure may provide a method of selectively applying one of a separable transform and a non-separable transform as a primary transform method. Alternatively, the present disclosure may provide a method of applying both a separable transform and a non-separable transform as a primary transform method. The separable transform-based primary transform method may include DCT type 2, DST type 7, DCT type 8, DCT type 5, DST type 4, DST type 1, IDT (identity transform), or any other transform method in which transforms are performed independently in the horizontal and vertical directions (e.g., a transform skip).

[0114] The present disclosure may include a non-separable primary transform (NSPT) as one of the primary transformation methods, and may provide a method for effectively signaling information regarding the activation / deactivation of the non-separable primary transform.

[0115] According to an embodiment of the present disclosure, one flag may be used to control non-separable linear transform activation for luma blocks and / or chroma blocks. For example, as shown in Table 4 below, one sps_nspt_enabled_flag may be signaled in the SPS syntax for non-separable linear transform activation.

[0116] [Table 4]

[0117] Here, when the value of sps_nspt_enabled_flag is 1 (i.e., the second value), the video encoding device 100 and / or the video decoding device 200 can perform one of separable transform or non-separable transform as a primary transform method. When the value of sps_nspt_enabled_flag is 0 (i.e., the first value), the video encoding device 100 and / or the video decoding device 200 can perform only separable transform as a primary transform method. In this case, the flag for activating / deactivating the non-separable primary transform (i.e., sps_nspt_enabled_flag) may be signaled by APS, PPS, VPS, DPS (decoding parameter set), picture header syntax, slice header syntax, etc. in addition to the SPS syntax.

[0118] According to another embodiment of the present disclosure, one flag can be used to control non-separable linear transform activation for luma blocks and / or chroma blocks. For non-separable linear transform activation, the present disclosure can signal one sps_nspt_enabled_flag in the SPS syntax as shown in Table 4 above.

[0119] In this case, when the value of sps_nspt_enabled_flag is 1 (i.e., the second value), the video encoding device 100 and / or the video decoding device 200 can selectively perform either the separable transform or the non-separable transform as the primary transform method. Alternatively, when the value of sps_nspt_enabled_flag is 1 (i.e., the second value), the video encoding device 100 and / or the video decoding device 200 can perform both the separable transform and the non-separable transform as the primary transform method.

[0120] In this case, whether to selectively perform one of the separable transform or the non-separable transform, or to perform both the separable transform and the non-separable transform as the primary transform method, may be explicitly determined by another flag. A flag indicating whether to selectively perform one of the separable transform or the non-separable transform, or to perform both the separable transform and the non-separable transform as the primary transform method, may be signaled in various high-level syntaxes such as SPS, APS, PPS, VPS, DPS, picture header syntax, slice header syntax, etc. Alternatively, a flag indicating whether to selectively perform one of the separable transform or the non-separable transform, or to perform both the separable transform and the non-separable transform as the primary transform method may be signaled in a low-level syntax such as CTU, CU, or TU. A flag indicating whether to selectively perform one of the separable transform or the non-separable transform, or to perform both the separable transform and the non-separable transform as the primary transform method may be signaled only when the value of sps_nspt_enabled_flag is 1 (i.e., the second value).

[0121] When the value of sps_nspt_enabled_flag is 0 (i.e., the first value), the video encoding device 100 and / or the video decoding device 200 can perform only the separation transform as the primary transform method. In this case, the flag for activating / deactivating the non-separation primary transform (i.e., sps_nspt_enabled_flag) may be signaled by the APS, PPS, VPS, DPS (decoding parameter set), picture header syntax, slice header syntax, etc. in addition to the SPS syntax.

[0122] Example 2 According to the present disclosure, the primary transform method from the spatial domain to the frequency domain may include a method of always applying only a separable transform, a method of always applying only a non-separable transform, a method of selectively applying either a separable transform or a non-separable transform, or a method of applying both a separable transform and a non-separable transform. Here, the separable transform-based primary transform method may include DCT type 2, DST type 7, DCT type 8, DCT type 5, DST type 4, DST type 1, IDT, or any other transform method that performs transforms independently in the horizontal and vertical directions (i.e., transform skip).

[0123] In various situations involving a non-separable linear transform (NSPT) as one of the linear transform methods, the present disclosure may provide a method for effectively signaling syntax elements related to a non-separable linear transform. For example, instead of applying a non-separable transform to all transform blocks, the video encoding device 100 and / or the video decoding device 200 may apply a non-separable transform only to specific blocks, taking into consideration coding efficiency and required computational complexity. Here, the specific blocks may be square blocks, non-square blocks, blocks with a given number of pixels in the transform block or less, or blocks with a given width and / or height. That is, when the conditions for the specific blocks are met, non-separable transform-related syntax elements (i.e., a flag indicating whether non-separable transform activation is possible, a flag indicating whether a non-separable transform should be applied, or a non-separable transform index) may be signaled.

[0124] In addition, the video encoding device 100 and / or the video decoding device 200 can determine whether to apply a non-separable transform based on at least one of a square block, a non-square block, a block in which the number of pixels in the transform block is less than or equal to a given value, or a block in which the width and / or height of the transform block is less than or equal to a given value.

[0125] According to an embodiment of the present disclosure, a non-separable transform may be applied to some or all of the square blocks (i.e., 4x4, 8x8, 16x16, 32x32, etc.). Therefore, a separable transform and / or a non-separable transform may be considered as a primary transform method for square blocks, and only a separable transform may be considered as a primary transform method for non-square blocks. In this way, the present disclosure may improve compression efficiency due to a non-separable transform for square blocks, and may prevent a decrease in compression performance and an increase in complexity due to a non-separable transform due to the transmission and / or storage of unnecessary syntax elements and side information for non-square blocks.

[0126] According to another embodiment of the present disclosure, a non-separable transform may be applied to some or all of the non-square blocks (i.e., 4x8, 8x4, 8x16, 16x8, 16x32, 32x16, etc.). Therefore, a separable transform and / or a non-separable transform may be considered as a primary transform method for non-square blocks, and only a separable transform may be considered as a primary transform method for square blocks. In this way, the present disclosure may improve compression efficiency due to a non-separable transform for non-square blocks, and may prevent a decrease in compression performance and an increase in complexity due to a non-separable transform for square blocks due to the transmission and / or storage of unnecessary syntax elements and side information.

[0127] According to another embodiment of the present disclosure, a non-separable transform may be applied only to blocks whose transform block contains a given number of pixels or less. For example, a non-separable transform may be applied only to blocks whose transform block contains 64 or less pixels. In this case, separable transform and / or non-separable transform may be considered as a primary transform method only for blocks such as 4x4, 4x8, 8x4, 8x8, 4x16, and 16x4. For blocks whose transform block contains more than 64 pixels, only separable transform may be considered as a primary transform method. In this way, the present disclosure can improve compression efficiency using non-separable transform for blocks whose transform block contains a given number of pixels or less, and can prevent degradation of compression performance and increase in complexity due to transmission and / or storage of unnecessary syntax elements and side information for blocks whose transform block contains more than a given number of pixels.

[0128] According to another embodiment of the present disclosure, a non-separable transform may be applied only to blocks whose transform block width and / or height are equal to or less than a predetermined value. For example, when a non-separable transform is applied only to blocks whose transform block width is equal to or less than 8, a separable transform and / or a non-separable transform may be considered as a primary transform method only for blocks such as 4x4, 4x8, 4x16, 8x4, 8x8, and 8x16. For blocks whose transform block width is greater than 8, only a separable transform may be considered as a primary transform method. In this way, the present disclosure can improve compression efficiency using a non-separable transform for blocks whose transform block width and / or height are equal to or less than a predetermined value, and can prevent degradation of compression performance and increase in complexity due to transmission and / or storage of unnecessary syntax elements and side information for blocks whose transform block width and / or height exceed a predetermined value.

[0129] Example 3 According to the present disclosure, the primary transform method from the spatial domain to the frequency domain may include a method of always applying only separable transforms, a method of always applying only non-separable transforms, a method of selectively applying either a separable transform or a non-separable transform, or a method of applying both a separable transform and a non-separable transform. Here, the separable transform-based primary transform method may include DCT type 2, DST type 7, DCT type 8, DCT type 5, DST type 4, DST type 1, IDT, or any other transform in which transforms are performed independently in the horizontal and vertical directions (i.e., a transform skip) scheme. In various situations where a non-separable primary transform (nspt) is included as one of the primary transform methods, the present disclosure may provide a method of effectively signaling syntax elements related to the non-separable primary transform.

[0130] The non-separable transform set and / or kernel may be determined taking into account the statistical characteristics (statistical distribution of residual data) of the residual signal based on information such as whether the prediction mode of the current block is intra prediction mode or inter prediction mode, the intra prediction mode of the current block, size information of the current block (i.e., block width and / or height, block shape, number of pixels in the block, etc.), explicitly signaled syntax elements, statistical characteristics of surrounding pixels (all or some of the pixels included in blocks adjacent to the upper and / or left side of the current block), whether a quadratic transform is used, or QP (quantization parameter).

[0131] According to an embodiment of the present disclosure, non-separable transform sets and / or kernels may be determined based on whether the prediction mode of a current block is an intra prediction mode or an inter prediction mode. In the case of an inter prediction mode, the amount of residual data is not large, so using a large number of non-separable transform sets and / or kernels may not provide significant coding efficiency. Therefore, to reduce memory consumption for storing non-separable linear transform kernels and reduce overhead due to signaling of non-separable linear transform syntax elements / side information, the present disclosure may have a smaller number of transform sets and / or kernels for inter prediction mode blocks than for intra prediction mode blocks.

[0132] According to another embodiment of the present disclosure, a current block may have a different number of non-separable transform sets and / or kernels depending on the intra prediction mode of the current block. In this case, a block to which a planar mode, a DC mode, or a wide-angle intra prediction (WAIP) mode is applied may have a smaller number of non-separable transform sets and / or kernels because residual characteristics are less diverse than blocks to which other intra prediction modes are applied. As a result, the video encoding device 100 and / or the video decoding device 200 may be able to achieve advantages such as improved coding efficiency, reduced memory consumption for storing non-separable linear transform kernels, and reduced overhead due to non-separable linear transform syntax elements / side information signaling.

[0133] According to other embodiments of the present disclosure, the video encoding device 100 and / or the video decoding device 200 may have a non-separable transform set and / or kernel based on at least one of block size information (i.e., block width and / or height, block shape, number of pixels in the block), explicitly signaled syntax elements, statistical characteristics of surrounding pixels (all or some of the pixels included in blocks adjacent to the top and / or left of the current block), whether to use a quadratic transform, or QP.

[0134] The residual characteristics of each block may vary depending on at least one of block size information (i.e., block width and / or height, block shape, number of pixels in a block), explicitly signaled syntax elements, statistical characteristics of neighboring pixels (all or some of the pixels included in the blocks adjacent to the top and / or left of the current block), whether or not a quadratic transform is used, and QP. As a result, the video encoding device 100 and / or the video decoding device 200 may be able to improve coding efficiency, reduce memory consumption for storing non-separable linear transform kernels, and reduce overhead due to signaling of non-separable linear transform syntax elements / side information.

[0135] When having different numbers of non-separable transform kernels according to embodiments of the present disclosure, the video encoding device 100 and / or the video decoding device 200 may signal / parse non-separable transform-related syntax elements in various ways.

[0136] According to an embodiment of the present disclosure, when there is one kernel for a non-separable transform and multiple separable transform kernels, the non-separable transform may be set to transform index 0 (i.e., first value), and the transform indexes for the separable transform may be set starting from 1 (i.e., second value). That is, when there is one kernel for a non-separable transform and multiple separable transform kernels, transform index 0 may indicate the non-separable transform kernel, and transform indexes 1 to N+1 may indicate N separable transform kernels, where N may be a natural number. The term "transform index" referred to below may be used as a concept including various indexes, such as "non-separable transform index," "separable transform index," "non-separable primary transform index," "index indicating a non-separable transform kernel," or "index indicating a separable transform kernel."

[0137] For example, if there are five separable transform possibilities, transform index 0 may indicate a non-separable transform, and transform indexes 1 to 5 may indicate pre-defined separable transforms. That is, if there are five separable transform possibilities, transform index 0 may indicate a non-separable transform kernel, and transform indexes 1 to 5 may indicate pre-defined separable transform kernels, respectively. Therefore, transform index values ​​0 to 5 are signaled from video encoding device 100 to video decoding device 200, and whether a non-separable transform or a separable transform is used can be determined based on the decoded transform index values. Furthermore, which separable transform has been used can be determined based on the decoded transform index values. It goes without saying that which separable transform kernel or non-separable transform kernel has been used can be determined based on the decoded transform index values.

[0138] According to another embodiment of the present disclosure, when there is one kernel of a non-separable transform and there are multiple separable transform kernels, the transform index of the separable transform may be set starting from 0 (i.e., the first value), and the transform index of the non-separable transform may be set to the maximum transform index (last transform index). That is, when there is one kernel of a non-separable transform and the number of separable transform kernels is N, the transform indexes 0 to N-1 may each indicate a predefined separable transform kernel, and the transform index N may indicate a non-separable transform kernel, where N may be a natural number.

[0139] For example, if there are five possible separable transforms, transform indices 0 to 4 may indicate predefined separable transforms, and transform index 5 may indicate a non-separable transform. That is, transform indices 0 to 4 may indicate predefined separable transform kernels, and transform index 5 may indicate a non-separable transform kernel. Thus, transform index values ​​0 to 5 are signaled from video encoding device 100 to video decoding device 200, and it is possible to determine whether a non-separable transform or a separable transform is to be used based on the decoded transform index values. Furthermore, it is possible to determine which separable transform is to be used based on the decoded transform index values. It goes without saying that it is possible to determine whether a separable transform kernel or a non-separable transform kernel is to be used based on the decoded transform index values.

[0140] According to another embodiment of the present disclosure, when there is one kernel for a non-separable transform and there are multiple separable transform kernels, an index indicating the non-separable transform may be set to a predefined value N, and the indices indicating the separable transform may be defined sequentially, starting from 0 to a predefined maximum index value (i.e., M), excluding N (where M>=N). That is, when there is one kernel for a non-separable transform and there are multiple separable transform kernels, an index indicating the non-separable transform kernel may be set to a predefined value N, and the indices indicating the separable transform kernels may be defined sequentially, starting from 0 to a predefined maximum index value (i.e., M), excluding N. Here, M and N may be natural numbers.

[0141] For example, if there are five available separable transforms and the index of the non-separable transform is defined as 3, the separable transform indices are assigned from 0 to 5 in a predetermined order, but the transform index 3 may be skipped. That is, the transform indices 0 to 2 and 4 to 5 can indicate separable transform kernels. In this case, the transform index values ​​0 to 5 are signaled from the video encoding device 100 to the video decoding device 200, and it is possible to determine whether a non-separable transform or a separable transform is to be used based on the decoded transform index values. Furthermore, it is possible to determine which separable transform kernel or non-separable transform kernel is to be used based on the decoded transform index values. It goes without saying that it is possible to determine which separable transform kernel or non-separable transform kernel is to be used based on the decoded transform index values.

[0142] According to another embodiment of the present disclosure, when there is one kernel of the non-separable transform and there are multiple separable transform kernels, whether to perform the non-separable transform may be first signaled using a flag indicating whether to perform the non-separable transform. That is, the video encoding device 100 and / or the video decoding device 200 may signal a flag indicating whether to perform the non-separable transform.

[0143] According to this embodiment, when a flag indicating whether a non-separable transform is performed is 0 (i.e., a first value), an index of the separable transform may be further signaled. The value of the flag indicating whether a non-separable transform is performed may be signaled by predicting a probability using a context coded bin. In this case, a context model for probability prediction may be configured using a block size, a block shape, an intra prediction mode, information of a previously coded block, etc.

[0144] According to the present disclosure, transform index values ​​may be binarized using fixed length code (FLC) or truncated binary code (TBC). In this case, the transform index values ​​may be encoded and / or decoded using context coding or bypass coding. Here, context coding may refer to coding performed by treating the transform index as a context coding bin. Bypass coding may refer to coding performed by treating the transform index as a bypass coding bin. Alternatively, the transform index may be represented using truncated unary (TU) binarization. The transform index represented using TU binarization may be encoded and / or decoded by predicting its occurrence probability using context coding, or may be encoded and / or decoded with the same probability using bypass coding.

[0145] According to another embodiment of the present disclosure, when there are two or more kernels of a non-separable transform and the number of separable transform kernels is plural, a flag indicating whether or not to perform a non-separable transform may be first signaled. When the flag indicating whether or not to perform a non-separable transform is 1 (i.e., a second value), the video encoding device 100 and / or the video decoding device 200 may further signal an index indicating a non-separable transform kernel. When the flag indicating whether or not to perform a non-separable transform is 0 (i.e., a first value), the video encoding device 100 and / or the video decoding device 200 may further signal an index indicating a separable transform kernel. In this case, the value of the index indicating the non-separable transform kernel or the value of the index indicating the separable transform kernel may be binarized by FLC or TBC. In this case, the transform index may be coded and / or decoded by context coding or bypass coding.

[0146] Alternatively, the transform index may be represented by TU (truncated unary) binarization. The transform index represented by TU binarization may be coded and / or decoded by predicting its occurrence probability through context coding, or may be coded and / or decoded with the same probability through bypass coding. In addition, a flag indicating whether to perform a non-separable transform may be coded and / or decoded by predicting its probability through context coding bins. In this case, a context model for probability prediction may use a block size, a block shape, an intra prediction mode, information on a previously coded block, etc.

[0147] According to another embodiment of the present disclosure, when there are two or more non-separable transform kernels and the number of separable transform kernels is plural, it is possible to signal the transform index at once without separately signaling a flag indicating whether to perform the non-separable transform. For example, when there are N non-separable transform kernels and M available separable transform kernels, a transform index value from 0 to N-1 indicates N non-separable transform kernels (0 to N-1), and a transform index value from N to N+M-1 indicates separable transform indexes 0 to (M-1).

[0148] That is, when there are N non-separable transform kernels and M usable separable transform kernels, if the value of a transform index is 0 to N-1, the transform index may be an index indicating the N non-separable transform kernels. Alternatively, when there are N non-separable transform kernels and M usable separable transform kernels, if the value of a transform index is N to N+M-1, the transform index may be an index indicating the M separable transform kernels. Here, N and M may be natural numbers. In addition to the above-mentioned cases, the method of mapping usable separable transform indexes and non-separable transform kernels from a transform index may be applied in various predefined forms.

[0149] According to another embodiment of the present disclosure, when there are two or more non-separable transform kernels and the number of separable transform kernels is plural, the video encoding device 100 and / or the video decoding device 200 may signal the transform index at once without separately signaling a flag indicating whether to perform the non-separable transform. For example, when there are N non-separable transform kernels and M available separable transform kernels, a transform index value ranging from 0 to M-1 indicates M separable transform kernels (0 to M-1), and a transform index value ranging from M to N+M-1 indicates non-separable transform indexes 0 to (N-1).

[0150] That is, when there are N non-separable transform kernels and M usable separable transform kernels, if the value of a transform index is 0 to M-1, the transform index may be an index indicating the M separable transform kernels. Alternatively, when there are N non-separable transform kernels and M usable separable transform kernels, if the value of a transform index is M to N+M-1, the transform index may be an index indicating the N non-separable transform kernels. Here, N and M may be natural numbers. In addition to the above-mentioned cases, the method of mapping usable separable transform indexes and non-separable transform kernels from a transform index may be applied in various predefined forms.

[0151] In the various embodiments described above, when defining a separate transform index, the video encoding device 100 and / or the video decoding device 200 may determine the transforms to be applied in the horizontal and vertical directions according to a predetermined rule using the value of a given separate transform index. Alternatively, the separate transform index may be signaled after being divided into a horizontal separate transform index and a vertical separate transform index, and the transform designated by each index may be applied to the corresponding direction.

[0152] The non-separable transform activation flag referred to in this disclosure may be signaled in high-level syntax (i.e., SPS, APS, PPS, VPS, DPS, picture header, slice header, etc.), or the flag indicating whether to perform a non-separable transform or the non-separable transform activation flag may not be signaled separately by the convention that a non-separable transform is always considered as a candidate for the primary transform.

[0153] The flag indicating whether or not to perform the non-separable transform may be a flag indicating whether or not to perform the non-separable transform as a primary transform. Also, the non-separable transform activation flag may be a flag indicating whether or not the non-separable transform is activated as a primary transform.

[0154] Example 4 According to the present disclosure, the primary transform method from the spatial domain to the frequency domain may include a method of always applying only a separable transform, a method of always applying only a non-separable transform, a method of selectively applying either a separable transform or a non-separable transform, or a method of applying both a separable transform and a non-separable transform. Here, the separable transform-based primary transform method may include DCT type 2, DST type 7, DCT type 8, DCT type 5, DST type 4, DST type 1, IDT, or any other transform method that performs transforms independently in the horizontal and vertical directions (i.e., transform skip).

[0155] Since non-separable transforms have higher computational complexity and / or memory requirements than separable transforms, various methods are needed to reduce computational complexity and power consumption. Therefore, the present disclosure proposes a method for signaling non-separable transform syntax elements for an efficient transform structure.

[0156] Since non-separable transforms require greater computational complexity or memory requirements than separable transforms, a method is needed to reduce complexity and power consumption by adjusting worst-case computational complexity. Therefore, when a non-separable linear transform is used, the use of a non-separable quadratic transform may be restricted. The video encoding device 100 and / or the video decoding device 200 may specify that syntax elements for a non-separable quadratic transform technique are coded and / or decoded after syntax elements for a non-separable linear transform technique. That is, the video encoding device 100 may not consider a non-separable quadratic transform when a non-separable linear transform is used. After parsing a syntax element for a non-separable linear transform technique, the video decoding device 200 may not parse a syntax element for a non-separable quadratic transform technique if the information indicates that a non-separable linear transform is used. That is, if a syntax element for a non-separable linear transform technique indicates that a non-separable linear transform is used, the video decoding device 200 does not need to parse a syntax element for a non-separable secondary transform technique. In this case, the video decoding device 200 can restrict the use of secondary transforms.

[0157] According to another embodiment of the present disclosure, the video encoding device 100 and / or the video decoding device 200 may specify that syntax elements for a non-separable primary transform technique are encoded and / or decoded after syntax elements for a non-separable secondary transform technique. The video encoding device 100 may not consider a non-separable primary transform when a non-separable secondary transform is used. After parsing a syntax element for a non-separable secondary transform technique, the video decoding device 200 may not parse a syntax element for a non-separable primary transform technique if the syntax element indicates the use of a non-separable secondary transform. In other words, if a syntax element for a non-separable secondary transform technique indicates the use of a non-separable secondary transform, the video decoding device 200 may not parse the syntax element for a non-separable primary transform technique. In this case, the video decoding device 200 may perform a separable transform in the primary transform.

[0158] For blocks with a small number of residual coefficients or blocks without relatively high-frequency components in the current transform block, it may be effective to use one non-separable transform kernel rather than multiple non-separable transform kernels. In this case, by reducing bit consumption for signaling non-separable transform syntax elements and / or kernels, the video encoding device 100 and / or the video decoding device 200 may achieve more efficient compression performance. Conversely, for blocks with a large number of residual coefficients or blocks with a large number of high-frequency components in the current transform block, it may be effective to use multiple non-separable transform kernels, even if it requires additional signaling, because these blocks may have diverse characteristics. Alternatively, it may be effective to use a separable transform instead of a non-separable linear transform based on the number of residual coefficients in the current transform block or the characteristics of the current transform block.

[0159] To utilize the characteristics of the residual coefficients of the current transform block, the video encoding device 100 and / or the video decoding device 200 may specify that syntax elements for a non-separable first-order transform technique be coded after residual coding (or transform coefficient coding) syntax elements. For example, the video encoding device 100 and / or the video decoding device 200 may code syntax elements for a first-order transform technique after residual coding syntax elements, but may or may not signal a syntax element indicating the use of a non-separable transform technique based on the number of residual coefficients in the transform block or the position of the last residual coefficient.

[0160] Specifically, if the number of residual coefficients or the position of the last residual coefficient is smaller than (or smaller than or equal to) a specific threshold, the video encoding device 100 and / or the video decoding device 200 may determine that the number of high-frequency components is relatively small and may not signal or may signal syntax elements associated with the non-separable transform technique only in this case. Alternatively, if the number of residual coefficients or the position of the last residual coefficient is larger than (or larger than or equal to) a specific threshold, the video encoding device 100 and / or the video decoding device 200 may determine that the number of high-frequency components is relatively large and may signal syntax elements associated with the non-separable transform technique only in this case.

[0161] 4 is a flowchart of a video encoding method according to the present disclosure. According to the present disclosure, the video encoding apparatus 100 may obtain a residual block for a current block (S410). The video encoding apparatus 100 may also encode transform information including information about a non-separable linear transform (S420). The video encoding apparatus 100 may generate transform coefficients by transforming the residual block based on the transform information (S430).

[0162] Here, the transform information may include a non-separable primary transform activation flag indicating whether a non-separable primary transform is available for the current block. The non-separable primary transform activation flag may be sps_nspt_enabled_flag. If the value of the non-separable primary transform activation flag is a first value (i.e., 0), a separable primary transform may be available for the current block. If the value of the non-separable primary transform activation flag is a second value (i.e., 1), either a non-separable primary transform or a separable primary transform may be available for the current block.

[0163] According to an embodiment of the present disclosure, the transform information may include a non-separable linear transform application flag indicating whether a non-separable linear transform is applied to the current block. The non-separable linear transform application flag may be coded based on at least one of the shape of the current block or the size of the current block. For example, the non-separable linear transform application flag may be coded based on at least one of a square block, a non-square block, a block in which the number of pixels in the transform block is equal to or less than a given value, or a block in which the width and / or height of the transform block is equal to or less than a given value.

[0164] According to an embodiment of the present disclosure, the transformation information may include a kernel index indicating one of a plurality of transformation kernels. The plurality of transformation kernels may include at least one non-separable transformation kernel and at least one separable transformation kernel. The "kernel index" according to the present disclosure may also be referred to as a "transform index" or other terms having an equivalent technical meaning.

[0165] According to an embodiment of the present disclosure, when the non-separable linear transform application flag indicates that a non-separable linear transform is applied to the current block, the video decoding apparatus 200 may encode a non-separable transform kernel index indicating one of a plurality of non-separable transform kernels. When the non-separable linear transform application flag indicates that a non-separable linear transform is not applied to the current block, a separable transform kernel index indicating one of a plurality of separable transform kernels may be encoded. Here, the number of non-separable transform kernels or separable transform kernels may be determined based on at least one of whether intra prediction is applied to the current block, an intra prediction mode, a block size, neighboring sample values, whether a secondary transform is used, or a quantization parameter.

[0166] According to an embodiment of the present disclosure, when a non-separable linear transform is applied to a current block, non-separable secondary transforms may be restricted. In this case, information about the non-separable linear transform may be transmitted before information about the non-separable secondary transform for the current block, and transmission of information about the non-separable secondary transform for the current block may be skipped.

[0167] According to an embodiment of the present disclosure, when a non-separable secondary transform is applied to a current block, the non-separable linear transform may be restricted. In this case, information about the non-separable secondary transform may be transmitted before information about the non-separable linear transform for the current block, and transmission of information about the non-separable linear transform for the current block may be skipped.

[0168] According to one embodiment of the present disclosure, information about the non-separable linear transform may be coded based on the characteristics of the residual coefficients. For example, a block with a small number of residual coefficients may use a single non-separable transform kernel instead of multiple non-separable transform kernels. As yet another example, a block with a large number of residual coefficients may use multiple non-separable transform kernels.

[0169] 5 is a flowchart of a video decoding method according to the present disclosure. According to the present disclosure, the video decoding apparatus 200 may obtain transform information including information about a non-separable linear transform (S510). Then, the video decoding apparatus 200 may obtain transform coefficients for a current block (S520). The video decoding apparatus 200 may generate a residual block by inverse transforming the transform coefficients based on the obtained transform information (S530).

[0170] Here, the transformation information may include a non-separable primary transform activation flag indicating whether a non-separable primary transform is possible for the current block. The non-separable primary transform activation flag may be sps_nspt_enabled_flag.

[0171] If the value of the non-separable primary transform activation flag is a first value (i.e., 0), a separable primary transform may be available for the current block. If the value of the non-separable primary transform activation flag is a second value (i.e., 1), either a non-separable primary transform or a separable primary transform may be available for the current block.

[0172] According to an embodiment of the present disclosure, the transformation information may include a non-separable linear transform application flag indicating whether a non-separable linear transform is applied to the current block. The non-separable linear transform application flag may be obtained based on at least one of the shape of the current block or the size of the current block. For example, the non-separable linear transform application flag may be obtained based on at least one of a square block, a non-square block, a block in which the number of pixels in the transformation block is equal to or less than a predetermined value, or a block in which the width and / or height of the transformation block is equal to or less than a predetermined value.

[0173] According to an embodiment of the present disclosure, the transformation information may include a kernel index indicating one of a plurality of transformation kernels. The plurality of transformation kernels may include at least one non-separable transformation kernel and at least one separable transformation kernel. Here, according to the present disclosure, the “kernel index” may be referred to as a “transform index” or another term having an equivalent technical meaning.

[0174] According to an embodiment of the present disclosure, when a non-separable linear transform application flag indicates that a non-separable linear transform is applied to a current block, the video decoding apparatus 200 may obtain a non-separable transform kernel index indicating one of a plurality of non-separable transform kernels. When the non-separable linear transform application flag indicates that a non-separable linear transform is not applied to the current block, a separable transform kernel index indicating one of a plurality of separable transform kernels may be obtained. Here, the number of non-separable transform kernels or separable transform kernels may be obtained based on at least one of whether intra prediction is applied to the current block, an intra prediction mode, a block size, neighboring sample values, whether a secondary transform is used, or a quantization parameter.

[0175] According to an embodiment of the present disclosure, when a non-separable linear transform is applied to a current block, non-separable secondary transforms may be restricted. In this case, information about the non-separable linear transform may be parsed before information about the non-separable secondary transform for the current block, and parsing of information about the non-separable secondary transform for the current block may be skipped.

[0176] According to an embodiment of the present disclosure, when a non-separable secondary transform is applied to a current block, the non-separable linear transform may be restricted. In this case, information about the non-separable secondary transform may be parsed before information about the non-separable linear transform for the current block, and parsing of information about the non-separable linear transform for the current block may be skipped.

[0177] According to one embodiment of the present disclosure, information about the non-separable linear transform may be coded based on the characteristics of the residual coefficients. For example, a block with a small number of residual coefficients may use a single non-separable transform kernel instead of multiple non-separable transform kernels. As yet another example, a block with a large number of residual coefficients may use multiple non-separable transform kernels.

[0178] Although the exemplary method of the present disclosure is expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order if necessary. To embody the method of the present disclosure, the exemplary steps may further include other steps, or some steps may be omitted and the remaining steps may be included, or some steps may be omitted and the remaining steps may be included.

[0179] In the present disclosure, the video encoding device 100 or the video decoding device 200 performing a predetermined operation (step) may perform the operation (step) to check the execution conditions or circumstances of the operation (step). For example, if it is described that the predetermined operation is performed when a predetermined condition is satisfied, the video encoding device 100 or the video decoding device 200 may perform the predetermined operation after performing an operation to check whether the predetermined condition is satisfied.

[0180] The various embodiments of the present disclosure do not enumerate all possible combinations but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.

[0181] Furthermore, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof. In the case of a hardware implementation, the implementation may be using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0182] Furthermore, the video decoding device 200 and the video encoding device 100 to which the embodiments of the present disclosure are applied may be included in a multimedia broadcasting transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video on demand (VoD) service providing device, an over-the-top (OTT) video device, an internet streaming service providing device, a three-dimensional (3D) video device, an image telephone video device, a medical video device, etc., and may be used to process a video signal or a data signal. For example, over-the-top (OTT) video devices may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0183] FIG. 6 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.

[0184] As shown in FIG. 6, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0185] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0186] The bitstream may be generated by a video encoding method and / or video encoding device 100 to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0187] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server may act as an intermediary informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server may control commands and responses between devices in the content streaming system.

[0188] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.

[0189] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, and head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signage.

[0190] Each server in the content streaming system may be operated as a distributed server, in which case data received by each server may be processed in a distributed manner.

[0191] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable media on which such software or instructions, etc., may be stored and executed on a device or computer. [Industrial Applicability]

[0192] The embodiments of the present disclosure can be used to encode / decode video.

[0193] [Claims at the time of international application] [Claim 1] A video decoding method performed by a video decoding device, comprising: obtaining transformation information including information about the non-separable linear transformation; obtaining transform coefficients for the current block; generating a residual block by inverse transforming the transform coefficients based on the transform information. [Claim 2] the transformation information includes a non-separable primary transformation activation flag indicating whether a non-separable primary transformation is possible for the current block; indicating that a separate primary transform is available for the current block based on the value of the non-separable primary transform activation flag being a first value; 2. The video decoding method of claim 1, further comprising: indicating that either a non-separable linear transform or a separable linear transform is available for the current block based on the value of the non-separable linear transform activation flag being a second value. [Claim 3] 2. The video decoding method of claim 1, wherein the transformation information includes a non-separable linear transform application flag indicating whether a non-separable linear transform is applied to the current block based on at least one of a type of the current block or a size of the current block. [Claim 4] the transformation information includes a kernel index indicating one of a plurality of transformation kernels; The video decoding method of claim 1 , wherein the plurality of transform kernels includes at least one non-separable transform kernel and at least one separable transform kernel. [Claim 5] The transformation information includes a non-separable linear transformation application flag indicating whether a non-separable linear transformation is applied to the current block; a non-separable transform kernel index indicating one of a plurality of non-separable transform kernels is obtained based on the non-separable linear transform application flag indicating that a non-separable linear transform is applied to the current block; 2. The video decoding method of claim 1, wherein a separate transform kernel index indicating one of a plurality of separate transform kernels is obtained based on the non-separable linear transform application flag indicating that a non-separable linear transform is not applied to the current block. [Claim 6] 5. The video decoding method of claim 4, wherein the number of the non-separable transform kernels or the number of the separate transform kernels is obtained based on at least one of whether intra prediction is applied to the current block, an intra prediction mode, a block size, neighboring sample values, whether a secondary transform is used, or a quantization parameter. [Claim 7] The video decoding method of claim 1 , wherein a secondary transform for the current block is restricted to be a non-separable transform based on a non-separable linear transform applied to the current block. [Claim 8] information about the non-separable linear transform is parsed before information about the non-separable secondary transform for the current block; The video decoding method of claim 7 , wherein parsing of information related to the non-separable secondary transform is skipped based on a non-separable linear transform being applied to the current block. [Claim 9] The video decoding method of claim 1 , wherein a linear transform for the current block is restricted to be a non-separable transform based on the application of a non-separable secondary transform to the current block. [Claim 10] information about a non-separable secondary transform for the current block is parsed before information about the non-separable linear transform; 10. The video decoding method of claim 9, wherein parsing of information related to the non-separable quadratic transform is skipped based on a non-separable quadratic transform being applied to the current block. [Claim 11] The video decoding method of claim 1 , wherein the information about the non-separable linear transform is coded based on a characteristic of a residual coefficient. [Claim 12] A video encoding method performed by a video encoding device, comprising: obtaining a residual block for the current block; encoding transformation information including information about the non-separable linear transformation; generating transform coefficients by transforming the residual block based on the transform information. [Claim 13] A computer-readable recording medium storing a bitstream generated by the video encoding method of claim 12. [Claim 14] A method for transmitting a bitstream generated by a video encoding method, comprising: obtaining a residual block for the current block; encoding transformation information including information about the non-separable linear transformation; generating transform coefficients by transforming the residual block based on the transform information.

Claims

1. A video decoding method performed by a video decoding device, comprising: obtaining transformation information including information about the non-separable linear transformation; obtaining transform coefficients for the current block; generating a residual block by inverse transforming the transform coefficients based on the transform information.

2. the transformation information includes a non-separable primary transformation activation flag indicating whether a non-separable primary transformation is possible for the current block; indicating that a separate linear transform is available for the current block based on the value of the non-separable linear transform activation flag being a first value; 2. The video decoding method of claim 1, wherein the non-separable linear transform activation flag is set to a second value, indicating that either a non-separable linear transform or a separable linear transform is available for the current block.

3. 2. The video decoding method of claim 1, wherein the transformation information includes a non-separable linear transform application flag indicating whether a non-separable linear transform is applied to the current block based on at least one of a type of the current block or a size of the current block.

4. the transformation information includes a kernel index indicating one of a plurality of transformation kernels; The video decoding method of claim 1 , wherein the plurality of transform kernels comprises at least one non-separable transform kernel and at least one separable transform kernel.

5. The transformation information includes a non-separable linear transformation application flag indicating whether a non-separable linear transformation is applied to the current block; a non-separable transform kernel index indicating one of a plurality of non-separable transform kernels is obtained based on the non-separable linear transform application flag indicating that a non-separable linear transform is applied to the current block; 2. The video decoding method of claim 1, wherein a separate transform kernel index indicating one of a plurality of separate transform kernels is obtained based on the non-separable linear transform application flag indicating that a non-separable linear transform is not applied to the current block.

6. The video decoding method of claim 4, wherein the number of the non-separable transform kernels or the number of the separate transform kernels is obtained based on at least one of whether intra prediction is applied to the current block, an intra prediction mode, a block size, neighboring sample values, whether a secondary transform is used, or a quantization parameter.

7. The video decoding method of claim 1 , wherein a secondary transform for the current block is restricted to be a non-separable transform based on a non-separable linear transform applied to the current block.

8. information about the non-separable linear transform is parsed before information about the non-separable secondary transform for the current block; The video decoding method of claim 7 , wherein parsing of information related to the non-separable secondary transform is skipped based on the fact that a non-separable linear transform is applied to the current block.

9. The video decoding method of claim 1 , wherein a linear transform for the current block is restricted to be a non-separable transform based on the fact that a non-separable secondary transform is applied to the current block.

10. information about a non-separable secondary transform for the current block is parsed before information about the non-separable linear transform; The video decoding method of claim 9 , wherein parsing of information related to the non-separable quadratic transform is skipped based on the fact that a non-separable quadratic transform is applied to the current block.

11. The video decoding method of claim 1 , wherein the information about the non-separable linear transform is coded based on a characteristic of a residual coefficient.

12. A video encoding method performed by a video encoding device, comprising: obtaining a residual block for the current block; encoding transform information including information about the non-separable linear transform; generating transform coefficients by transforming the residual block based on the transform information.

13. A computer-readable recording medium storing a bitstream generated by the video encoding method of claim 12.

14. A method for transmitting a bitstream generated by a video encoding method, comprising: obtaining a residual block for the current block; encoding transform information including information about the non-separable linear transform; generating transform coefficients by transforming the residual block based on the transform information.