Image encoding / decoding method and device, and recording medium storing bit stream
By applying the inseparable main transform and dimensionality reduction transform kernel, the image encoding and decoding process is optimized, the problem of insufficient compression efficiency of high-resolution and high-quality images is solved, and the encoding efficiency and image quality are improved.
Patent Information
- Application Number
- CN202380092429.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-09
- Filing Date
- 2023-12-08
- Publication Date
- 2025-09-05
AI Technical Summary
Existing image compression technologies are inefficient in encoding high-resolution and high-quality images, especially when processing high-definition and ultra-high-definition images, which are difficult to compress and decode effectively.
The non-separable primary transform (NSPT) and the dimensionality-reduced non-separable primary transform kernel are adopted. The transform index is notified by signal, the transform kernel is determined, and the forward and backward non-separable primary transforms are combined for transformation and inverse transformation to optimize the image encoding and decoding process.
It improves the transformation performance and coding efficiency, improves the image compression effect, and is suitable for encoding and decoding of high-resolution and high-quality images.
Smart Images

Figure CN120604513A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and a recording medium storing a bit stream. Background Art
[0002] Recently, demands for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images have been increasing in various application fields, and thus, high-efficiency image compression technology is being discussed.
[0003] There are various technologies, such as inter-frame prediction technology that uses video compression technology to predict pixel values included in the current picture from pictures before or after the current picture, intra-frame prediction technology that predicts pixel values included in the current picture by using pixel information in the current picture, entropy coding technology that assigns short symbols to values with high frequency of occurrence and long symbols to values with low frequency of occurrence, etc. These image compression technologies can be used to effectively compress image data and transmit or store it. Summary of the Invention
[0004] Technical issues
[0005] The present disclosure provides a method and apparatus for performing transformation by using an inseparable main transform.
[0006] The present disclosure provides a method and apparatus for performing a transform by using a non-separable primary transform kernel with reduced dimensionality.
[0007] The present disclosure provides a method and apparatus for determining an inseparable primary transform kernel based on encoding parameters.
[0008] Technical Solution
[0009] According to the image decoding method and apparatus of the present disclosure, residual information can be obtained from a bitstream, transform coefficients of a current block are derived based on the residual information, at least one of dequantization or inverse transformation is performed on the transform coefficients of the current block to derive residual samples of the current block, and the current block is reconstructed based on the residual samples of the current block.
[0010] In the image decoding method and apparatus according to the present disclosure, an inverse transform may be performed based on a backward non-separable primary transform (NSPT), and a transform kernel of the NSPT may be determined based on a transform index indicating any one of a plurality of transform kernel candidates belonging to a transform set. Here, the transform index may be signaled based on at least one of a first condition for a luma component or a second condition for a chroma component.
[0011] In the image decoding method and device according to the present disclosure, the first condition for the luminance component may indicate that there are no non-zero transform coefficients in a predetermined area within the luminance component block of the current block, and the second condition for the chrominance component may indicate that there are no non-zero transform coefficients in a predetermined area within the chrominance component block of the current block.
[0012] In the image decoding method and device according to the present disclosure, the predetermined area within the luminance component block can be a remaining area excluding the first area from the luminance component block, and the first area can be an area consisting of the same number of samples as the number of transform coefficients input to the inseparable main transform.
[0013] In the image decoding method and apparatus according to the present disclosure, the predetermined area within the chroma component block may be a remaining area excluding the first area from the chroma component block, and the first area may be an area consisting of the same number of samples as the input length of the inseparable main transform corresponding to the size of the chroma component block.
[0014] In the image decoding method and apparatus according to the present disclosure, when the tree type of the current block is a single tree and the inseparable main transform is applied to the luminance component block of the current block but not to the chrominance component block of the current block, the transform index may be signaled based on a situation where a first condition for the luminance component and a second condition for the chrominance component are satisfied.
[0015] In the image decoding method and apparatus according to the present disclosure, when the tree type of the current block is a single tree and the inseparable main transform is applied to the luminance component block of the current block but not to the chrominance component block of the current block, the transform index can be signaled based on the situation where the first condition for the luminance component is satisfied, without checking whether the second condition for the chrominance component is satisfied.
[0016] In the image decoding method and apparatus according to the present disclosure, when the tree type of the current block is a single tree and an inseparable main transform is applied to the luminance component block and the chrominance component block of the current block, a transform index may be signaled based on a situation where a first condition for the luminance component and a second condition for the chrominance component are satisfied.
[0017] In the image decoding method and apparatus according to the present disclosure, a transform index may be signaled based on at least one of the following: whether there is a non-zero transform coefficient at a sample position other than the upper left sample position in at least one of the three color component blocks within the current block, or whether at least one of the three color component blocks is a transform skip block.
[0018] The image encoding method and apparatus according to the present disclosure may derive residual samples of a current block, perform at least one of transformation or quantization on the residual samples of the current block to derive transform coefficients of the current block, and encode the transform coefficients of the current block.
[0019] In the image encoding method and apparatus according to the present disclosure, transform may be performed based on a non-separable forward primary transform (NSPT), and a transform kernel of the NSPT may be determined as any one of a plurality of transform kernel candidates belonging to the transform kernel.
[0020] In the image decoding method and apparatus according to the present disclosure, a transform kernel indicating any one of a plurality of transform kernel candidates may be encoded based on at least one of a first condition for a luma component or a second condition for a chroma component.
[0021] A computer-readable digital storage medium storing encoded video / image information is provided, which enables a decoding device according to the present disclosure to perform an image decoding method.
[0022] A computer-readable digital storage medium storing video / image information generated according to the image encoding method according to the present disclosure is provided.
[0023] Provided are a method and apparatus for transmitting video / image information generated according to the image encoding method according to the present disclosure.
[0024] Beneficial effects
[0025] The present disclosure may improve transform performance by using a non-separable main transform as the main transform.
[0026] The present disclosure may improve transformation performance by performing transformation using a dimensionally reduced inseparable primary transform kernel.
[0027] The present disclosure may improve coding efficiency by efficiently determining or signaling a non-separable primary transform kernel based on coding parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A video / image encoding system according to the present disclosure is shown.
[0029] Figure 2 A schematic block diagram shows an encoding device to which embodiments of the present disclosure are applied and which performs encoding of a video / image signal.
[0030] Figure 3 A schematic block diagram shows a decoding device to which embodiments of the present disclosure are applied and which performs decoding of a video / image signal.
[0031] Figure 4 An image decoding method performed by a decoding device (300) according to an embodiment of the present disclosure is shown.
[0032] Figure 5 The intra prediction mode and its prediction direction according to the present disclosure are exemplarily shown.
[0033] Figure 6 A schematic configuration of a decoding device (300) for performing an image decoding method according to the present disclosure is shown.
[0034] Figure 7 An image encoding method performed by an encoding device (200) according to an embodiment of the present disclosure is shown.
[0035] Figure 8 A schematic configuration of an encoding device (200) for performing an image encoding method according to the present disclosure is shown.
[0036] Figure 9 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown. DETAILED DESCRIPTION
[0037] Since the present disclosure can be modified in various ways and has multiple embodiments, specific embodiments will be shown in the drawings and described in detail in the specific embodiments. However, it is not intended to limit the present disclosure to specific embodiments, but rather it should be understood to include all changes, equivalents, and alternatives included within the spirit and technical scope of the present disclosure. When describing the various figures, similar reference numerals are used for similar components.
[0038] Terms such as first, second, etc. can be used to describe various components, but components should not be limited by these terms. Terms are only used to distinguish one component from other components. For example, a first component can be referred to as a second component, and similarly, a second component can be referred to as a first component without departing from the scope of the present disclosure. Terms include any one or more combinations of related terms in a plurality of related terms.
[0039] When a component is referred to as being “connected” or “linked” to another component, it should be understood that it can be directly connected or linked to the other component, but another component may exist in between. On the other hand, when a component is referred to as being “directly connected” or “directly linked” to another component, it should be understood that another component does not exist in between.
[0040] The terms used in this application are only used to describe specific embodiments and are not intended to limit the present disclosure. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In this application, it should be understood that terms such as "including" or "having" are intended to specify the presence of features, quantities, steps, operations, components, parts, or combinations thereof described in this specification, but do not preclude the presence or addition of one or more other features, quantities, steps, operations, components, parts, or combinations thereof.
[0041] The present disclosure relates to video / image coding. For example, the methods / implementations disclosed herein may be applied to methods disclosed in the Versatile Video Coding (VVC) standard. In addition, the methods / implementations disclosed herein may be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0042] This specification proposes various embodiments of video / image encoding, and unless otherwise specified, the embodiments may be performed in combination with each other.
[0043] As used herein, video may refer to a collection of images over time. A picture generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms part of a picture during encoding. A slice / tile may include at least one Coding Tree Unit (CTU). A picture may be composed of at least one slice / tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and tile row of a picture. A tile column is a rectangular area of CTUs with a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs with a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile may be arranged contiguously according to a CTU raster scan, while tiles within a picture may be arranged contiguously according to a tile raster scan. A slice may include an integer number of complete tiles of a picture that may be exclusively included in a single NAL unit, or an integer number of consecutive complete CTU rows within a tile. In addition, a picture may be divided into at least two sub-pictures. A sub-picture may be a rectangular area of at least one slice within a picture.
[0044] A pixel or picture element (pel) may represent the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.
[0045] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the corresponding area. A unit may include a luma block and two chroma (e.g., CB, CR) blocks. In some cases, a unit may be used interchangeably with terms such as block or area. In general, an M×N block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0046] As used herein, "A or B" may mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" may be interpreted as "A and / or B." For example, as used herein, "A, B, or C" may mean "only A," "only B," "only C," or "any combination of A, B, and C."
[0047] As used herein, a slash ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B, or C."
[0048] Herein, “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” In addition, herein, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted in the same manner as “at least one of A and B.”
[0049] In addition, herein, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” In addition, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”
[0050] In addition, brackets used herein may indicate "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction." In other words, "prediction" herein is not limited to "intra-frame prediction," and "intra-frame prediction" may be provided as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction."
[0051] In this document, technical features described separately in one figure can be implemented separately or simultaneously.
[0052] Figure 1 A video / image encoding system according to the present disclosure is shown.
[0053] Reference Figure 1 , a video / image encoding system may include a first device (source device) and a second device (sink device).
[0054] The source device can transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0055] The video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a device for capturing video / images and a device for generating video / images. The device for capturing video / images may include at least one camera, a video / image archive including previously captured video / images, etc. The device for generating video / images may include a computer, a tablet computer, a smartphone, etc., and may generate the video / images (electronically). For example, a virtual video / image may be generated by a computer, etc., and in this case, the process of capturing video / images may be replaced by a process of generating relevant data.
[0056] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0057] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include components for generating a media file in a predetermined file format and may include components for transmitting via a broadcast / communication network. The receiving unit may receive / extract the bitstream and transmit it to the decoding device.
[0058] The decoding device may decode the video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0059] The renderer may render the decoded video / image, and the rendered video / image may be displayed by a display unit.
[0060] Figure 2 A rough block diagram showing an encoding device to which embodiments of the present disclosure can be applied and which performs encoding of a video / image signal is shown.
[0061] Reference Figure 2, the encoding device 200 may be composed of an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-mentioned image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.
[0062] The image splitter 210 may split the input image (or picture or frame) input to the encoding device 200 into at least one processing unit. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from the coding tree unit (CTU) or the largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure.
[0063] For example, one coding unit can be split into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure can be applied first, and then the binary tree structure and / or the ternary tree structure can be applied. Alternatively, the binary tree structure can be applied before the quadtree structure. The encoding process according to this specification can be performed based on the final coding unit that is no longer split. In this case, based on encoding efficiency and the like according to image characteristics, the maximum coding unit can be directly used as the final coding unit, or if necessary, the coding unit can be recursively split into coding units of a deeper depth, and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later.
[0064] As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or split from the final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0065] In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." In general, an M×N block may represent a set of transform coefficients or samples consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample may be used as a term for forming a picture (or image) corresponding to a pixel or picture element.
[0066] The encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, the unit in the encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as a subtractor 231.
[0067] The predictor 220 may perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a prediction block including prediction samples of the current block. The predictor 220 may determine whether to apply intra prediction or inter prediction in units of the current block or CU. The predictor 220 may generate various information about the prediction (e.g., prediction mode information, etc.) and send it to the entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0068] The intra-frame predictor 222 can predict the current block by referencing samples within the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or can be located a specific distance away from the current block. In intra-frame prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include 33 directional modes or 65 directional modes. However, this is an example, and more or fewer directional modes can be used depending on the configuration. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0069] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. Motion information can also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and reference pictures including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, for skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. For skip mode, unlike merge mode, it may not be possible to send a residual signal. For motion vector prediction (MVP) mode, the motion vector of the neighboring block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0070] The predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but can also apply intra prediction and inter prediction simultaneously. This can be referred to as inter-frame intra-frame combined prediction (CIIP) mode. In addition, the predictor can predict the block based on the intra block copy (IBC) prediction mode, or can predict the block based on the palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) for games. IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction because it derives reference blocks within the current picture. In other words, IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information about the palette table and palette index. The prediction signal generated by the predictor 220 can be used to generate a reconstruction signal or a residual signal.
[0071] The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen–Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is represented as a graph. CNT refers to a transform obtained by generating a prediction signal using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks of the same size, or to non-square blocks of variable size.
[0072] The quantizer 233 may quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in block form into a 1D vector form based on the coefficient scanning order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the 1D vector form.
[0073] The entropy encoder 240 may perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 may encode information required for video / image reconstruction (e.g., values of syntax elements, etc.) other than the quantized transform coefficients together with or separately.
[0074] Coding information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of Network Abstraction Layer (NAL) units. The video / image information may also include information about various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Furthermore, the video / image information may also include general constraint information. Herein, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded and included in the bitstream through the above-described encoding process. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit (not shown) for transmitting the signal output from the entropy encoder 240 and / or the storage unit (not shown) for storing the signal may be configured as internal / external components of the encoding device 200, or the transmitting unit may also be included in the entropy encoder 240.
[0075] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235. The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed (similar to when skip mode is applied), the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current picture, and can also be used for inter-frame prediction of the next picture through filtering as described later. In addition, luminance mapping and chroma scaling (LMCS) can be applied in the picture coding and / or reconstruction process.
[0076] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270 (specifically, the DPB of the memory 270). Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 260 can generate various information about filtering and send it to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0077] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. When inter-frame prediction is applied therethrough, the encoding device can avoid prediction mismatch in the encoding device 200 and the decoding device, and can also improve encoding efficiency.
[0078] The DPB of the memory 270 can store the modified reconstructed picture to be used as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the previously reconstructed picture. The stored motion information can be sent to the inter-frame predictor 221 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 270 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 222.
[0079] Figure 3 A rough block diagram shows a decoding device to which embodiments of the present disclosure can be applied and which performs decoding of a video / image signal.
[0080] Reference Figure 3 , the decoding apparatus 300 may be configured by including an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322.
[0081] According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by a hardware component (e.g., a decoder chipset or a processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0082] When a bit stream including video / image information is input, the decoding device 300 may generate a signal in response to the bit stream received in the video / image processing. Figure 2 The video / image information is processed in the encoding device of the reconstructed image. For example, the decoding device 300 can derive the unit / block based on the block segmentation related information obtained from the bit stream. The decoding device 300 can perform decoding using the processing unit applied in the encoding device. Therefore, the processing unit of decoding can be a coding unit, and the coding unit can be divided from the coding tree unit or the larger coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. At least one transform unit can be derived from the coding unit. And, the reconstructed image signal decoded and output by the decoding device 300 can be played by a playback device.
[0083] The decoding device 300 can receive the data in the form of a bit stream from Figure 2The received signal is output by the encoding device, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or syntax elements described later in this document can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. In more detail, the CABAC entropy decoding method can receive bins corresponding to various syntax elements from the bitstream, use the syntax element information to be decoded, the decoded information of the neighboring blocks and the block to be decoded, or the information of the symbol / bin decoded in the previous step to determine the context model, perform arithmetic decoding of the bin by predicting the probability of occurrence of the bin according to the determined context model, and generate symbols corresponding to the values of the various syntax elements. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information about the decoded symbol / bin for the context model of the next symbol / bin. Among the information decoded in the entropy decoder 310, the information about the prediction is provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (i.e., quantized transform coefficients and related parameter information) on which entropy decoding is performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. In addition, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300 , or the receiving unit may be a component of the entropy decoder 310 .
[0084] In addition, the decoding device according to the present specification may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0085] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 may dequantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) and obtain the transform coefficients.
[0086] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0087] The predictor 320 may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor 320 may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and determine a specific intra / inter prediction mode.
[0088] The predictor 320 can generate a prediction signal based on various prediction methods described later. For example, the predictor 320 can not only apply intra prediction or inter prediction to predict a block, but can also apply intra prediction and inter prediction at the same time. This can be called inter-frame intra-frame combined prediction (CIIP) mode. In addition, the predictor can predict the block based on the intra block copy (IBC) prediction mode, or can predict the block based on the palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction because it derives a reference block within the current picture. In other words, IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and notified with a signal.
[0089] The intra-frame predictor 331 can predict the current block by referencing samples within the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or at a specific distance from the current block. In intra-frame prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode to apply to the current block by using the prediction modes applied to neighboring blocks.
[0090] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 may configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the prediction information may include information indicating the inter-frame prediction mode of the current block.
[0091] The adder 340 may add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed (similar to when the skip mode is applied), the prediction block can be used as the reconstructed block.
[0092] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering as described later, or may be used for inter-frame prediction of the next picture. In addition, luminance mapping and chroma scaling (LMCS) may be applied in the picture decoding process.
[0093] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and send the modified reconstructed picture to the memory 360 (specifically, the DPB of the memory 360). Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0094] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the block in the previously reconstructed picture. The stored motion information can be sent to the inter-frame predictor 332 for use as the motion information of the spatially adjacent blocks or the motion information of the temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.
[0095] Herein, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may also be applied equivalently or correspondingly to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively.
[0096] Figure 4 An image decoding method performed by a decoding device (300) according to an embodiment of the present disclosure is shown.
[0097] Reference Figure 4 , the transform coefficient of the current block can be derived from the bitstream (S400). That is, the bitstream may include residual information of the current block, and the transform coefficient of the current block can be derived by decoding the residual information.
[0098] Reference Figure 4 , residual samples of the current block may be derived by performing at least one of dequantization and inverse transformation on the transformation coefficients of the current block ( S410 ).
[0099] When adaptive multi-transform selection (MTS) is applied, inverse transform can be performed based on at least one of DCT-2, DST-7, or DCT-8. Here, DCT-2, DST-7, DCT-8, etc. can be referred to as transform type, transform kernel, or transform core.
[0100] In this disclosure, the term "inverse transform" may refer to a separable transform. However, this is not limited to this, and the term "inverse transform" may refer to an inseparable transform, or may include both separable and inseparable transforms. Furthermore, the term "inverse transform" in this disclosure refers to a primary transform, but is not limited thereto, and may be applied to a secondary transform by modifying the primary transform to the same or similar form.
[0101] For example, as an inverse transform method, only DCT-2 and an inseparable transform may be used, or an inseparable transform may be used in addition to at least one of DCT-2, DST-7, or DCT-8, or an inseparable transform may replace the transform kernel of one or more of DCT-2, DST-7, or DCT-8.
[0102] As a more specific embodiment, when there are (DCT-2, DCT-2), (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), (DCT-8, DCT-8) as transform core candidates for separable transforms, a non-separable transform can replace or be added to one or more of these five transform core candidates. Here, the notation (transform 1, transform 2) indicates that transform 1 is applied in the horizontal direction and transform 2 is applied in the vertical direction. When a non-separable transform replaces some transform core candidates, the remaining transform core candidates except (DCT-2, DCT-2) and (DST-7, DST-7) can be replaced by a non-separable transform. However, the above-mentioned transform core candidates are merely examples, and other types of DCT and / or DST may be included, and transform skipping may be included as transform core candidates.
[0103] Non-separable transform may refer to transform or inverse transform based on a non-separable transform matrix. That is, unlike a separable transform that performs horizontal transform and vertical transform independently by separating vertical transform and horizontal transform, a non-separable transform may perform horizontal transform and vertical transform simultaneously.
[0104] For example, when non-separable transform is performed on a 4×4 block, input data X of the non-separable transform is as shown in Equation 1 below.
[0105] [Formula 1]
[0106]
[0107] When input data X is represented in vector form, vector X' can be represented as follows.
[0108] [Formula 2]
[0109] X′=[X 00 ,X 01 ,X 02 ,X 03 ,X 10 ,X 11 ,X 12 ,X 13 ,X 20 ,X 21 ,X 22 ,X 23 ,X 30 ,X 31 ,X 32 ,X 33 ] T
[0110] In this case, the non-separable transformation can be performed as shown in Equation 3 below.
[0111] [Formula 3]
[0112] F=T·X′
[0113] In Equation 3, F represents a transform coefficient vector, T represents a 16×16 non-separable transform matrix, and · represents the multiplication of a matrix and a vector.
[0114] The 16×1 transform coefficient vector F can be derived using Equation 3, and F can be reconfigured into a 4×4 block according to a predetermined scanning order. The scanning order can be horizontal scanning, vertical scanning, diagonal scanning, z scanning, raster scanning, or a predefined scanning order.
[0115] The non-separable transform set and / or the transform kernel of the non-separable transform can be configured differently based on at least one of the prediction mode (e.g., intra mode, inter mode, etc.), the width, height or number of pixels of the current block, the position of the sub-block within the current block, the syntax elements explicitly signaled, the statistical characteristics of the neighboring samples, whether to use the secondary transform or the quantization parameter (QP).
[0116] Specifically, for intra mode, the predefined intra prediction modes can be grouped into n inseparable transform sets corresponding to each other, each inseparable transform set can include k transform kernel candidates. Here, n and k can be arbitrary constants according to the rules (conditions) defined identically for the encoding device and the decoding device.
[0117] The number of inseparable transform sets and / or the number of transform core candidates included in the inseparable transform set can be configured differently according to the width and / or height of the current block. For example, for a 4×4 block, n1 inseparable transform sets and k1 transform core candidates can be configured. For a 4×8 block, n2 inseparable transform sets and k2 transform core candidates can be configured. In addition, the number of inseparable transform sets and the number of transform core candidates included in each inseparable transform set can be configured differently according to the product of the width and height of the current block. For example, when the product of the width and height of the current block is equal to or greater than 256, n3 inseparable transform sets and k3 transform core candidates can be configured, otherwise, n4 inseparable transform sets and k4 transform core candidates can be configured. That is, since the degree of change in the statistical characteristics of the residual signal varies according to the block size, the number of inseparable transform sets and transform core candidates can be configured differently to reflect this.
[0118] When the current block is divided into multiple sub-blocks, the statistical characteristics of the residual signal may be different for each sub-block, so the number of inseparable transform sets and transform core candidates may be configured differently. For example, when a 4×8 or 8×4 block is divided into two 4×4 sub-blocks and an inseparable transform is applied to each sub-block, n5 inseparable transform sets and k5 transform core candidates may be configured for the top left 4×4 sub-block, and n6 inseparable transform sets and k6 transform core candidates may be configured for other 4×4 sub-blocks.
[0119] The number of inseparable transform sets and transform core candidates can be configured differently based on a syntax element explicitly signaled. As a syntax element, information indicating one of a plurality of inseparable transform configurations can be used. For example, when three inseparable transform configurations are supported (i.e., n7 inseparable transform sets and k7 transform core candidates, n8 inseparable transform sets and k8 transform core candidates, and n9 inseparable transform sets and k9 transform core candidates), the syntax element can have values 0, 1, and 2, and the inseparable transform configuration applied to the current block can be determined based on the value of the signaled syntax element.
[0120] The number of inseparable transform sets and transform core candidates may be configured differently depending on whether a secondary transform is applied and / or which secondary transform is applied. For example, when a secondary transform is not applied, a set including n 10 A set of inseparable transformations and k 10 When applying the secondary transformation, you can apply n 11 A set of inseparable transformations and k 11 The inseparable transformation configuration of the transformation kernel candidates.
[0121] Based on the quantization parameter (QP) and / or the range to which the QP value belongs, different inseparable transform configurations may be applied. For example, when the QP value has a small value, a configuration including n 12 A set of inseparable transformations and k 12 On the other hand, when the QP value has a large value, a non-separable transform configuration including n transform kernel candidates can be applied. 13 A set of inseparable transformations and k 13 The non-separable transform configuration of the transform kernel candidates is selected. When the QP value is less than or equal to a threshold (e.g., 32), the case is classified as having a smaller QP value, otherwise, the case is classified as having a larger QP value. Alternatively, the QP value range can be divided into three or more, and a different non-separable transform configuration can be applied to each range.
[0122] For relatively large blocks, instead of using a non-separable transform corresponding to the width and height of the block, the block can be divided into multiple sub-blocks, and a non-separable transform corresponding to the width and height of the sub-block can be used. For example, when a non-separable transform is performed on a 4×8 block, the 4×8 block can be divided into two 4×4 sub-blocks, and a 4×4 block-based non-separable transform can be used for each 4×4 sub-block. Alternatively, an 8×16 block can be divided into two 8×8 sub-blocks, and an 8×8 block-based non-separable transform can be used.
[0123] A non-separable transform set can be determined based on the intra prediction mode of the current block and a mapping table. The mapping table can define the mapping relationship between predefined intra prediction modes and the non-separable transform set. The predefined intra prediction modes may include two non-directional modes and 65 directional modes. Generally, a non-separable transform has a larger transform kernel size than a separable transform. This means that the computational complexity required for the transform process is higher and the memory required to store the transform kernel is larger. Furthermore, while a separable transform may only consider statistical characteristics existing in the horizontal and / or vertical directions, a non-separable transform can consider statistical characteristics in two dimensions, including the horizontal and vertical directions, thereby providing better compression efficiency. Because the statistical characteristics and residual diversity vary depending on the directionality of the intra prediction mode, there may be cases where a non-separable transform is absolutely necessary, while there may be intra prediction modes where only a separable transform is sufficient to identify residual characteristics. Therefore, by pre-defining which transform to use based on the intra prediction mode in the encoding and decoding devices, the transform process can be designed with optimized complexity and memory requirements. Non-directional modes include planar mode numbered 0 and DC mode numbered 1, and directional modes include intra prediction modes numbered 2 to 66. However, this is merely an example, and the present disclosure may also be applied to a case where the number of predefined intra prediction modes is different.
[0124] Due to application of wide angle intra prediction (WAIP), the predefined intra prediction modes may further include intra prediction modes from -14 to -1 and intra prediction modes from 67 to 80.
[0125] Figure 5 The intra-frame prediction mode and its prediction direction according to the present disclosure are exemplarily shown. Figure 5 Modes -14 to -1, 2 to 33, and modes 35 to 80 are symmetrical with respect to mode 34 in terms of prediction direction. For example, modes 10 and 58 are symmetrical with respect to the direction corresponding to mode 34, and mode -1 is symmetrical with mode 67. Therefore, for vertical directivity modes that are symmetrical with horizontal directivity modes with respect to mode 34, the input data can be transposed and used. Transposing the input data means that the rows and columns in the M×N two-dimensional block of input data are changed to columns and rows, respectively, to form N×M data.
[0126] For example, when using a 4×4 block, the 16 data that form the 4×4 block can be appropriately arranged to form a 16×1 one-dimensional vector for the non-separable transform. In this case, the one-dimensional vector can be formed in row-major order or column-major order. The residual samples obtained from the non-separable transform can be arranged in the above order to form a two-dimensional block.
[0127] For modes -14 to -1 and 2 to 33, when the data arrangement order forming the 16×1 input vector is the row-major order, for modes 35 to 80, the input vector can be formed according to the column-major order.
[0128] Mode 34 cannot be considered a horizontal or vertical directivity mode, but in this disclosure, it is classified as belonging to the horizontal directivity mode. That is, for modes -14 to -1 and 2 to 33, the input data arrangement method of the horizontal directivity mode (i.e., row priority) is used, and for the vertical directivity mode that is symmetrical about mode 34, the input data can be transposed and used.
[0129] For non-square blocks, the symmetry in square blocks (i.e., the symmetry between mode P and mode (68-P) in an N×N block (2<=P<=33) or the symmetry between mode Q and mode (66-Q) (-14<=Q<=-1)) cannot be utilized. Therefore, in addition to the symmetry based solely on the intra-frame prediction mode, the symmetry between block shapes that are in a transposed relationship with each other, that is, the symmetry between a K×L block and an L×K block, can also be utilized. Specifically, there is a symmetry relationship between a K×L block predicted by mode P and an L×K block predicted by mode (68-P). Alternatively, there is a symmetry relationship between a K×L block predicted by mode Q and an L×K block predicted by mode (66-Q).
[0130] Since the K×L block with mode 2 and the L×K block with mode 66 can be considered symmetrical to each other, the same transform kernel can be applied to the K×L block and the L×K block. If the inseparable transform set of the intra prediction mode of the K×L block is mapped, then in order to apply the inseparable transform to the L×K block, the inseparable transform set can be derived by a mapping table corresponding to the K×L block based on mode (68-P) (instead of mode P applied to the L×K block). Alternatively, the inseparable transform set can be derived by a mapping table corresponding to the K×L block based on mode (66-Q) (instead of mode Q applied to the L×K block).
[0131] For example, to apply a non-separable transform to an L×K block, a non-separable transform set may be selected based on mode 2 rather than mode 66. Furthermore, for a K×L block, input data may be read in a predetermined order (e.g., row-first order or column-first order) to form a 1D vector, and then the corresponding non-separable transform may be applied. For an L×K block, input data may be read in a transposed order to form a 1D vector, and then the corresponding non-separable transform may be applied. That is, when a K×L block is read in row-first order, an L×K block may be read in column-first order. Conversely, when a K×L block is read in column-first order, an L×K block may be read in row-first order.
[0132] In addition, when mode 34 is applied to a K×L block, a set of inseparable transforms may be determined based on mode 34, and input data may be read in a predetermined order to form a 1D vector and corresponding inseparable transforms may be performed. When mode 34 is applied to an L×K block, a set of inseparable transforms may be determined based on mode 34, but input data may be read in a transposed order to form a 1D vector and corresponding inseparable transforms may be performed.
[0133] In the present disclosure, the method of determining an inseparable transform set and the method of forming input data are described based on K×L blocks. However, the above-mentioned symmetry of the K×L blocks can be utilized to perform inseparable transforms based on L×K blocks. Alternatively, blocks whose width is greater than their height can be restricted to be used as reference blocks. Alternatively, the symmetry can be restricted to not be used in the case of non-square blocks. In this case, the non-square blocks can use inseparable transform sets and / or transform kernel candidates with different numbering than square blocks, and can use a different mapping table to select inseparable transform sets than square blocks.
[0134] An example of a mapping table for selecting a set of inseparable transforms is as follows:
[0135] [Table 1]
[0136] predModeIntra TrSetIdx predModeIntra<0 4 0<=predModeIntra<=1 0 2<=predModeIntra<=12 1 13<=predModeIntra<=23 2 24<=predModeIntra<=44 3 45<=predModeIntra<=55 2 56<=predModeIntra<=66 1 67<=predModeIntra<=80 4
[0137] Table 1 shows an example of assigning an inseparable transform set to each intra prediction mode when there are five inseparable transform sets. The value of predModeIntra indicates the value of the intra prediction mode that takes WAIP into account, and TrSetIdx is an index indicating a specific inseparable transform set. Table 1 shows that the same inseparable transform set is applied to modes located in a symmetric direction according to the intra prediction mode. Table 1 is merely an example of using five inseparable transform sets and does not limit the total number of inseparable transform sets used for inseparable transforms.
[0138] Alternatively, as shown in Table 2, for compression performance, the inseparable transform may not be applied to WAIP.
[0139] [Table 2]
[0140] predModeIntra TrSetIdx 0<=predModeIntra<=1 0 2<=predModeIntra<=12 1 13<=predModeIntra<=23 2 24<=predModeIntra<=44 3 45<=predModeIntra<=55 2 56<=predModeIntra<=66 1
[0141] Alternatively, as shown in Table 3, instead of configuring a separate inseparable transform set for WAIP, inseparable transform sets corresponding to adjacent intra prediction modes may be shared.
[0142] [Table 3]
[0143] predModeIntra TrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0 2<=predModeIntra<=12 1 13<=predModeIntra<=23 2 24<=predModeIntra<=44 3 45<=predModeIntra<=55 2 56<=predModeIntra<=80 1
[0144] The inseparable transform set may include multiple transform core candidates, and one of the multiple transform core candidates may be selectively used. To this end, an index signaled via the bitstream may be used. Alternatively, one of the multiple transform core candidates may be implicitly determined based on the context information of the current block. Here, the context information may refer to the size of the current block or whether an inseparable transform is applied to a neighboring block. Here, the size of the current block may be defined as the width, height, the maximum / minimum value of the width and height, the sum of the width and height, or the product of the width and height.
[0145] Hereinafter, a method of determining a transform kernel for inverse transform of a current block will be described in detail.
[0146] Implementation Method 1
[0147] As described above, inverse transforms can be divided into separable transforms and non-separable transforms. A separable transform means performing transforms in the horizontal and vertical directions on a two-dimensional block, respectively, while a non-separable transform means performing a single transform on samples constituting the entire or partial two-dimensional block. When expressing a separable transform, it can be expressed as a pair of a horizontal transform and a vertical transform, and in this disclosure, it will be expressed as (horizontal transform, vertical transform).
[0148] Multiple transform sets may be defined for inverse transform of the current block. Each transform set may include one or more transform kernel candidates.
[0149] For example, one of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), or (DCT-8, DCT-8) can be applied as a separable transform, and the above four transform core candidates can be regarded as a transform set. In addition, (DCT-2, DCT-2) can be regarded as a transform set. A transform skip in which no transform is applied can also be regarded as a transform set, and (DCT-2, DCT-2) and a transform skip can be regarded as a transform set. In the present disclosure, a transform core can refer to one transform (e.g., DCT-2, DST-7), or can refer to two transform pairs (e.g., (DCT-2, DCT-2)).
[0150] As another example of a transform set, there may be the aforementioned inseparable transform set. In the present disclosure, an inseparable transform applied as a main transform may be represented as an inseparable main transform (NSPT). In NSPT, multiple inseparable transform sets may be configured, and each inseparable transform set may include one or more transform cores as transform core candidates. In the case of NSPT, one of the multiple inseparable transform sets is selected based on the intra prediction mode, and the multiple inseparable transform sets used for NSPT may be represented as an NSPT set list. This is as described above, and its detailed description will be omitted here.
[0151] A group of one or more transform sets available for the current block can be configured from a plurality of predefined transform sets. The group of one or more transform sets can be configured in predetermined region units to which the current block belongs, hereinafter referred to as a set. Here, the predetermined region unit can be at least one of a picture, a slice, a coding tree unit row (CTU row), or a coding tree unit (CTU).
[0152] For example, the transform set consisting of (DCT-2, DCT-2) is called S1, and the transform set consisting of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8) is called S2. In addition, the above NSPT set list may include N inseparable transform sets, which are respectively called S 3,1 、S 3,2 ,...,S 3,N Here, N may be 35, but is not limited thereto.
[0153] When the intra prediction mode based on the current block is selected S 3,13 As an inseparable transform set for NSPT, the transform kernel applied to the current block can belong to S1, S2 or S 3,13 In this case, the set available for the current block can be represented as {S1, S2, S 3,13}.
[0154] As described above, since the set according to the present disclosure is a group of one or more transform sets that can be used for the current block, the set can be configured differently based on the context of the current block. Here, the context may include at least one of shape, size, or intra prediction mode. If a total of K contexts are defined, K sets can be generated, and each set can be represented as C i (i=1, 2, ..., N). For example, when the block sizes to which NSPT is applicable are 4×4, 8×8, 16×16, and 32×32 and one of a total of 35 inseparable transform sets is selected based on the intra prediction mode, if different transform kernels are applied to each block size, a total of 4×35=140 contexts can be defined.
[0155] A set may be configured based on the context of the current block. In this case, a process may be performed to select one of multiple transform sets belonging to the set and one of multiple transform core candidates belonging to the selected transform set. Here, the selection of the transform set and the transform core candidate may be performed implicitly based on the context of the current block, or based on an explicitly signaled index. Alternatively, the process of selecting one of the multiple transform sets belonging to the set and the process of selecting one of the multiple transform core candidates belonging to the selected transform set may be performed separately. For example, an index for selecting a transform set may be first signaled, and one of the multiple transform sets belonging to the set may be selected based on the index. Then, an index indicating one of the multiple transform core candidates belonging to the transform set may be signaled, and one of the transform core candidates may be selected from the transform set based on the signaled index. The transform core of the current block may be determined based on the selected transform core candidate. Alternatively, the selection of a transform set from the set may be performed implicitly based on the context of the current block, and the selection of a transform core candidate from the selected transform set may be performed based on the signaled index. Alternatively, selection of a transform set from a set can be performed based on a signaled index, and selection of a transform core candidate from the selected transform set can be performed implicitly based on the context of the current block. Alternatively, selection of a transform set from a set can be performed implicitly based on the context of the current block, and selection of a transform core candidate from the selected transform set can also be performed implicitly based on the context of the current block. Of course, when the number of transform sets in a set is one, the index used to select a transform set may not be signaled. Similarly, when the number of transform core candidates in the selected transform set is one, the index used to indicate a transform core candidate may not be signaled. Alternatively, an index indicating one of all transform core candidates belonging to the current set may be signaled. In this case, the process of selecting a transform set from the set may be omitted. In this case, all transform sets belonging to the set may be shuffled based on priority. For example, when assigning smaller-length binary codes (e.g., truncated unary codes) to smaller-valued indices, it may be advantageous to assign smaller-valued indices to transform core candidates that are more conducive to improving encoding performance. When all transform kernel candidates belonging to a set are shuffled according to priority, different shuffling may be applied to each set. In addition, instead of shuffling all transform kernel candidates belonging to a set, only some of them may be selectively shuffled.
[0156] Implementation Method 2
[0157] A transform kernel for inverse transform of the current block may be determined based on MTS (Multiple Transform Select).
[0158] The MTS according to the present disclosure may use at least one of DST-7, DCT-8, DCT-5, DST-4, DST-1, or IDT (Identity Transform) as a transform core. In addition, the MTS according to the present disclosure may also include a transform core of DCT-2.
[0159] In the present disclosure, multiple MTS sets for MTS can be defined. Based on the size of the current block and / or the intra prediction mode, one of the multiple MTS sets can be determined. For example, when determining an MTS set, 16 transform block sizes can be considered, and for the directional mode, the symmetry between the shape of the transform block and the intra prediction mode can be considered. For WAIP (wide angle intra prediction) mode (i.e., -1 to -14 (or -15), 67 to 80 (or 81)), the MTS set corresponding to mode 2 can be applied for mode -1 to -14 (or -15), and the MTS set corresponding to mode 66 can be applied for mode 67 to 80 (or 81). A separate MTS set can be assigned for MIP (matrix-based intra prediction) mode.
[0160] For example, MTS sets according to transform block size and intra prediction mode may be assigned / defined as shown in Table 4 below.
[0161] [Table 4]
[0162]
[0163] Table 4 shows the assignment of MTS sets according to 16 transform block sizes and intra prediction modes. The number of predefined MTS sets is 80, and the index indicating one of the 80 MTS sets may have a value of 0 to 79, as shown in Table 4.
[0164] [Table 5]
[0165]
[0166]
[0167]
[0168] Table 5 shows the transform core candidates included in each MTS set described in Table 4. Each MTS set may consist of six transform core candidates. The transform core candidate index has a value from 0 to 5 and may indicate one of the six transform core candidates. Here, each transform core candidate may be a combination of a horizontal transform core and a vertical transform core for separable transform, and 25 transform core candidates with indices from 0 to 24 may be defined.
[0169] [Table 6]
[0170]
[0171] Table 6 shows an example of the 25 transform core candidates described in Table 5. Specifically, the horizontal and vertical transforms of the transform core candidates are represented as (horizontal transform, vertical transform). For each transform core candidate index, the horizontal / vertical transform when the intra prediction mode is less than 35 can be opposite to the horizontal / vertical transform when the intra prediction mode is greater than or equal to 35. When the intra prediction mode value is greater than or equal to 35, a mode symmetric with respect to mode 34 can be derived, and an MTS set can be selected from Table 4 based on this mode. Furthermore, the symmetry of the block shape can be considered. When the original transform block has a W×H size, by making it symmetric, the original transform block can be considered to have an H×W size, and an MTS set can be selected from Table 4. Here, the value of the intra prediction mode can be the value of a modified intra prediction mode. That is, as the WAIP mode value, for values from -14 (or -15) to -1, it is modified to mode 2, and for values from 67 to 80 (or 81), it is modified to mode 66. For the remaining modes, the value of the original intra prediction mode can be set to the value of the modified intra prediction mode. In this case, since the extended mode of WAIP is also configured symmetrically with respect to the mode 34, the symmetry with respect to the mode 34 can be applied to all directional modes except the planar mode and the DC mode.
[0172] For example, when predicting a 16×32 block based on mode 54, mode 14 (=68-54) can be derived as a mode symmetric to mode 54, and the block size can be considered as 32×16. In this case, the MTS set with index 72 can be selected, as defined in Table 4.
[0173] When the MIP mode is applied, the MTS set assigned to the MIP mode may be selected based on the size of the current block, without considering the symmetry of the block shape. Alternatively, when the MIP mode is applied, the MTS set assigned to the MIP mode may be selected based on the symmetric block size, considering the symmetry of the block shape. For example, when the MIP mode is applied for an 8×16 block, the 8×16 block may be regarded as a 16×8 block symmetrical thereto, and the MTS set with index 49 may be selected as defined in Table 4. Alternatively, when the MIP mode is applied, the intra prediction mode may be regarded as a planar mode. In this case, the MTS set assigned to the MIP mode may be selected based on the size of the current block, without considering the symmetry of the block shape. Alternatively, the MTS set assigned to the MIP mode may be selected based on the symmetric block size, considering the symmetry of the block shape.
[0174] For the MIP mode, a flag can be used to indicate whether the MIP mode is applied in the transposed mode. When the MIP mode is applied to the current block of M×N and the flag indicates that the transposed mode is applied, the intra-frame prediction mode can be regarded as the plane mode, and the current block of M×N can be regarded as the N×M block. That is, from Table 4, the MTS set corresponding to the block size of N×M and the plane mode can be selected. As described in Table 6, when the value of the intra-frame prediction mode is greater than or equal to 35, the horizontal transform and the vertical transform are exchanged, but since the intra-frame prediction mode of the current block is regarded as the plane mode, the horizontal transform and the vertical transform of the transform core candidate may not be exchanged. Alternatively, when the MIP mode is applied to the current block of M×N and the flag indicates that the transposed mode is applied, the intra-frame prediction mode may not be regarded as the plane mode, and the current block of M×N can be regarded as the N×M block. That is, from Table 4, the MTS set corresponding to the block size of N×M and the MIP mode can be selected.
[0175] In Table 5, the transform core candidate selected by the transform core candidate index can be set as the transform core of the current block. Alternatively, based on the size of the current block, at least one of the horizontal transform or the vertical transform of the selected transform core candidate can be changed to another transform core. For example, when the transform core candidate index is 3 and both the width and height of the current block are less than or equal to 16, at least one of the horizontal transform or the vertical transform of the transform core candidate corresponding to the transform core candidate index 3 can be changed to another transform core. In this case, the horizontal transform and the vertical transform can be changed independently of each other. When the difference between the value of the intra-frame prediction mode and the value of the horizontal mode of the current block (or the absolute value of the difference) is less than or equal to a predetermined threshold, the vertical transform of the selected transform core candidate can be changed to an IDT (identity transform). When the difference between the value of the intra-frame prediction mode and the value of the vertical mode of the current block (or the absolute value of the difference) is less than or equal to a predetermined threshold, the horizontal transform of the selected transform core candidate can be changed to an IDT (identity transform). Here, the threshold can be determined based on the width and height of the current block, as shown in Table 7 below.
[0176] [Table 7]
[0177]
[0178] Table 7 is used to change the horizontal transform and / or vertical transform of a transform core candidate selected by a transform core candidate index to another transform core, and defines a threshold value according to the size of a transform block.
[0179] The six transform core candidates that make up an MTS set can be distinguished by transform core candidate indices from 0 to 5, as defined in Table 5. The transform core candidate index can be signaled via the bitstream. A flag indicating whether the MTS set is available / applied (MTS enable flag or MTS flag) can be signaled, and when the flag indicates that the MTS set is available / applied, the transform core candidate index can be signaled. The MTS flag can consist of a bin, and one or more contexts for CABAC-based entropy coding (hereinafter referred to as CABAC context) can be assigned to the bin. For example, different CABAC contexts can be assigned to non-MIP mode and MIP mode, respectively.
[0180] Based on the context of the current block described above, the number of transform core candidates available for the current block can be set differently. For example, as the context of the current block, the sum of the absolute values of all or some transform coefficients in the current block can be considered. The sum of the absolute values of the transform coefficients is called AbsSum. When AbsSum is less than or equal to T1, only one transform core candidate corresponding to transform core candidate index 0 may be available. When AbsSum is greater than T1 and less than or equal to T2, four transform core candidates corresponding to transform core candidate indices 0 to 3 may be available. When AbsSum is greater than T2, six transform core candidates corresponding to transform core candidate indices 0 to 5 may be available. Here, T1 may be 6 and T2 may be 32, but this is just an example.
[0181] When AbsSum is less than or equal to T1, since the number of transform core candidates available for the current block is 1, the transform core candidate corresponding to the transform core candidate index 0 can be set as the transform core of the current block without signaling the transform core candidate index. When AbsSum is greater than T1 and less than or equal to T2, since four transform core candidates are available, one of the four transform core candidates can be selected based on the transform core candidate index with two bins. That is, the transform core candidate indices 0 to 3 can be signaled as 00, 01, 10, and 11, respectively. For these two bins, the MSB (most significant bit) can be signaled first, and the LSB (least significant bit) can be signaled later. Different CABAC contexts can be assigned to each bin. For example, a CABAC context other than the CABAC context assigned for the MTS flag can be assigned to each bin of the two bins. Alternatively, bypass coding can be applied without assigning a CABAC context to the two bins. When AbsSum is greater than T2, the transform core candidate index has a value of 0 to 5, so it is impossible to represent the transform core candidate index with only two bins. In this case, the transform core candidate index can be represented by assigning two or more bins, such as truncated binary coding. For each bin assigned by the truncated binary coding method, a CABAC context can be assigned, or bypass coding can be applied without assigning a CABAC context. Alternatively, the CABAC context can be assigned to some of the multiple bins (e.g., the first bin, or the first bin and the second bin), and bypass coding can be applied to the remaining bins.
[0182] Implementation 3
[0183] A transform core of the current block may be determined based on a transform set including one or more transform core candidates. The transform core of the current block may be derived as one of the one or more transform core candidates belonging to the transform set.
[0184] The process of determining the transform core of the current block may include at least one of 1) determining a transform set for the current block or 2) selecting a transform core candidate from the transform set for the current block. The process of determining the transform set may be a process of selecting one of a plurality of transform sets predefined in the same manner in the encoding device and the decoding device. Alternatively, the process of determining the transform set may be a process of configuring one or more transform sets applicable to the current block from a plurality of transform sets predefined in the same manner in the encoding device and the decoding device, and selecting one of the configured transform sets. Alternatively, the process of determining the transform set may be a process of configuring a transform set based on the transform core candidates applicable to the current block from a plurality of transform core candidates predefined in the same manner in the encoding device and the decoding device.
[0185] When the transform set of the current block includes a plurality of transform core candidates, a process of selecting one of the plurality of transform core candidates for the current block may be performed. However, when the transform set of the current block includes one transform core candidate (i.e., when the number of transform core candidates available for the current block is 1), the transform core of the current block may be set as the corresponding transform core candidate.
[0186] The transform set according to the present disclosure may refer to the (inseparable) transform set in the above-mentioned embodiment 1, or may refer to the MTS set in embodiment 2. Alternatively, the transform set may be defined separately from the (inseparable) transform set in embodiment 1 or the MTS set in embodiment 2. In this case, the transform set may include one or more specific transform kernels as transform kernel candidates. One specific transform kernel may be defined as a pair of transform kernels for horizontal transform and a transform kernel for vertical transform, or may be defined as one transform kernel that is equally applied to horizontal transform and vertical transform.
[0187] In an embodiment of the present disclosure, the process of applying NSPT (an inseparable transform applied as a main transform) is described in detail. NSPT can be applied to the entire or partial transform block. Based on the forward NSPT, the residual samples present in the area where the NSPT is applied can be used as a 1D vector input of the NSPT. In other words, the residual samples present in the entire or part of a single transform block (referred to as a region of interest (ROI) in the present disclosure) can be collected as a 1D vector and configured as input. Then, when the forward NSPT is applied, the main transform coefficients can be obtained. Conversely, when the backward NSPT is applied to the main transform coefficients, a 1D vector output can be obtained. The residual samples of the ROI can be obtained by arranging the respective element values of the corresponding output vector at a determined position within the 2D transform block.
[0188] For a non-separable transform kernel used for NSPT, the matrix dimensions can be determined based on the size of the ROI. In this disclosure, a transform kernel may be referred to as a transform type or transform matrix, and a non-separable transform kernel used for NSPT may be referred to as an NSPT kernel. For example, when the current block is an M×N transform block, the ROI is the entire M×N transform block, and a square NSPT is applied, the corresponding transform matrix dimensions may be MN×MN. For example, when the ROI is the entire 8×8 transform block, the NSPT kernel dimensions may be 64×64.
[0189] According to an embodiment of the present disclosure, when NSPT is applied to the residual generated by intra-frame prediction, the NSPT kernel can be adaptively determined according to the intra-frame prediction mode. Since the statistical characteristics of the residual block can vary according to the intra-frame prediction mode, compression efficiency can be improved by adaptively determining the NSPT kernel according to the intra-frame prediction mode.
[0190] A shared NSPT core applied to at least one intra prediction mode can be configured. As described above, an inseparable transform set can be determined based on the intra prediction mode of the current block and a mapping table. The mapping table can define the mapping relationship between predefined intra prediction modes and inseparable transform sets. The predefined intra prediction modes can include two non-directional modes and 65 directional modes.
[0191] As an embodiment, intra-frame prediction modes can be grouped into intra-frame prediction mode groups. One NSPT core can be assigned to the intra-frame prediction mode group, or multiple NSPT cores can be assigned. In other words, an inseparable transform set (NSPT set) including at least one NSPT core can be assigned to the intra-frame prediction mode group. The inseparable transform set can be mapped to the intra-frame prediction mode, and one of the N NSPT cores included in the inseparable transform set can be selected.
[0192] As an example, the intra prediction group may include adjacent prediction modes (e.g., modes 17, 18, and 19). In addition, the intra prediction group may include modes having symmetry. For example, in the above Figure 5 , the directional mode may be symmetric around the diagonal mode (i.e., intra-prediction mode 34). In this case, two symmetric modes may be configured as a group (or a pair). For example, mode 18 and mode 50 may be included in the same group because they are symmetric around mode 34. However, for modes having symmetry, a process of transposing the 2D input block and then configuring a one-dimensional input vector may be added before applying the forward NSPT kernel. For example, when the intra-prediction mode is less than or equal to 34, a one-dimensional input vector may be derived from the corresponding input block in row-first order without transposing the 2D input block. When the intra-prediction mode is greater than 34, a one-dimensional input vector may be configured by first transposing the 2D input block and then reading the corresponding input block in row-first order, or by keeping the 2D input block as is and reading the corresponding input block in column-first order.
[0193] Table 8 below shows a mapping table for allocating NSPT sets according to intra prediction modes. Referring to Table 8, a total of 35 NSPT sets from 0 to 34 can be defined. The NSPT set allocated to the nearest general directional mode can be allocated to the extended WAIP mode (ie, Figure 5 In other words, NSPT set 2 may be allocated to the extended WAIP mode.
[0194] [Table 8]
[0195]
[0196] The NSPT set may include at least one NSPT core (or core candidate). In other words, the NSPT set may include N NSPT core candidates. As an example, N may be set to a value equal to or greater than 1, such as 1, 2, 3, 4, etc. The core applied to the current block among the at least one NSPT core included in the NSPT set may be signaled using an index. In the present disclosure, the corresponding index may be referred to as an NSPT index. As an example, the NSPT index may have values of 0, 1, 2, ..., N-1.
[0197] In addition, as an embodiment, when the number of NSPT core candidates is 1, the NSPT index value can be fixed to 0. In this case, the NSPT index can be inferred without being separately signaled. In addition, a flag indicating whether NSPT is applied can be signaled separately from the NSPT index. In the present disclosure, the corresponding flag can be referred to as an NSPT flag.
[0198] When the NSPT flag value is 1, NSPT may be applied. When the NSPT flag value is 0, NSPT may not be applied. When the NSPT flag is not signaled, the NSPT flag value may be inferred to be 0. As an example, when the NSPT flag value is 1, an NSPT index may be applied. One of the N core candidates included in the NSPT set selected by the intra prediction mode may be specified based on the signaled NSPT index.
[0199] In an embodiment, the entropy encoding method of the NSPT index may be defined in various ways by considering the number (N) of NSPT cores included in the NSPT set. For example, as a method of mapping values 0 to N-1 to a bin string (i.e., a binarization method), truncated unary binarization, truncated binarization, and fixed-length binarization methods may be used.
[0200] For example, when the number N of core candidates configuring the NSPT set is 2, one of the two candidates can be specified using one bin. For example, 0 can indicate the first candidate and 1 can indicate the second candidate. In addition, when the value of N is 3 and truncated unary binarization is applied, two bins can be used to specify the candidates. For example, the first, second, and third candidates can be binarized to 0, 10, and 11, respectively, and notified using a signal. As an embodiment, the binarized bins can be encoded using context coding or bypass coding.
[0201] In the present disclosure, a reduced primary transform (RPT) method using a dimensionality reduction transform kernel as a primary transform is described. As described above, when forward NSPT is applied, samples belonging to a 2D residual block can be arranged (or rearranged) into a 1D vector according to row-first order (or column-first order). Thereafter, the transform matrix used for NSPT can be multiplied by the arranged vector. When the corresponding 2D residual block is an M×N block (M is the horizontal length and N is the vertical length), the length of the rearranged 1D vector can be M*N. In other words, the corresponding 2D residual block can also be represented as an M*N×1-dimensional column vector. In the present disclosure, for convenience, M*N can be represented as MN. In this case, the dimension of the corresponding transform matrix can be MN×MN. In summary, forward NSPT can operate in a manner such that an MN×1 transform coefficient vector is obtained by multiplying the left side of an MN×1 vector by the corresponding MN×MN transform matrix.
[0202] When RPT is applied, r transform coefficients can be obtained by multiplying an r×MN matrix instead of an MN×MN matrix as the forward NSPT transform matrix described above. Here, r represents the number of rows of the transform matrix, and MN represents the number of columns of the transform matrix. According to an embodiment of the present disclosure, the value of r can be set to be less than or equal to MN. In other words, the existing forward NSPT transform matrix includes MN rows, and each row is a 1×MN row vector and a transform basis vector of the corresponding NSPT transform matrix. The corresponding transform coefficient can be obtained by multiplying each transform basis vector by an MN×1 sample column vector.
[0203] Since the existing forward NSPT transform matrix consists of MN row vectors, MN transform coefficients (i.e., MN×1 transform coefficient column vectors) can be obtained by applying the forward NSPT. In addition, for the forward RPT, the transform matrix can be composed of r transform basis vectors instead of MN transform basis vectors. Therefore, when the forward RPT is applied, r transform coefficients (i.e., r×1 transform coefficient column vectors) can be obtained instead of MN.
[0204] The RPT kernel can be configured by selecting r transform basis vectors as partial transform basis vectors for configuring the MN×MN forward NSPT kernel. In the present disclosure, the transform kernel can be referred to as a transform type or a transform matrix, and the inseparable transform kernel for NSPT can be referred to as an RPT kernel. In other words, when selecting r 1×MN row vectors from the MN×MN forward NSPT kernel, it may be advantageous to select the most important transform basis vectors from the perspective of coding performance. Specifically, in terms of energy concentration through transformation, by multiplying the forward NSPT transform matrix, more energy can be concentrated on the transform coefficients that appear first. In other words, the transform basis vectors located on the top side of the forward NSPT transform matrix can generate transform coefficients with greater energy. Taking this into account, the r×MN forward RPT kernel can be configured (or derived) by taking r from the top side of the forward NSPT kernel.
[0205] The RPT according to the present disclosure only uses a portion (i.e., r) of the transform coefficients obtained by applying the existing NSPT, so the energy of the original signal may be partially lost. In other words, distortion between the original signal and the natural signal may occur through the corresponding processing. However, since only r transform coefficients are generated instead of MN transform coefficients by applying the RPT, the number of bits required to encode the corresponding transform coefficients can be reduced. Therefore, for signals (e.g., image residual signals) where a large amount of energy is concentrated in a small number of transform coefficients, the gain obtained by reducing signaling bits can be significantly large, thereby improving encoding performance.
[0206] The backward NSPT is a transformation matrix and can be the transposed matrix of the forward NSPT kernel described above. In this case, the input data can be a transform coefficient signal rather than a sample signal such as a residual signal. Specifically, when the forward NSPT transform matrix is G and the sample signal rearranged as a 1D vector is x, the transform coefficient vector obtained by multiplying the corresponding transform matrix by the left side can be expressed as shown in Equation 4 below.
[0207] [Formula 4]
[0208] y=Gx
[0209] Referring to Equation 4, x and y can be MN×1 column vectors. G can be in the form of an MN×MN matrix. The backward NSPT process can be expressed as shown in Equation 5 below using the same variables.
[0210] [Formula 5]
[0211] x=G T y
[0212] In formula 5, G TIt means the transposed matrix of G. The forward RPT operation and the backward RPT operation according to the present disclosure can also be expressed by these two formulas. However, when RPT is applied, y is an r×1 column vector instead of an MN×1 column vector, and G is an r×MN matrix instead of an MM×MN matrix. In other words, even when RPT is applied instead of NSPT, the dimension of the sample signal (e.g., image residual signal) does not change, which may mean that the original number of sample signals (i.e., MN sample signals) can be reconstructed using only r transform coefficients through backward RPT. In other words, the original MN sample signals can be reconstructed by encoding only r transform coefficients that are less than MN, which can improve encoding performance.
[0213] In an embodiment of the present disclosure, an RPT structure is proposed that defines the value of r by considering the statistical characteristics of the residual block and derives a residual block of the existing transform block size from the residual block of a reduced size determined according to the defined r value. If another additional transform (i.e., a secondary transform) is applied to predict the statistical distribution of the main transform coefficients, quantization is applied to the main transform coefficients, so that the quantized non-zero coefficients can be concentrated in a relatively low-frequency domain. Therefore, the reduced secondary transform of the statistical distribution of the main transform coefficients can relatively simply define the statistical characteristics of the main transform coefficients in the form of setting the value of r for a given low-frequency domain. However, as a technique for defining the value of r by considering the statistical characteristics of samples within the residual block whose characteristics are very different from the distribution of the main transform coefficients, the RPT according to the present disclosure is fundamentally different from the reduced secondary transform. Below, various embodiments of determining the RPT kernel as a dimensionality reduction transform matrix are described. In other words, the method of determining or defining the value of r in the RPT is described below.
[0214] In an embodiment of the present disclosure, the value of r in the RPT can be determined by considering the worst-case complexity allowed by the transform system. As an embodiment, the worst-case complexity can be calculated based on the number of multiplications per sample. MN*r multiplications are required to apply the RPT in both the forward and backward directions based on an M×N block. Since a 2D block consists of a total of MN samples, the number of multiplications per sample can be calculated as (MN*r) / MN=r. Therefore, the value of r can be configured to maintain less than or equal to the maximum number of multiplications allowed per sample. For example, when the maximum possible number of multiplications per sample is set to 16 for a 16×16 block, the value of r can be determined to be less than or equal to 16. In other words, the forward RPT kernel can be set to 16×256.
[0215] In another embodiment, memory usage can be viewed as a measure of worst-case complexity. As an example, the memory size allowed per core can be set. For example, when each core coefficient requires p bytes (in this disclosure, the individual elements that configure the transform core are referred to as core coefficients) and memory usage is set to less than or equal to q bytes per core, the value of r can be set to less than or equal to q / (MN*p). For example, when for a forward RPT core of 16×16 blocks, p is 1 byte, and memory usage is set to less than or equal to 8KB per core (q=8KB=2 13 byte), the value of r can be set to less than or equal to 32.
[0216] Additionally, as another example, memory usage and / or the number of multiplications per sample can be considered as a measure of worst-case complexity. For example, when the maximum possible number of multiplications per sample is set to 16 for a 16×16 block, and memory usage is set to less than or equal to 8KB per core (core coefficients are represented as 1 byte), the value of r can be set to less than or equal to 16.
[0217] In addition, in an embodiment, the value of r configuring the RPT core can be determined by specific information. In other words, the value of r configuring the RPT core can be determined based on predefined coding parameters. For example, the value of r can be determined based on the size of the block. In other words, the RPT core can be variably determined based on the size of the block. Here, the block can be at least one of a coding block, a transform block, and a prediction block. In addition, for example, the value of r can be determined based on prediction information. Here, the prediction information may include information about inter / intra prediction, intra prediction mode information, etc. In addition, for example, the value of r can be determined based on information notified by a signal (the value of a syntax element). For example, the value of r can be variably determined according to a quantization parameter value. In addition, in terms of complexity improvement, a fixed value predefined as the r value can be used, and the predefined fixed value can be determined based on the information notified by the signal.
[0218] When the sample signal is multiplied by the RPT kernel r×MN, r transform coefficients can be obtained. The obtained r transform coefficients can be arranged according to a predefined scanning order of the transform coefficients (for example, a forward / backward zigzag scanning order, a forward / backward horizontal scanning order, a forward / backward vertical scanning order, a forward / backward diagonal scanning order, a scanning order specified based on an intra-frame prediction mode, etc.). When the transform coefficients obtained by applying the forward RPT are arranged according to such a scanning order (for example, a scanning order in units of coefficient groups (CGs) can also be applied), if the value of r is less than MN, the interior of the M×N block may not be completely filled with r transform coefficients, and thus blank space may appear. As an embodiment of the present disclosure, the above-mentioned blank space can be predicted in the following manner taking into account the characteristics of the residual signal.
[0219] - Values that can fill empty spaces by using the values of available neighboring pixels.
[0220] - The value of the blank space may be filled based on the values of available neighboring pixels and the intra prediction mode. For example, the value of the blank space may be predicted by performing intra prediction based on the values of available neighboring pixels and the intra prediction mode.
[0221] - Values that can fill empty spaces by using predefined fixed values (eg, 0).
[0222] - Values of empty spaces can be filled from available neighboring pixels by using a predetermined intra prediction mode (eg, planar mode).
[0223] In the present disclosure, filling the blank space with 0 in the above example may be referred to as zeroing. When filling the blank space with 0, the following embodiments may be applied. When a non-zero transform coefficient is detected (or parsed) in the corresponding blank space portion during parsing of the transform coefficients on the decoding device side, it may be considered (or inferred) that RPT is not applied. In other words, when a non-zero transform coefficient exists in a predefined area representing the corresponding blank space, it may be considered that RPT is not applied. In this case, signaling (or parsing) of a flag indicating whether RPT is applied and / or specifying an index of one of multiple RPT core candidates may not be performed. As an example, when a non-zero transform coefficient exists in a predefined area representing the corresponding blank space, a predefined variable value may be updated, and it may be inferred that RPT is not applied based on the updated variable value.
[0224] In embodiments of the present disclosure, whether to apply the RPT may be determined based on the size and / or form of the block. Furthermore, the RPT kernel may be determined differently depending on the size and / or form of the block. Since the value of r may differ depending on the size and / or form of the block (i.e., for each M×N block), the blank space may differ depending on the size and / or form of the block. Therefore, the region used to check whether non-zero transform coefficients are detected may be defined differently for blocks of different sizes and / or forms. In other words, the zeroing region may be determined differently.
[0225] As an example, when a 16×64 matrix is applied as a forward RPT matrix for an 8×8 block, the value of r may be 16. In this case, when the CG is a 4×4 sub-block, only the upper left 4×4 block may be filled with non-zero RPT transform coefficients, and the remaining three 4×4 sub-blocks (i.e., the upper right, lower left, and lower right sub-blocks) as blank spaces may be filled with a value of 0. In this case, when non-zero transform coefficients are detected in the corresponding remaining three 4×4 sub-block regions during the decoding process, it may be considered that RPT is not applied. Also, as described above, a flag indicating whether RPT is applied or an index specifying one of a plurality of RPT core candidates may not be signaled.
[0226] In addition, as an example, when a 32×128 matrix is applied as the forward RPT matrix of a 16×8 block (i.e., the value of r is 32) and the CG is a 4×4 sub-block, only the two CGs in the scanning order can be filled with non-zero RPT transform coefficients. For example, the upper left 4×4 sub-block and the 4×4 sub-block adjacent to the bottom of the upper left sub-block can be filled with corresponding RPT transform coefficients. The area filled with 0 as the blank space can be determined as the remaining area except for the corresponding two 4×4 sub-blocks. The RPT kernel can be determined differently according to the size and / or form of the block, and as described, the blank space can be determined differently for 8×8 blocks and 16×8 blocks.
[0227] As an embodiment, when the value of r is a multiple of the CG size and the transform coefficients are scanned in units of CG, if a non-zero transform coefficient is detected in a CG belonging to a blank space, the flag and / or index related to the RPT may not be signaled. In other words, the transform coefficients inside the CG may be scanned in a specified order for each CG, and the transform coefficients inside the CG may be scanned in the same manner after moving to the next CG in units of CG in the scanning order. In existing image compression technology, since a flag indicating whether there is a non-zero transform coefficient in the corresponding CG is first signaled for each CG, it is possible to determine whether to apply the RPT using only the corresponding information, which can reduce signaling overhead and related implementation complexity.
[0228] As described above, when RPT is applied, if a non-zero transform coefficient is detected in a 0-padded blank space region, RPT may not be applied. In this case, signaling of information related to RPT may be omitted. However, since it is impossible to determine whether RPT is applied when no non-zero transform coefficient is detected in the corresponding blank space region, a flag indicating whether RPT is applied may be parsed after parsing (or signaling) the relevant transform coefficients to ultimately determine whether RPT is applied.
[0229] As an embodiment, a forward secondary transform may be additionally applied to the transform coefficients generated by applying the RPT. Alternatively, a forward secondary transform may be additionally applied to the region in the M×N block where the corresponding generated transform coefficients are located. In the present disclosure, from the perspective of the forward secondary transform, the corresponding region or a portion of the corresponding region may be referred to as an ROI. For the backward direction, the backward secondary transform may be applied first, and then the backward RPT may be applied. Specifically, a region in which r transform coefficients generated by applying the forward RPT are arranged or a portion of the corresponding region may be set as an ROI to apply the forward secondary transform. In this case, when a 16×64 forward RPT transform matrix is applied to an 8×8 region, the 16 transform coefficients generated may be located in the upper left 4×4 sub-block, and the corresponding sub-block region may be set as an ROI to apply the forward secondary transform to the corresponding ROI.
[0230] In addition, the RPT core can adjust the coefficient value by considering operations such as integer operations or fixed-point operations. In other words, the RPT core can be configured to perform transformations (rather than theoretical orthogonal transformations or non-orthogonal transformations (here, orthogonal transformations and non-orthogonal transformations represent transformations in which the norm of each transformation basis vector is 1)) through integer operations (or fixed-point operations) in an actual encoding and decoding system by appropriately scaling the kernel coefficients belonging to the corresponding kernel. Even when RPT is applied, it can be reflected equally as multiplying the scaling factor when applying a separable transform in the existing image compression technology. In this case, a separable transform or an inseparable transform (including RPT) can be performed while maintaining other processing other than the transform (for example, quantization processing and dequantization processing).
[0231] The integerized coefficients of the RPT kernel can be obtained by multiplying the transformation basis vector by the above-mentioned scaling value. As an embodiment, multiplying by the scaling value may include applying operations such as rounding, flooring, ceilinging, etc. to each kernel coefficient. In other words, the integerized RPT kernel obtained by the above method can be defined and used for transform / inverse transform processing. As described above, when the kernel coefficients of the scaled integer value are obtained by operations such as rounding, flooring, ceilinging, etc., the maximum value and the minimum value can be obtained for all the kernel coefficients, and therefore the number of bits sufficient to represent all the kernel coefficients can be obtained from the maximum value and the minimum value. For example, when the maximum value is less than or equal to 127 and the minimum value is greater than or equal to -128, all integer kernel coefficients can be represented by 8 bits (particularly by a 2's complement expression, etc.).
[0232] Typically, when the maximum value is less than or equal to (2 (N-1) -1) and the minimum value is greater than or equal to -2 (N-1) When all integer kernel coefficients can be represented by N bits. When the maximum value is greater than (2 (N-1) -1) or the minimum value is less than -2 (N-1) When all kernel coefficients need to be multiplied by 2, it may not be possible to represent all integer kernel coefficients in N bits. In this case, 1) all kernel coefficients can be multiplied by a scaling value to adjust them to fall within the range of N bits, or 2) the number of bits required to represent the kernel coefficients can be increased (i.e., N+1 bits or more). -p (p>=1) To represent them with N bits, we can multiply them by 2 - p to compensate them so that they can be integrated into the existing encoding / decoding process. As an implementation method, multiply by 2 p This can be achieved by additionally performing a left shift of p bits or by reducing the right shift amount applied in the quantization or dequantization process by p.
[0233] The above method can be used to represent all kernel coefficients with 8 bits, 9 bits, 10 bits, etc. Of course, the scaling value of the kernel coefficient can be set differently for each block size or kernel, and the number of bits used to represent the kernel coefficient can be set differently.
[0234] The NSPT described above may be applied based on at least one of the size, tree type, or component type of the current block. For example, whether to apply the NSPT may be determined based on at least one of the size, tree type, or component type of the current block. An NSPT index may be signaled based on at least one of the size, tree type, or component type of the current block. An NSPT set or NSPT core may be derived based on at least one of the size, tree type, or component type of the current block.
[0235] The allowed transform block sizes predefined in the decoding device can be roughly divided into two groups. Either of the two groups (hereinafter referred to as the first group) can refer to a set of block sizes to which NSPT is applicable. The first group can be composed of any allowed transform block size, or can be composed of two or more block sizes among the allowed block sizes. The block size to which NSPT is applicable can be defined as a block size in which at least one of the width and height is less than or equal to a predetermined threshold. Alternatively, the block size to which NSPT is applicable can be defined as a block size in which the product of the width and height is less than or equal to a predetermined threshold. Alternatively, the block size to which NSPT is applicable can be defined as a block size in which the maximum value of the width and height is less than or equal to a predetermined threshold. The threshold can be an integer of 4, 8, 16, 32, 64, 128 or more.
[0236] The other of the two groups (hereinafter referred to as the second group) may refer to a set of block sizes to which NSPT is not applied. The above-mentioned separable primary transform may be applied to the block sizes belonging to the second group. In addition, the non-separable secondary transform may be applied to all or part of the block sizes belonging to the second group.
[0237] For example, when the size of the current block belongs to the first group, a backward NSPT may be applied to the (dequantized) transform coefficients of the current block. When the size of the current block belongs to the second group, a backward separable primary transform may be applied to the (dequantized) transform coefficients of the current block. Alternatively, when the size of the current block belongs to the second group, a backward non-separable secondary transform (e.g., a low-frequency non-separable transform LFNST) may be first applied to the (dequantized) transform coefficients of the current block, and a backward separable primary transform (e.g., DCT-2) may be applied to the transform coefficients obtained therefrom.
[0238] For example, as a set of block sizes to which NSPT is applicable, the first group may be defined as a set of 4×4, 4×8, 8×4, and 8×8. Alternatively, the first group may be defined as a set of 4×8, 8×4, and 8×8. Alternatively, the first group may be defined as a set of 4×8 and 8×4. Alternatively, the first group may be defined as a set of 4×4, 4×8, 4×16, 8×4, 8×8, and 16×4. Alternatively, the first group may be defined as a set of 4×8, 4×16, 8×4, 8×8, and 16×4. Alternatively, the first group may be defined as a set of 4×8, 4×16, 8×4, and 16×4. Alternatively, the first group may be defined as a set of 4×4, 4×8, 8×4, 8×8, 8×16, 16×8, and 16×16. Alternatively, the first group may be defined as a set of 4×4, 4×8, 8×4, 8×8, 8×16, and 16×8. Alternatively, the first group may be defined as a set of 4×8, 8×4, 8×8, 8×16, and 16×8. Alternatively, the first group may be defined as a set of 4×8, 8×4, 8×16, and 16×8. Alternatively, the first group may be defined as a set of 4×4, 4×8, 8×4, 8×8, 8×16, 16×8, 16×16, 16×32, 32×16, and 32×32. Alternatively, the first group may be defined as a set of 4×4, 4×8, 8×4, 8×8, 8×16, 16×8, 16×16, 16×32, and 32×16. Alternatively, the first group may be defined as a set of 4×8, 8×4, 8×8, 8×16, 16×8, 16×16, 16×32, and 32×16. Alternatively, the first group may be defined as a set of 4×8, 8×4, 8×16, 16×8, 16×16, 16×32, and 32×16.
[0239] As shown in the example, NSPT can be applied to M×N blocks and N×M blocks, which are non-square blocks. For example, NSPT can be applied to 4×8 blocks and 8×4 blocks. Alternatively, NSPT can be applied to 4×16 blocks and 16×4 blocks, or NSPT can be applied to 8×16 blocks and 16×8 blocks, or NSPT can be applied to 16×32 blocks and 32×16 blocks.
[0240] By applying NSPT to specific block sizes belonging to the first group, a more complex transform can be performed, and encoding performance can be improved. When applying forward LFNST, the transform coefficients of the main transform can be zeroed for the remaining area (i.e., the region of interest (ROI)) except for the area where LFNST is applied. In addition, LFNST can consist of a small number of transform basis vectors. In this case, performance degradation may occur when a separable main transform such as DCT-2 and an inseparable secondary transform such as LFNST are applied to the corresponding block size instead of NSPT. When NSPT is applied instead of LFNST for the corresponding case, the zeroing process is omitted, and encoding performance can be improved compared to the case where LFNST is applied. Furthermore, the method of applying NSPT can be expected to improve performance. NSPT or LFNST can be applied by leveraging the following symmetry. Here, for LFNST, the symmetry is used to perform a transposition operation on the corresponding input block only for the ROI region. On the other hand, for NSPT, the symmetry is used to perform a transposition operation on the entire block. Therefore, for NSPT, a more complex symmetry can be used to train and apply the corresponding NSPT kernel, thus expecting performance improvements.
[0241] In addition, when LFNST is applied to an 8×8 block instead of NSPT, from the perspective of forward transform, a 32×64 transform matrix can be applied instead of a 16×64 transform matrix. Here, the 16×64 transform matrix can be configured by sampling the top 16 rows in the 32×64 transform matrix. When LFNST based on the 16×64 transform matrix is applied to an 8×8 block, 16 multiplications are required per sample to apply LFNST, but when a 32×64 transform matrix is used, 32 multiplications are required per sample to apply LFNST. However, when a 32×64 transform matrix is used like this, improvements in coding performance can be expected.
[0242] When the tree type of the current block is single tree, NSPT may be applied to the luma component of the current block, but may not be applied to the chroma components of the current block. When the tree type of the current block is dual tree, NSPT may be applied to both the luma component and the chroma components of the current block.
[0243] Alternatively, regardless of whether the tree type of the current block is a single tree, NSPT may be applied to the luma component of the current block, but may not be applied to the chroma components of the current block. Alternatively, regardless of whether the tree type of the current block is a single tree, NSPT may be applied to both the luma component and the chroma components of the current block.
[0244] As an example, when the tree type of the current block is a single tree, NSPT is allowed for the luma component and the chroma component, and the size of the current block belongs to the first group, one NSPT index can be signaled, and the luma component and the chroma component of the current block can share the corresponding NSPT index. Here, the NSPT index can be an index for selecting any one transform kernel candidate for NSPT. When the sizes of the luma block and the chroma block of the current block belong to the first group, the transform kernel candidate selected by the same NSPT index can be applied to the luma component and the chroma component. When the tree type of the current block is a single tree and NSPT is applied only to the luma component, LFNST may not be applied to the chroma component of the current block and a separate transform may be applied. Alternatively, when the tree type of the current block is a single tree and NSPT is applied only to the luma component, LFNST may be applied to the chroma components of the current block.
[0245] A single tree may have a high correlation between the luma component and the chroma component. In this case, by applying the NSPT only to the luma component or by applying the transform kernel candidate selected by one NSPT index to both the luma and chroma components, unnecessary signaling can be reduced and compression efficiency can be improved. On the other hand, with a non-single tree, the luma component and the chroma component each have independent partitioning and coding structures. In this case, by signaling the NSPT index of each component, the characteristics of each component can be reflected and compression efficiency can be improved.
[0246] The NSPT core for NSPT can be derived based on at least one of the symmetry between intra prediction modes or the symmetry between block shapes. As an example, the NSPT core can be derived as an NSPT core corresponding to at least one of a mode symmetrical with the intra prediction mode of the current block or a block shape symmetrical with the block shape of the current block. Alternatively, the NSPT core can be derived based on an NSPT set including one or more NSPT core candidates, wherein the NSPT set can be derived as an NSPT set corresponding to at least one of a mode symmetrical with the intra prediction mode of the current block or a block shape symmetrical with the block shape of the current block. Any one of the one or more NSPT core candidates belonging to the NSPT set can be set as the NSPT core of the current block. To this end, an NSPT index specifying any one of the one or more NSPT core candidates belonging to the NSPT set can be used. The NSPT index can be signaled via the bitstream or can be derived based on the above-mentioned symmetry.
[0247] Symmetry may exist between at least two intra prediction modes among the intra prediction modes predefined in the decoding apparatus. Hereinafter, for convenience of description, the symmetry around the upper left diagonal mode (ie, mode 34) is described. Figure 5, there is symmetry between the directional modes. Except for the planar mode numbered 0 and the DC mode numbered 1, all modes have a predicted direction. Modes 2 to 66 may be referred to as normal directional modes (which may be expressed as [2, 66]), and modes -14 to -1 (which may be expressed as [-14, -1]) and modes 67 to 80 (which may be expressed as [67, 80]) may be referred to as wide directional modes. The wide directional mode may include at least one of a mode with a value less than -14 or a mode with a value greater than 80. Reference Figure 5 , all modes except mode 0 and mode 1 are symmetric around mode 34. Specifically, for mode [2, 66], mode x is symmetric with mode (68-x), and between mode [-14, -1] and mode [67, 80], mode x is symmetric with mode (66-x). The same symmetry relationship can be established between mode [N, -1] and mode [67, 66-N]. Here, N can be an integer less than or equal to -14.
[0248] In addition, regarding the symmetry between block shapes, M×N blocks and N×M blocks can be defined as blocks that are symmetrical to each other. Here, M and N can be the same as or different from each other. Alternatively, when the ratio of the width to the height of the M1×N1 block (M1 / N1) is the same as the ratio of the height to the width of the M2×N2 block (N2 / M2), the M1×N1 block and the M2×N2 block can be defined as blocks that are symmetrical to each other. Alternatively, when the ratio of the width to the height of the M1×N1 block (M1 / N1) is the same as the ratio of the height to the width of the M2×N2 block (M2 / N2), the M1×N1 block and the M2×N2 block can be defined as blocks that are symmetrical to each other.
[0249] In a square block, mutually symmetric patterns can share at least one of an NSPT set, an NSPT index, or an NSPT core. In other words, at least one of an NSPT set, an NSPT index, or an NSPT core of any symmetric pattern can be equally applied to another symmetric pattern.
[0250] As an example, patterns that are symmetric to each other can share one NSPT core. However, for any one symmetric pattern, the corresponding NSPT core can be applied to the input data, and for the other symmetric pattern, the corresponding NSPT core can be applied after applying a transpose operation to the input data. Specifically, when pattern x belongs to pattern [2,33], a 1D vector can be configured for an M×M block as input data in column-first order for pattern x, and the NSPT core can be applied to the corresponding 1D vector. Here, configuring a 1D vector according to column-first order can read input data in column units from an M×M block as input data to obtain M columns, and arrange them in sequence to configure a 1D vector. On the other hand, a 1D vector can be configured in row-first order for a pattern (68-x) that is symmetric to pattern x, and the corresponding same NSPT core can be applied to the corresponding 1D vector. Here, configuring a 1D vector according to row-first order can read input data in row units from an M×M block as input data to obtain M rows, and arrange them in sequence to configure a 1D vector. When mode x belongs to mode [N, -1] (N≤-14), a 1D vector can be configured in row-first order for mode (66-x) symmetric to mode x, and the same NSPT kernel as that of mode x can be applied to the corresponding 1D vector. Column-first order or row-first order can be applied to mode 0 and mode 1, and column-first order or row-first order can also be applied to mode 34. In addition, row-first order can be applied to intra-frame prediction modes belonging to mode [2, 33], and column-first order can be applied to modes symmetric to the corresponding intra-frame prediction modes. Row-first order can be applied to intra-frame prediction modes belonging to mode [N, -1], and column-first order can be applied to modes symmetric thereto.
[0251] For non-square blocks, in addition to the symmetry between intra prediction modes, the symmetry between block shapes can also be considered. Non-square blocks with width and height of M and N, respectively, can be considered to have a symmetric relationship with non-square blocks with width and height of N and M, respectively. As an example, in mode [2,66], there can be symmetry between mode x of an M×N block and mode (68-x) of an N×M block. Similarly, when mode x of an M×N block belongs to mode [N,-1] (N≤-14), there can be symmetry between mode x of an M×N block and mode (66-x) of an N×M block.
[0252] The method of configuring a 1D vector from an input data block is as described above. In other words, when column priority is applied to pattern x, row priority can be applied to a pattern symmetrical thereto. Alternatively, when row priority is applied to pattern x, column priority can be applied to a pattern symmetrical thereto. Specifically, when column priority is applied to pattern x, M columns can be obtained by reading input data in units of columns from an M×N block as input data, and they can be arranged in sequence to configure a 1D vector. Here, each column can have a length N. For a pattern symmetrical to pattern x, N rows can be obtained by reading input data in units of rows from an M×N block as input data, and they can be arranged in sequence to configure a 1D vector. Here, each row can have a length M. Alternatively, when row priority is applied to pattern x, N rows can be obtained by reading input data in units of rows from an M×N block as input data, and they can be arranged in sequence to configure a 1D vector. Here, each row can have a length M. For a pattern symmetrical to the pattern x, M columns can be obtained by reading input data in units of columns from an M×N block as input data, and they can be sequentially arranged to configure a 1D vector. Here, each column can have a length of N.
[0253] When the current block is an M×N block with mode x and the above-mentioned symmetry is used for the current block, the NSPT set and / or NSPT core of the current block can be determined based on at least one of an intra-frame prediction mode symmetric to mode x or an N×M block size symmetric to the M×N block size. Here, the NSPT core can be set to an NSPT core of an N×M block instead of an NSPT core of an M×N block. In other words, when symmetry is used for the current block, the NSPT set and / or NSPT core of a block having symmetry with the current block can be used in the same manner. As described above, a 1D vector can be configured from the input data block according to a predetermined priority that can correspond to the input of the NSPT core.
[0254] In addition, there may be a restriction that symmetry is used only when the value of the intra-frame prediction mode of the current block is greater than 34. In other words, when the value of the intra-frame prediction mode of the current block is greater than 34, a transposition operation may be applied when configuring a 1D vector from an input data block, and an NSPT set or NSPT kernel corresponding to a block shape and / or pattern that is symmetric with the current block may be used. Specifically, when the intra-frame prediction mode of the current block belongs to mode [N, -1] and mode [2, 34], symmetry may not be used for the current block. On the other hand, when the intra-frame prediction mode of the current block belongs to mode [35, 66] and mode [67, 66-N], symmetry may be used for the current block. Here, N may be an integer less than or equal to -14.
[0255] The derivation of the NSPT set or NSPT kernel based on symmetry may be adaptively performed based on the size of the current block. For example, for a 4×4 block and an 8×8 block, the NSPT set or NSPT kernel may be derivable based on symmetry, whereas for a 4×8 block and an 8×4 block, the NSPT set or NSPT kernel may not be derivable based on symmetry.
[0256] Depending on whether symmetry is used, the number of available NSPT sets may be different. As an example, when symmetry is used, the number of available NSPT sets may be 35, and when symmetry is not used, the number of available NSPT sets may be 67.
[0257] Table 9 below relates to an example of determining an NSPT set using symmetry, and shows a mapping relationship between NSPT sets and intra prediction modes when the number of available NSPT sets is 35.
[0258] [Table 9]
[0259] Intra prediction mode NSPT set index X<0 2 0≤X≤34 X 35≤X≤66 68-X X>66 2
[0260] Referring to Table 9, when the value (X) of the intra prediction mode of the current block is less than 0, the NSPT set of the current block may be determined as the NSPT set with an NSPT set index of 2 among the 35 NSPT sets. When the value (X) of the intra prediction mode of the current block is greater than or equal to 0 and less than or equal to 34, the NSPT set of the current block may be determined as the NSPT set with an NSPT set index of X among the 35 NSPT sets. When the value (X) of the intra prediction mode of the current block is greater than or equal to 35 and less than or equal to 66, the NSPT set of the current block may be determined as the NSPT set with an NSPT set index of (68-X) among the 35 NSPT sets. When the value (X) of the intra prediction mode of the current block is greater than or equal to 35 and less than or equal to 66, the NSPT set of the current block may be the same as the NSPT set with a value of (68-X) corresponding to a mode symmetric to the intra prediction mode of the current block. Similarly, when the value (X) of the intra prediction mode of the current block is greater than 66, the NSPT set of the current block may be determined as the NSPT set with an NSPT set index of 2 among the 35 NSPT sets. When the value (X) of the intra prediction mode of the current block is greater than 66, the NSPT set of the current block may be the same as the NSPT set corresponding to a mode symmetric to the intra prediction mode of the current block.
[0261] Table 10 below relates to an example of determining an NSPT set without using symmetry, and shows a mapping relationship between NSPT sets and intra prediction modes when the number of available NSPT sets is 67.
[0262] [Table 10]
[0263] Intra prediction mode NSPT set index X<0 2 0≤X≤66 X X>66 66
[0264] Referring to Table 10, when the value (X) of the intra prediction mode of the current block is less than 0, the NSPT set of the current block may be determined as the NSPT set with an NSPT set index of 2 among the 67 NSPT sets. When the value (X) of the intra prediction mode of the current block is greater than or equal to 0 and less than or equal to 66, the NSPT set of the current block may be determined as the NSPT set with an NSPT set index of X among the 67 NSPT sets. Similarly, when the value (X) of the intra prediction mode of the current block is greater than 66, the NSPT set of the current block may be determined as the NSPT set with an NSPT set index of 66 among the 67 NSPT sets.
[0265] Symmetry can be used to save the memory size required to store the transformation kernel while maintaining performance depending on the application of the transformation. For example, when symmetry is exploited to use 35 NSPT sets instead of 67 NSPT sets, the memory size required to store the NSPT kernel can be significantly reduced.
[0266] The number of available NSPT sets and / or the number of NSPT core candidates belonging to an NSPT set may vary depending on the block size. For example, the number of available NSPT sets for a 4×4 block may be 35, the number of available NSPT sets for a 4×8 block and an 8×4 block may be 19, and the number of available NSPT sets for an 8×8 block may be 10. The NSPT set for a 4×4 block may consist of three NSPT core candidates, the NSPT sets for a 4×8 block and an 8×4 block may consist of three or two NSPT core candidates, and the NSPT set for an 8×8 block may consist of one NSPT core candidate.
[0267] As block size increases, the size of the transform kernel can increase. Therefore, the number of available NSPT sets and / or the number of NSPT kernel candidates belonging to an NSPT set can be reduced to save the memory size required to store the transform kernel. In addition, as block size increases, the characteristics of the residual signal within the corresponding block tend to become more generalized. Therefore, reducing the number of available NSPT sets and / or the number of NSPT kernel candidates belonging to an NSPT set can help maintain compression efficiency while reducing implementation complexity by reflecting these statistical characteristics.
[0268] The NSPT kernel can be configured with 8-bit precision. The coefficient range within the NSPT kernel can be greater than or equal to -128 and less than or equal to 127. When the corresponding precision increases to more than 8 bits, the result value obtained by the matrix multiplication can be shifted right by the increased precision. For example, when the value obtained after matrix multiplication based on the NSPT kernel with 8-bit precision is shifted right by S bits and stored in the buffer, if the kernel coefficient is configured with N-bit precision, it can be shifted right by (S+(N-8)) bits and stored in the buffer.
[0269] When the NSPT core is configured with 8-bit precision, it can prevent excessive increase in internal precision in encoders / decoders performing the transform, thereby reducing implementation complexity in terms of memory requirements and number of operations while minimizing reduction in compression efficiency.
[0270] When backward NSPT is applied to a current block of size N×N, the size of the NSPT kernel (or NSPT matrix) can be expressed as MN×r. Here, MN may refer to the product of the width and height of the current block. This may refer to the output length of the NSPT or the number of residual samples generated by the NSPT. In addition, r may refer to the input length of the NSPT or the number of (dequantized) transform coefficients to which the NSPT is applied. r may be an integer greater than or equal to 0 and less than or equal to MN. The following is an example of an NSPT matrix of MN×r according to block size.
[0271] The NSPT matrix for a 4×4 block can consist of a 16×16 matrix. The NSPT matrix for a 4×8 block and an 8×4 block can consist of a 32×20 matrix, a 32×16 matrix, a 32×24 matrix, a 32×28 matrix, or a 32×32 matrix. The NSPT matrix for an 8×8 block can consist of a 64×16 matrix, a 64×24 matrix, a 64×32 matrix, a 64×40 matrix, a 64×48 matrix, a 64×56 matrix, or a 64×64 matrix. The NSPT matrix for a 4×16 block and a 16×4 block can consist of a 64×16 matrix, a 64×24 matrix, a 64×32 matrix, a 64×40 matrix, a 64×48 matrix, a 64×56 matrix, or a 64×64 matrix. The NSPT matrix for 8×16 and 16×8 blocks can be composed of a 128×96 matrix, a 128×64 matrix, a 128×48 matrix, or a 128×32 matrix. The NSPT matrix for 16×16 blocks can be composed of a 256×128 matrix, a 256×96 matrix, or a 256×64 matrix. The NSPT matrix for 16×32 and 32×16 blocks can be composed of a 512×256 matrix or a 512×128 matrix. The NSPT matrix for 32×32 blocks can be composed of a 1024×512 matrix, a 1024×256 matrix, or a 1024×128 matrix.
[0272] Alternatively, a 16×16 matrix can be applied to a 4×N block and an N×4 block. Here, N may be an integer greater than or equal to 4. A 64×16 matrix can be applied to an 8×8 block. A 64×32 matrix can be applied to an 8×N block and an N×8 block. Here, N may be an integer greater than or equal to 16. A 96×32 matrix can be applied to a 16×N block and an N×16 block. Here, N may be an integer greater than or equal to 16.
[0273] Alternatively, the value of r in the MN×r NSPT matrix can be determined according to predetermined criteria. The criteria here can be (1) ensuring that the sum of the computational complexity of the primary transform and the computational complexity of the secondary transform is less than or equal to a specific level, and (2) ensuring that the number of multiplications per sample required for the NSPT operation is less than or equal to a specific number.
[0274] Based on the inverse transform, when a separable primary transform is performed for an M×N block through matrix multiplication, (M+N) multiplications are required per sample to perform the corresponding primary transform. In addition, when LFNST is applied to a specific region of interest (ROI) area, (P*Q) / (M*N) multiplications are required per sample when assuming that the inverse transformed LFNST matrix is a P×Q matrix. Here, the P×Q matrix may refer to a matrix with P columns and Q rows.
[0275] When NSPT is applied to an M×N block instead of DCT-2 transform (or separable transform, such as KLT) and LFNST, the value of r that ensures that the number of per-sample multiplications when the corresponding NSPT is applied is less than or equal to the number of per-sample multiplications when the DCT-2 transform and LFNST are applied can be determined as follows.
[0276] [Formula 6]
[0277] r≤(M+N+(P*Q) / (M*N))
[0278] When the value of r is set to the maximum value (ie, r=M+N+(P*Q) / (M*N)) while satisfying the above Equation 6, the value of r in the NSPT matrix of each block size can be set as follows.
[0279] For the NSPT of a 4×4 block, the value of r is 24 (r=4+4+((16×16) / (4×4))), but the value of r must be less than or equal to 16, so the value of r can be set to 16.
[0280] For the NSPT of 4×8 blocks and 8×4 blocks, the value of r is 20 (r=4+8+((16×16) / (4×8))), so the value of r can be set to 20.
[0281] For the NSPT of an 8×8 block, the value of r is 32 (r=8+8+((64×16) / (8×8))), so the value of r may be set to 32.
[0282] For NSPT of 4×16 blocks and 16×4 blocks, the value of r is 24 (r=4+16+((16×16) / (4×16))), so the value of r may be set to 24.
[0283] For the NSPT of 8×16 blocks and 16×8 blocks, the value of r is 40 (r=8+16+((64×32) / (8×16))), so the value of r can be set to 40.
[0284] For the NSPT of a 16×16 block, the value of r is 44 (r=16+16+((96×32) / (16×16))), so the value of r may be set to 44.
[0285] For NSPTs of 16×32 blocks and 32×16 blocks, the value of r is 54 (r=16+32+((96×32) / (16×32))), so the value of r may be set to 54.
[0286] For the NSPT of a 32×32 block, the value of r is 67 (r=32+32+((96×32) / (32×32))), so the value of r can be set to 67.
[0287] The value of r according to the above-described predetermined standard does not consider zeroing. In other words, when forward LFNST is applied, the transform coefficients of the main transform in the remaining area other than the area where LFNST is applied are zeroed, so the actual amount of computation required for applying DCT-2 and LFNST can be less than the above-described amount of computation. Therefore, when zeroing is considered, the value of r can be set to a value smaller than the value of r according to the predetermined standard.
[0288] Since zeroing is not performed for a 4×4 block, the value of r may be set to a value less than or equal to 16.
[0289] For a 4×8 block, zeroing can be performed on the remaining area except for the upper left 4×4 block based on the forward transform, and a 16×16 matrix as the forward LFNST matrix can be applied to the upper left 4×4 block. When this zeroing is performed, the number of per-sample multiplications required in the forward separable main transform is 8 ((4×4×8)+(4×8×4)) / (4×8)=8), and the number of per-sample multiplications required in the LFNST is 8 ((16×16) / 32=8). Therefore, when the separable main transform and LFNST are replaced with NSPT, the value of r can be set to a value less than or equal to 16, which is the sum of the number of per-sample multiplications in the separable main transform and the number of per-sample multiplications in the LFNST. Since the same amount of calculation is required even when the backward separable main transform and LFNST are applied, the value of r can be set to a value less than or equal to 16.
[0290] For an 8×4 block, zeroing can be performed on the remaining area except for the upper left 4×4 block based on the forward transform, and a 16×16 matrix as the forward LFNST matrix can be applied to the upper left 4×4 block. When this zeroing is performed, the number of per-sample multiplications required in the forward separable main transform is 6 ((4×8×4)+(4×4×4) / (8×4)=6), and the number of per-sample multiplications required in the LFNST is 8 ((16×16) / 32=8). Therefore, when the separable main transform and LFNST are replaced with NSPT, the value of r can be set to a value less than or equal to 14, which is the sum of the number of per-sample multiplications in the separable main transform and the number of per-sample multiplications in the LFNST. Since the same amount of calculation is required even when the backward separable main transform and LFNST are applied, the value of r can be set to a value less than or equal to 14.
[0291] For 8x8 blocks, zeroing may not be performed for the separable primary transform, and in this case, the value of r may be set to a value less than or equal to 32.
[0292] For an 8×16 block, zeroing can be performed on the remaining area except for the upper left 8×8 block based on the forward transform, and a 64×32 matrix as the forward LFNST matrix can be applied to the upper left 8×8 block. When this zeroing is performed, the number of per-sample multiplications required in the forward separable main transform is 16 ((8×8×16)+(8×16×8) / (8×16)=16), and the number of per-sample multiplications required in the LFNST is 16 ((64×32) / 128=16). Therefore, when the separable main transform and LFNST are replaced with NSPT, the value of r can be set to a value less than or equal to 32, which is the sum of the number of per-sample multiplications in the separable main transform and the number of per-sample multiplications in the LFNST. Since the same amount of calculation is required even when the backward separable main transform and LFNST are applied, the value of r can be set to a value less than or equal to 32.
[0293] For a 16×8 block, zeroing can be performed on the remaining area except for the upper left 8×8 block based on the forward transform, and a 64×32 matrix as the forward LFNST matrix can be applied to the upper left 8×8 block. When this zeroing is performed, the number of per-sample multiplications required in the forward separable main transform is 12 ((8×16×8)+(8×8×8) / (16×8)=12), and the number of per-sample multiplications required in the LFNST is 16 ((64×32) / 128=16). Therefore, when the separable main transform and LFNST are replaced with NSPT, the value of r can be set to a value less than or equal to 28, which is the sum of the number of per-sample multiplications in the separable main transform and the number of per-sample multiplications in the LFNST. Since the same amount of calculation is required even when the backward separable main transform and LFNST are applied, the value of r can be set to a value less than or equal to 28.
[0294] For a 16×16 block, zeroing can be performed on the remaining area except for the upper left 12×12 block based on the forward transform, and a 96×32 matrix as the forward LFNST matrix can be applied to the upper left 12×12 block. When this zeroing is performed, the number of per-sample multiplications required in the forward separable main transform is 21 ((12×16×16)+(12×16×12) / (16×16)=21), and the number of per-sample multiplications required in the LFNST is 12 ((96×32) / 256=12). Therefore, when the separable main transform and LFNST are replaced with NSPT, the value of r can be set to a value less than or equal to 33, which is the sum of the number of per-sample multiplications in the separable main transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable main transform and LFNST, the value of r can be set to a value less than or equal to 33.
[0295] As described above, the r value in the NSPT matrix for an M×N block may differ from the r value in the NSPT matrix for an N×M block. As an example, the backward NSPT matrix for a 4×8 block may be a 32×16 matrix, and the backward NSPT matrix for an 8×4 block may be a 32×14 matrix. In this case, the symmetry between the M×N block and the N×M block can be used to determine the NSPT matrix.
[0296] Assume that the current block is an M×N block with pattern x. When the above symmetry is applied to the current block, instead of applying the NSPT matrix corresponding to the M×N block size or pattern x, an NSPT matrix corresponding to at least one of a pattern symmetric to pattern x or an N×M block size symmetric to the M×N block size can be applied. In this case, the NSPT matrix corresponding to the N×M block size can be applied to the current block as is. Alternatively, the NSPT matrix corresponding to the N×M block size is applied, but for the value of r, the r value in the NSPT matrix corresponding to the M×N block size can be used.
[0297] As an example, the backward NSPT matrix of a 4×8 block can be a 32×16 matrix (i.e., the r value in the NSPT matrix is 16), and the backward NSPT matrix of an 8×4 block can be a 32×14 matrix (i.e., the r value in the NSPT matrix is 14). When the current block is an 8×4 block with a pattern x, an NSPT matrix of at least one of a pattern symmetric to the pattern x or a 4×8 block symmetric to the 8×4 block can be applied to the current block. In this case, as the backward NSPT matrix of the 4×8 block, the 32×16 matrix can be used as is, or a 32×14 matrix with the r value in the backward NSPT matrix of the 8×4 block can be used. Here, the 32×14 matrix can be derived by sampling 14 rows from the left in the 32×16 matrix. In this way, when the 32×14 matrix is applied to the current block with a block size of 8×4, the above-mentioned predetermined criteria are met.
[0298] Conversely, when the current block is a 4×8 block with mode x, an NSPT matrix of at least one of a mode symmetric to mode x or an 8×4 block symmetric to a 4×8 block can be applied to the current block. In this case, a 32×14 matrix for an 8×4 block can be applied to the current block instead of a 32×16 matrix for a 4×8 block. Thus, NSPT can be performed using fewer multiplications than allowed for a 4×8 block.
[0299] For the NSPT of M×N blocks and N×M blocks, when the r values satisfying the above predetermined conditions are r1 and r2 respectively, the backward NSPT matrices of the M×N blocks and N×M blocks can be set to MN×max(r1, r2).
[0300] As an example, the backward NSPT matrix of a 4×8 block can be a 32×16 matrix (i.e., the r value in the NSPT matrix is 16), and the backward NSPT matrix of an 8×4 block can be a 32×14 matrix (i.e., the r value in the NSPT matrix is 14). When the current block is a 4×8 block with a pattern x, an NSPT matrix of at least one of a pattern symmetric to the pattern x or an 8×4 block symmetric to the 4×8 block can be applied to the current block. In this case, a 32×16 matrix as the backward NSPT matrix of the 8×4 block can be used. When the backward NSPT matrix is not configured like MN×max(r1,r2), the 32×14 matrix will be used as the NSPT matrix of the 8×4 block. However, when the NSPT matrix of the 4×8 block and the NSPT matrix of the 8×4 block are configured as a 32×max(16,14) matrix, the 32×16 matrix can be fully applied.
[0301] Conversely, when the current block is an 8×4 block with mode x, an NSPT matrix of at least one of a mode symmetric to mode x or a 4×8 block symmetric to the 8×4 block can be applied to the current block. In this case, a 32×16 matrix can be used as the backward NSPT matrix for the 4×8 block, or a 32×14 matrix can be used. Here, the 32×14 matrix can be derived by sampling 14 rows from the left of the 32×16 matrix.
[0302] When the NSPT matrix is configured as described above, it is possible to apply a transform consisting of a maximum number of transform basis vectors while satisfying predetermined conditions, thereby maximizing encoding performance.
[0303] In the above embodiment, the value of r can be set to a multiple of 16. As an example, for the backward NSPT of 4×8 blocks and 8×4 blocks, a 32×16 matrix can be applied instead of a 32×20 matrix. The transform coefficients of the transform block can be encoded in units of predetermined coefficient groups (CGs). Here, CG can be defined as a group of 16 transform coefficients, and as an example, the CG can be a sub-block of a size such as 4×4, 2×8 or 8×2. There may be no non-zero transform coefficient within a CG, and in this case, the encoding process of the transform coefficient can be skipped for the corresponding CG. Therefore, when the value of r is set to a multiple of 16, there is an advantage of reducing the implementation complexity.
[0304] The transform coefficients may be derived by applying the forward NSPT to an M×N block of residual samples. In this case, due to zeroing, the number of derived transform coefficients may be less than or equal to the value (M*N). In other words, the forward NSPT matrix may be defined as an r×(M*N) matrix, where r may refer to the output length of the NSPT or the number of transform coefficients derived by the NSPT, and (M*N) may refer to the input length of the NSPT or the number of residual samples to which the NSPT is applied.
[0305] The derived transform coefficients may be arranged in an M×N block according to a predetermined scanning order, and regions not filled with transform coefficients may be filled with 0s (i.e., set to zero). Therefore, in a process of scanning transform coefficients in a decoding apparatus, when a non-zero transform coefficient is found in a region that should be filled with 0s if NSPT is applied (or when the scanning position of the last significant coefficient in an M×N block is greater than or equal to r), it is considered that NSPT is not applied to the corresponding M×N block, and the NSPT index may not be signaled.
[0306] One or more r-values may be defined for block sizes to which NSPT is applicable. For example, one or more r-values may be defined for each block size to which NSPT is applicable. Alternatively, one r-value may be defined for each block size to which NSPT is applicable, and the r-value for any block size to which NSPT is applicable may be different from the r-value for another block size. Alternatively, one r-value may be defined for some block sizes to which NSPT is applicable, and at least two r-values may be defined for the remaining block sizes.
[0307] When multiple r values are available, an index specifying any of the multiple r values or r itself can be signaled. The corresponding index can be signaled in a high-level syntax (HLS) such as VPS, SPS, PPS, PH, or SH, or can be signaled at the block level such as CTU, CU, or TU. When the value of r is a value within a specific range, bits sufficient to cover the corresponding range can be allocated and signaled. For example, when the value of r is within the range of 1 to 256, 8 bits can be designated as a fixed length and signaled.
[0308] The transform core of the current block may be determined based on any one of the above-described embodiments 1 to 3. Alternatively, the transform core of the current block may be determined based on a combination of at least two of the embodiments 1 to 3 within the scope in which the inventions according to the above-described embodiments 1 to 3 do not conflict with each other.
[0309] A transform index for inverse transform of the current block may be signaled. Here, the transform index may specify any one of one or more transform kernels (or transform matrices) belonging to a transform set. Here, the transform index may refer to an NSPT index that specifies any one of one or more NSPT kernels belonging to an NSPT set. Alternatively, the transform index may refer to an LFNST index that specifies any one of one or more LFNST kernels belonging to an LFNST set.
[0310] Whether a transform index corresponds to an NSPT index can be determined based on whether the current block size is any of the block sizes belonging to the first group described above. Assume that the block sizes applicable to NSPT and LFNST are different. In this case, when the current block size belongs to the first group, the transform index signaled for the current block may correspond to an NSPT index, and the NSPT core may be determined from the NSPT set based on the corresponding transform index. On the other hand, when the current block size does not belong to the first group, the transform index signaled for the current block may correspond to an LFNST index, and the LFNST core may be determined from the LFNST set based on the corresponding transform index. When the current block size does not belong to the first group, this may mean that the current block size belongs to the second group described above. Alternatively, when the current block size does not belong to the first group, this may mean that the current block size corresponds to a block size applicable to LFNST among the block sizes belonging to the second group. In this way, the NSPT index and LFNST index can be configured as an integrated syntax rather than as separate syntaxes.
[0311] As an example, assume that the block sizes for which NSPT is applicable, belonging to the first group, are 4×4, 4×8, 8×4, and 8×8. For the block sizes belonging to the first group, NSPT can be applied instead of LFNST. Specifically, NSPT can be applied instead of a combination of a separable primary transform (e.g., DCT-2, separable KLT) and LFNST. NSPT indices can be signaled for the four block sizes belonging to the first group, and LFNST indices can be signaled for the remaining block sizes (for which LFNST is permitted).
[0312] In this way, when the NSPT index and the LFNST index are signaled as one syntax, the amount of information to be encoded can be reduced. In addition, the implementation complexity can be reduced by equally applying at least one of binarization, CABAC context, or initial value of entropy coding to the NSPT / LFNST index.
[0313] Alternatively, the NSPT index and the LFNST index may be signaled separately as separate syntaxes. In this case, the implementation complexity may partially increase, but the compression performance may be improved by performing optimized entropy coding for each index.
[0314] When the number of LFNST core candidates belonging to the LFNST set and the number of NSPT core candidates belonging to the NSPT set are the same, the same binarization can be applied to the LFNST index and the NSPT index. The same CABAC context (or CABAC context increment) can be assigned to the bins of the LFNST index and the NSPT index.
[0315] Different binarization and / or CABAC contexts can be used for LFNST indexes and NSPT indexes. Different CABAC initial values can be assigned to LFNST indexes and NSPT indexes. As an example, either LFNST indexes and NSPT indexes can be binarized based on fixed-length binarization, and the other can be binarized based on truncated unary binarization. Even when the binarization of the LFNST index and the NSPT index is the same, different CABAC contexts and / or CABAC initial values can be assigned. When the number of LFNST core candidates belonging to the LFNST set and the number of NSPT core candidates belonging to the NSPT set are different from each other, different binarization and / or CABAC contexts can be used for the LFNST index and the NSPT index.
[0316] The number of NSPT core candidates belonging to the NSPT set may be set differently for each block size. Alternatively, the block sizes belonging to the first group may be divided into a plurality of subgroups. In this case, the number of NSPT core candidates belonging to the NSPT set may be set differently for each of the plurality of subgroups. Here, at least one of the plurality of subgroups may include a plurality of different block sizes.
[0317] Depending on the number of NSPT core candidates belonging to the NSPT set, the binarization applied to the NSPT index may be different.
[0318] As an example, when the number of NSPT core candidates in the NSPT set of a specific block size is 3, the NSPT index can have any value from 0 to 3. When the value of the NSPT index is 0, it can indicate that NSPT is not applied to the current block. When the value of the NSPT index is not 0, it can indicate the NSPT core candidate corresponding to the corresponding NSPT index among the three NSPT core candidates. A bin can be allocated to distinguish between the case where NSPT is applied and the case where NSPT is not applied. The case where the value of the corresponding bin is 0 can correspond to the case where the value of the NSPT index is 0. On the other hand, the case where the value of the corresponding bin is 1 can correspond to the case where the value of the NSPT index is 1, 2 or 3. In this case, truncated unary binarization can be applied to distinguish the three NSPT core candidates. In other words, two bins can be allocated to distinguish the three NSPT core candidates as 0, 10 and 11.
[0319] When the number of NSPT core candidates in the NSPT set of a specific block size is 2, the NSPT index can have any value from 0 to 2. When the value of the NSPT index is 0, it can indicate that NSPT is not applied to the current block. When the value of the NSPT index is not 0, it can indicate the NSPT core candidate corresponding to the corresponding NSPT index among the two NSPT core candidates. A bin can be allocated to distinguish between the case where NSPT is applied and the case where NSPT is not applied. The two NSPT core candidates can be distinguished by allocating a bin representing either of the two NSPT core candidates.
[0320] When the number of NSPT core candidates in the NSPT set for a specific block size is 1, the NSPT index can have either a value of 0 or 1. When the value of the NSPT index is 0, it can indicate that NSPT is not applied to the current block. When the value of the NSPT index is 1, it can indicate one NSPT core candidate. In this case, whether to apply NSPT and NSPT core candidates can be specified using only one bin.
[0321] The inverse transform of the current block can be a separable main transform and / or an LFNST-based inverse transform. In other words, backward LFNST can be applied to all or part of the (dequantized) transform coefficients of the current block, and then the backward separable main transform can be applied to the transform coefficients derived by LFNST to derive residual samples. As an example, backward LFNST can be applied to the (dequantized) transform coefficients belonging to a partial area of the current block. Here, the partial area refers to the area where the forward LFNST is applied, which is hereinafter referred to as the region of interest (ROI) area. The transform coefficients derived by LFNST can be arranged in a predetermined scanning order in the ROI area. The predetermined scanning order can be a row-first order or a column-first order. The backward separable main transform can be applied to the transform coefficients derived by LFNST and the transform coefficients belonging to the remaining area of the current block other than the ROI area. Alternatively, in the forward transform process, when zeroing is performed on the remaining area except the ROI area within the current block (i.e., when the transform coefficients within the remaining area are set to 0), the backward separable main transform can be applied to the transform coefficients derived through LFNST.
[0322] LFNST defined for 4×8 blocks (hereinafter referred to as LFNST4×8) or LFNST defined for 8×4 blocks (hereinafter referred to as LFNST8×4) can be applied to the current block. When LFNST4×8 is applied to the current block, the ROI region of the current block can be a 4×8 region within the current block. Here, the 4×8 region is the region that includes the upper left sample of the current block, which can mean a block with a width and height of 4 and 8, respectively. When LFNST8×4 is applied to the current block, the ROI region of the current block can be an 8×4 region within the current block. Here, the 8×4 region is the region that includes the upper left sample of the current block, which can mean a block with a width and height of 8 and 4, respectively.
[0323] LFNST4×8 can be applied only when the current block is a 4×8 block. Alternatively, LFNST4×8 can be applied when the current block is 4×N and N is greater than or equal to 8 (for example, when the current block is a 4×8, 4×16, 4×32, or larger block). Alternatively, LFNST4×8 can be applied when the current block is 4×N and N is greater than or equal to 8 and less than or equal to a predetermined threshold. As an example, when the threshold is 16, LFNST4×8 can be applied when the current block is a 4×8 or 4×16 block.
[0324] LFNST 8×4 can be applied only when the current block is an 8×4 block. Alternatively, LFNST 8×4 can be applied when the current block is N×4 and N is greater than or equal to 8 (for example, when the current block is an 8×4, 16×4, 32×4, or larger block). Alternatively, LFNST 8×4 can be applied when the current block is N×4 and N is greater than or equal to 8 and less than or equal to a predetermined threshold. As an example, when the threshold is 16, LFNST 8×4 can be applied when the current block is an 8×4 or 16×4 block.
[0325] Specifically, when LFNST 4×8 or LFNST 8×4 is applied to the current block, an LFNST with an input length less than or equal to 32 and an output length of 32 (hereinafter referred to as LFNST-32) may be used. The input length may refer to the number of transform coefficients input to the LFNST. Here, the transform coefficients input to the LFNST may refer to all transform coefficients derived based on the residual information of the current block, or may refer to transform coefficients to which forward LFNST is applied among the derived transform coefficients. The output length may refer to the number of transform coefficients derived by the LFNST.
[0326] For example, when the size of the current block is 4×N, LFNST-32 may be applied to all or part of the transform coefficients belonging to a 4×8 region serving as the ROI region within the current block. The transform coefficients derived through LFNST-32 may be arranged in a predetermined scanning order within the 4×8 region within the current block. Alternatively, when the size of the current block is N×4, LFNST-32 may be applied to all or part of the transform coefficients belonging to an 8×4 region serving as the ROI region within the current block. The transform coefficients derived through LFNST-32 may be arranged in a predetermined scanning order within the 8×4 region within the current block.
[0327] However, as described above, LFNST4×8 and LFNST8×4 can be adaptively applied based on the size of the current block. When LFNST4×8 and LFNST8×4 are not applied to the current block, LFNST with an input length less than or equal to 16 and an output length of 16 (hereinafter referred to as LFNST16) can be used. LFNST16 can be applied to all or part of the transform coefficients belonging to a 4×4 region within the current block. The 4×4 region within the current block is the region including the upper left sample of the current block, which may mean a block with a width and height of 4 and 4, respectively. The transform coefficients derived by LFNST16 can be arranged in a predetermined scanning order in the 4×4 region within the current block.
[0328] From the perspective of forward transform, transform coefficients can be derived by applying a separate main transform to the residual samples of the current block. Forward LFNST can be applied to all derived transform coefficients or only to the transform coefficients belonging to a partial region (i.e., the ROI region) within the current block. The transform coefficients derived by LFNST can be arranged in the ROI region within the current block according to a predetermined scanning order. When the number of transform coefficients output from the LFNST is less than the number of transform coefficients input to the LFNST, there may be areas within the ROI region that are not filled with transform coefficients derived by LFNST. The corresponding areas can be filled with transform coefficients of 0. In addition, the remaining areas in the current block other than the ROI region (or areas to which LFNST is not applied) can be filled with transform coefficients derived by a separate main transform, or can be filled with transform coefficients set to 0.
[0329] Specifically, when LFNST 4×8 or LFNST 8×4 is applied to the current block or the ROI region of the current block, an LFNST with an input length of 32 and an output length less than or equal to 32 (hereinafter referred to as LFNST-32) may be used. Here, the input length may refer to the number of transform coefficients input to the LFNST. The input transform coefficients may be coefficients derived through a separable main transform. The input length may refer to the number of transform coefficients belonging to the ROI region of the current block. The output length may refer to the number of transform coefficients derived through the LFNST.
[0330] For example, when the size of the current block is 4×N, LFNST-32 may be applied to the transform coefficients belonging to a 4×8 region, which is the ROI region within the current block. The transform coefficients derived by LFNST-32 may be arranged in a predetermined scanning order within the 4×8 region within the current block. Alternatively, when the size of the current block is N×4, LFNST-32 may be applied to the transform coefficients belonging to an 8×4 region, which is the ROI region within the current block. The transform coefficients derived by LFNST-32 may be arranged in a predetermined scanning order within the 8×4 region within the current block.
[0331] However, as described above, LFNST4×8 and LFNST8×4 can be adaptively applied based on the size of the current block. When LFNST4×8 and LFNST8×4 are not applied to the current block, LFNST with an input length of 16 and an output length of 16, respectively, can be used (hereinafter referred to as LFNST16). LFNST16 can be applied to transform coefficients belonging to a 4×4 region within the current block. The 4×4 region within the current block is the region including the upper left sample of the current block, which can mean a block with a width and height of 4 and 4, respectively. The transform coefficients derived by LFNST16 can be arranged in a predetermined scanning order within the 4×4 region within the current block.
[0332] When applying LFNST 4×8 or LFNST 8×4 to the current block, the symmetry between the above block shapes can be used. As an example, the LFNST set and / or LFNST kernel for LFNST of the current block can be determined based on the symmetry between the 4×N block and the N×4 block. The 4×N block (or N×4 block) can use the LNFST set and / or LFNST kernel corresponding to the N×4 block (or 4×N block) with which it has symmetry. The above symmetry can be used for blocks with N of 8 and can be used as is, and the symmetry described above for the ROI region within the current block (i.e., 4×8 region or 8×4 region) can be used for blocks with N greater than or equal to 16.
[0333] In addition, LFNST for 8×16 blocks (hereinafter referred to as LFNST8×16) and LFNST for 16×8 blocks (hereinafter referred to as LFNST16×8) can be defined. When LFNST8×16 is applied to the current block, the ROI region of the current block can be a 4×8, 4×16, 8×8, or 8×16 region within the current block. Here, the ROI region is the region that includes the top-left sample of the current block, which can refer to a block with corresponding width and height. When LFNST16×8 is applied to the current block, the ROI region of the current block can be an 8×4, 16×4, 8×8, or 16×8 region within the current block. Here, the ROI region is the region that includes the top-left sample of the current block, which can refer to a block with corresponding width and height. The ROI region can be preset identically for both LFNST8×16 and LFNST16×8 for encoding and decoding devices. The same ROI region can be set for blocks with a width of 8 and a height of N, respectively. Alternatively, at least one block having a width of 8 and a height of N may have a different ROI region from another block. Similarly, the same ROI region may be set for blocks having a width of N and a height of 8. Alternatively, at least one block having a width of N and a height of 8 may have a different ROI region from another block.
[0334] LFNST 8×16 can be applied only when the current block is an 8×16 block. Alternatively, LFNST 8×16 can be applied when the current block is 8×N and N is greater than or equal to 16 (for example, when the current block is an 8×16, 8×32, 8×64, or larger block). Alternatively, LFNST 8×16 can be applied when the current block is 8×N and N is greater than or equal to 16 and less than or equal to a predetermined threshold. As an example, when the threshold is 32, LFNST 8×16 can be applied when the current block is an 8×16 or 8×32 block.
[0335] LFNST16×8 can be applied only when the current block is a 16×8 block. Alternatively, LFNST16×8 can be applied when the current block is N×8 and N is greater than or equal to 16 (for example, when the current block is a 16×8, 32×8, 64×8, or larger block). Alternatively, LFNST16×8 can be applied when the current block is N×8 and N is greater than or equal to 16 and less than or equal to a predetermined threshold. As an example, when the threshold is 32, LFNST16×8 can be applied when the current block is a 16×8 or 32×8 block.
[0336] Specifically, when LFNST 8×16 or LFNST 16×8 is applied to the current block, an LFNST having an input length less than or equal to K and an output length of K (hereinafter referred to as LFNST-K) may be used. As described above, the input length may refer to the number of transform coefficients input to the LFNST. Here, the transform coefficients input to the LFNST may refer to all transform coefficients derived based on the residual information of the current block, or may refer to transform coefficients to which forward LFNST is applied among the derived transform coefficients. The output length may refer to the number of transform coefficients derived by the LFNST.
[0337] K can be determined based on the size of the ROI region of the current block. As an example, when the ROI region is 4×8 or 8×4, an LFNST with an input length less than or equal to 32 and an output length of 32 (i.e., LFNST-32) can be used. When the ROI region is 4×16, 16×4, or 8×8, an LFNST with an input length less than or equal to 64 and an output length of 64 (i.e., LFNST-64) can be used. When the ROI region is 8×16 or 16×8, an LFNST with an input length less than or equal to 128 and an output length of 128 (i.e., LFNST-128) can be used. In other words, when LFNST 8×16 or LFNST 16×8 is applied to an 8×N or N×8 current block, LFNST-K can be applied to the ROI region within the current block. Here, K can be 32, 64, or 128, but this is only an example, and K can be an integer greater than 128.
[0338] When the size of the current block is 8×N, LFNST-K may be applied to all or part of the transform coefficients belonging to the ROI region within the current block. The transform coefficients derived by LFNST-K may be arranged in a predetermined scanning order within the ROI region within the current block. Alternatively, when the size of the current block is N×8, LFNST-K may be applied to all or part of the transform coefficients belonging to the ROI region within the current block. The transform coefficients derived by LFNST-K may be arranged in a predetermined scanning order within the ROI region within the current block.
[0339] However, as described above, LFNST8×16 and LFNST16×8 can be adaptively applied based on the size of the current block. When LFNST8×16 and LFNST16×8 are not applied to the current block, LFNST with an input length of 16 and an output length of 48 (hereinafter referred to as LFNST48) can be used. LFNST48 can be applied to the transform coefficients belonging to the 4×4 region within the current block. The 4×4 region within the current block is a region including the upper left sample of the current block, which may mean a block with a width and height of 4 and 4, respectively. The transform coefficients derived by LFNST48 can be arranged in a predetermined scanning order in three 4×4 regions within the current block. The three 4×4 regions can be composed of a first 4×4 region including the upper left sample of the current block, a second 4×4 region adjacent to the right side of the first 4×4 region, and a third 4×4 region adjacent to the bottom of the first 4×4 region.
[0340] From the perspective of forward transform, transform coefficients can be derived by applying a separate main transform to the residual samples of the current block. Forward LFNST can be applied to all derived transform coefficients or only to the transform coefficients belonging to a partial region (i.e., the ROI region) within the current block. The transform coefficients derived by LFNST can be arranged in the ROI region within the current block according to a predetermined scanning order. When the number of transform coefficients output from the LFNST is less than the number of transform coefficients input to the LFNST, there may be areas within the ROI region that are not filled with transform coefficients derived by LFNST. The corresponding areas can be filled with transform coefficients of 0. In addition, the remaining areas in the current block other than the ROI region (or areas to which LFNST is not applied) can be filled with transform coefficients derived by a separate main transform, or can be filled with transform coefficients set to 0.
[0341] Specifically, when LFNST 8×16 or LFNST 16×8 is applied to the current block or the ROI region of the current block, an LFNST with an input length of K and an output length less than or equal to K (i.e., referred to as LFNST-K) may be used. Here, the input length may refer to the number of transform coefficients input to the LFNST. The input transform coefficients may be coefficients derived through a separable main transform. The input length may refer to the number of transform coefficients belonging to the ROI region of the current block. The output length may refer to the number of transform coefficients derived through the LFNST.
[0342] As an example, when the size of the current block is 8×N and N is greater than or equal to 16, LFNST-K may be applied to the transform coefficients of the ROI region within the current block. The transform coefficients derived by LFNST-K may be arranged in the ROI region within the current block according to a predetermined scanning order. Alternatively, when the size of the current block is N×8 and N is greater than or equal to 16, LFNST-K may be applied to the transform coefficients of the ROI region within the current block. The transform coefficients derived by LFNST-K may be arranged in the ROI region within the current block according to a predetermined scanning order.
[0343] However, as described above, LFNST8×16 and LFNST16×8 can be adaptively applied based on the size of the current block. When LFNST8×16 and LFNST16×8 are not applied to the current block, LFNST with an input length of 48 and an output length of 16, respectively, can be used (i.e., referred to as LFNST48). LFNST48 can be applied to the transform coefficients of three 4×4 regions within the current block. These three 4×4 regions can consist of a first 4×4 region including the upper left sample of the current block, a second 4×4 region adjacent to the right of the first 4×4 region, and a third 4×4 region adjacent to the bottom of the first 4×4 region. The transform coefficients derived by LFNST48 can be arranged in a predetermined scanning order in the 4×4 regions within the current block.
[0344] When applying LFNST8×16 or LFNST16×8 to the current block, the symmetry between the above block shapes can be used. As an example, the LFNST set and / or LFNST kernel of LFNST-K for the current block can be determined based on the symmetry between the 8×N block and the N×8 block. The 8×N block (or N×8 block) can use the LNFST set and / or LFNST kernel corresponding to the N×8 block (or 8×N block) with which it has symmetry. The above symmetry can be used for blocks with N=16 and can be used as is, and the symmetry described above for the ROI region within the current block (i.e., 8×16 region or 16×8 region) can be used for blocks with N greater than or equal to 32.
[0345] As described above, LFNSTM×2M for M×2M blocks and LFNST2M×M for 2M×M blocks can be defined. LFNSTM×2M can be applied to M×N blocks, and LFNST2M×M can be applied to N×M blocks. Here, N can be greater than or equal to (2*M). ROI regions can be defined for LFNSTM×2M and LFNST2M×M, respectively, and LFNST with predetermined input lengths and output lengths can be defined / used. The size of the ROI region can be determined based on the minimum value of the width and height of the current block (i.e., M). The size of the ROI region can vary depending on the value of M. At least one of the input length and output length of LFNST can be determined based on at least one of the minimum value of the width and height of the current block or the size of the ROI region.
[0346] Hereinafter, a method for signaling the transform index of the inverse transform of the current block will be described. The inverse transform here may refer to a backward inseparable transform. The inseparable transform may refer to the aforementioned LFNST or NSPT, and the transform index may refer to an LFNST index or an NSPT index.
[0347] When the current block is encoded as a single tree, a non-separable transform may be applied to the luma component of the current block, but not to the chroma components of the current block. In this case, a transform index for the current block may be signaled based on at least one of a first condition for the luma component and a second condition for the chroma components. For example, when both the first condition for the luma component and the second condition for the chroma components are satisfied, the transform index for the current block may be signaled; otherwise (i.e., when either the first condition for the luma component or the second condition for the chroma components is not satisfied), no signaling may be performed. Alternatively, when the first condition for the luma component is satisfied, the transform index for the current block may be signaled without checking whether the second condition for the chroma components is satisfied; otherwise, no signaling may be performed. When no transform index is signaled, the non-separable transform may be set to not be applied to the current block. For example, the transform index for the current block may be derived as 0.
[0348] The first condition of the luminance component according to the present disclosure may mean that there are no non-zero transform coefficients in a predetermined area within the luminance component block of the current block. Here, the predetermined area within the luminance component block can be defined as the remaining area in the luminance component block except the first area, which can be referred to as the second area to distinguish it from the first area. The first area can be defined as an area consisting of the same number of samples (or sample positions) as the number of transform coefficients input to the backward inseparable transform (or the number of transform coefficients output by the forward inseparable transform). The first area may include the upper left sample position of the luminance component block. The first area may include at least one non-zero transform coefficient. The first area may be the area to which the last significant coefficient in the luminance component block belongs. The width and height of the first area may be less than or equal to the width and height of the luminance component block, respectively.
[0349] As described above, the number of transform coefficients to which the backward inseparable transform is applied may be determined based on the size of the current block. The size of the current block may be the same as that of the luminance component block. The size of the current block may be defined as a combination of width (W) and height (H), for example, W×H. However, without limitation thereto, the size of the current block may be defined as any one of width or height, a minimum / maximum value of width and height, or a product of width and height.
[0350] As an example, in an encoding device, when an inseparable transform is applied to an M×N block, r transform coefficients less than or equal to (M*N) may be output. This may mean that the forward inseparable transform has an output length of r. The r output transform coefficients may be arranged sequentially within the M×N block from the upper left sample position (i.e., the DC position) of the M×N block according to a predetermined scanning order. The region within the current block consisting of r transform coefficients may correspond to the above-mentioned first region. Then, zeroing may be applied to the sample positions after the rth within the M×N block (i.e., the second region within the current block). Zeroing may refer to assigning 0 to the sample positions within the second region. In this way, when an inseparable transform is applied in the encoding device, there are no non-zero transform coefficients for the sample positions to which zeroing is applied.
[0351] In the process of deriving transform coefficients based on residual information in the decoding device, if a non-zero transform coefficient is found in a sample position (i.e., a sample position after the rth one or the second region) to which a non-separable transform would be applied when the encoding device applies it, this means that the non-separable transform is not applied to the M×N block. In this case, the transform index of the non-separable transform may not be signaled for the M×N block. In this case, the non-separable transform may be set not to be applied to the M×N block, and as an example, the corresponding transform index may be derived as 0.
[0352] The second condition regarding the chroma component according to the present disclosure may mean that there are no non-zero transform coefficients in a predetermined area within the chroma component block of the current block. Here, the predetermined area within the chroma component block can be defined as the remaining area in the chroma component block except for the first area, which can be referred to as the second area to distinguish it from the first area. The first area can be defined as an area consisting of the same number of samples (or sample positions) as the input length of the backward inseparable transform corresponding to the size of the chroma component block. The corresponding input length may mean the number of transform coefficients input to the backward inseparable transform. The first area may include the upper left sample position of the chroma component block. The first area may include at least one non-zero transform coefficient. The first area may be the area to which the last significant coefficient in the chroma component block belongs. The width and height of the first area may be less than or equal to the width and height of the chroma component block, respectively.
[0353] Although the inseparable transform is not applied to the chroma component block, the input length of the backward inseparable transform can be determined based on the size of the chroma component block. Here, the input length of the inseparable transform can be determined as the input length of the inseparable transform determined based on the size of the chroma component block. Alternatively, the matrix size (or input / output length) of the inseparable transform can be determined based on the size of the luminance component block corresponding to the chroma component block, and half of the input length of the corresponding inseparable transform can be determined as the input length of the inseparable transform. The method of determining the inseparable transform matrix according to the block size is as described above, and a detailed description is omitted here.
[0354] Specifically, the non-separable transform is not applied to the chroma components, but zeroing may be applied to the transform coefficients of predetermined areas within the chroma component block in the same manner as when the non-separable transform is applied to the chroma components. In other words, the encoding device may retain only the transform coefficients of some areas within the chroma component block and set the transform coefficients of the remaining areas to 0. Here, the aforementioned some areas may refer to areas where the transform coefficients corresponding to the non-separable transform output are arranged if the non-separable transform is applied to the chroma component block.
[0355] As an example, assume that it is encoded in a 4:2:0 color format, the tree type of the current block is a single tree, and the sizes of the luma component block and the chroma component block are 8×16 and 4×8, respectively. Both 8×16 and 4×8 are block sizes that allow inseparable transforms. In this case, the inseparable transform can be applied to the luma component block, but not to the chroma component block. However, zeroing can also be applied to the chroma component block in the same manner as when the inseparable transform is applied.
[0356] Specifically, when it is assumed that forward NSPT is performed on a 4×8 block based on a 20×32 NSPT matrix, only 20 transform coefficients can be output through NSPT. The 20 transform coefficients will be arranged sequentially from the upper left sample position in the 4×8 block according to a predetermined scanning order, and zeroing will be applied to the remaining 12 sample positions. When zeroing is applied to the chroma component block in the same manner, the pre-output transform coefficients can be retained as is from the upper left sample position in the chroma component block to the 20th sample position according to the predetermined scanning order, and 0 can be allocated from the 21st sample position to the 32nd sample position. In addition, the chroma component block can be composed of a Cb component block and a Cr component block, and zeroing can be applied to the two component blocks in the same manner.
[0357] The zeroing of the chroma component block may be zeroing according to NSPT or zeroing according to LFNST.
[0358] Specifically, when the size of the chroma component block (or the size of the luminance component block corresponding to the chroma component block) corresponds to a block size to which NSPT is applicable, a zeroing method according to NSPT may be applied to the chroma component block. Specifically, when the size of the chroma component block (or the size of the luminance component block corresponding to the chroma component block) corresponds to a block size to which LFNST is applicable, a zeroing method according to LFNST may be applied to the chroma component block.
[0359] Depending on the tree type of the current block, either the size of the luma component block or the size of the chroma component block may be selectively used to determine whether to apply the zeroing method based on NSPT or the zeroing method based on LFNST to the chroma component block. For example, when the tree type of the current block is a single tree, the determination may be based on the size of the luma component block, and when the tree type of the current block is a dual tree, the determination may be based on the size of the chroma component block.
[0360] When it is determined that both the NSPT and the LFNST are applicable to a chroma component block, zeroing may be applied based on either the output length of the NSPT or the output length of the LFNST. When both the NSPT and the LFNST are applicable to a chroma component block, zeroing may be applied based on the output length of the inseparable transform with a preset priority. For example, the NSPT may have a higher priority than the LFNST. The LFNST may have a higher priority than the NSPT. Alternatively, when both the NSPT and the LFNST are applicable to a chroma component block, zeroing may be applied based on the minimum or maximum value of the output lengths of the NSPT and the LFNST. For example, when both the NSPT and the LFNST are applied to a chroma component block, and the output length of the forward NSPT corresponding to the chroma component block is 20 and the output length of the forward LFNST corresponding to the chroma component block is 16, only the transform coefficient at the 16th sample position from the top left sample position in the chroma component block may be retained, and the transform coefficients at the remaining sample positions may be set to 0.
[0361] As described above, when the current block is encoded in a single tree, an inseparable transform is applied only to the luma component, and an inseparable transform is not applied to the chroma component but is applied to zero, the first condition of the luma component and the second condition of the chroma component can be checked, and when the first condition and the second condition are met, the transform index of the current block (particularly the luma component block) can be signaled.
[0362] Alternatively, when the tree type of the current block is single-tree and a non-separable transform is applied only to the luma component, the above-described zeroing may not be applied to the chroma components. In this case, the first condition for the luma component may be checked, while the second condition for the chroma components may not be checked. When the first condition for the luma component is met, the transform index of the current block (particularly the luma component block) may be signaled. On the other hand, when the first condition for the luma component is not met (i.e., when non-zero transform coefficients exist in the second region within the luma component block), no transform index may be signaled for the current block. In this case, the corresponding transform index may be derived as 0.
[0363] Alternatively, when the tree type of the current block is a single tree and a non-separable transform is applied to the luma component and the chroma component, both the first condition for the luma component and the second condition for the chroma component may be checked. In this case, when both the first and second conditions are met, the transform index of the current block may be signaled. When neither the first or second conditions are met, the transform index of the current block may not be signaled.
[0364] When the color format is 4:2:0, the tree type of the current block is single tree, and the size of the luma component block is M×N, the size of the chroma component block can be (M / 2)×(N / 2). It is assumed that the intra subpartitioning (ISP) mode is not applied to the current block. When an inseparable transform is applied to the M×N block of the luma component, an inseparable transform such as NSPT or LFNST can be applied to the chroma components, or NSPT and LFNST may not be applied. Specifically, it can be divided as follows.
[0365] 1) When LFNST is applied to an M×N block of the luminance component
[0366] 1-a) Apply LFNST to the (M / 2)×(N / 2) transform block of the chroma component
[0367] 1-b) Apply NSPT to the (M / 2)×(N / 2) transform block of the chroma component
[0368] 1-c) Both NSPT and LFNST are not applied to the (M / 2)×(N / 2) transform block of the chroma component (in this case, the DCT-2 based horizontal / vertical main transform is applied to the chroma component)
[0369] 2) When NSPT is applied to an M×N transform block of the luminance component
[0370] 2-a) Apply LFNST to the (M / 2)×(N / 2) transform block of the chroma component
[0371] 2-b) Apply NSPT to the (M / 2)×(N / 2) transform block of the chroma component
[0372] 2-c) Both NSPT and LFNST are not applied to the (M / 2)×(N / 2) transform block of the chroma component (in this case, the DCT-2 based horizontal / vertical main transform is applied to the chroma component)
[0373] Assume that NSPT and LFNST are applicable when both M and N are greater than or equal to 4, and NSPT is applicable to 4x4, 4x8, 8x4, and 8x8 blocks. In this case, when the tree type of the current block is a single tree, the following may occur.
[0374] When the luma component block is a 16×8 block, the chroma component block can be an 8×4 block. It can be configured to ensure that LFNST is applied to the luma component and NSPT is applied to the chroma components. Alternatively, it can be configured to ensure that LFNST is applied to the luma component and that neither NSPT nor LFNST is applied to the chroma components. In this case, even when neither NSPT nor LFNST is applied to the chroma components, zeroing can be applied in the same manner as when NSPT or LFNST is applied as described above. In addition, since the chroma component block is an 8×4 block in this example, zeroing according to NSPT can be applied.
[0375] When the luma component block is an 8×8 block, the chroma component block can be a 4×4 block. It can be configured to ensure that NSPT is applied to the luma component and NSPT is applied to the chroma components. Alternatively, it can be configured to ensure that NSPT is applied to the luma component and that neither NSPT nor LFNST is applied to the chroma components. In this case, even when neither NSPT nor LFNST is applied to the chroma components, zeroing can be applied in the same manner as when NSPT or LFNST is applied as described above. In addition, since the chroma component block is a 4×4 block in this example, zeroing according to NSPT can be applied.
[0376] When the luma component block is an 8×4 block, the chroma component block may be a 4×2 block. It may be configured to ensure that NSPT is applied to the luma component and that both NSPT and LFNST are not applied to the chroma components. In this case, zeroing according to NSPT or LFNST may not be applied to the chroma components. Instead, the transform coefficients from the upper left sample position to the nth sample position in the chroma component block according to a predetermined scanning order may remain as they are, and zeroing may be applied to the sample positions after the nth. Here, n is a value predefined equally for the encoding device and the decoding device, and may be an integer greater than or equal to 2. As an example, n may be 4. The fourth sample position from the upper left sample position may belong to a 2×2 area that is an area including the upper left sample of the chroma component block.
[0377] When the luma component block is a 32×32 block, the chroma component block can be a 16×16 block. It can be configured to ensure that LFNST is applied to the luma component and that LFNST is applied to the chroma components. Alternatively, it can be configured to ensure that NFNST is applied to the luma component and that neither NSPT nor LFNST is applied to the chroma components. In this case, even when neither NSPT nor LFNST is applied to the chroma components, zeroing can be applied in the same manner as when NSPT or LFNST is applied as described above. Furthermore, since the chroma component blocks are 16×16 blocks in this example, zeroing according to LFNST can be applied.
[0378] When the tree type of the current block is a single tree and intra subpartitioning (ISP) mode is applied to the current block, or when the tree type of the current block is a dual tree, the width and height of the chroma component block may not be half the width and height of the luma component block, respectively. In this case, for the luma component, whether to apply NSPT or LFNST and / or the matrix size (or input / output length) of the corresponding inseparable transform can be determined based on the size of the corresponding transform block. For the chroma components, whether to apply NSPT or LFNST and / or the matrix size (or input / output length) of the corresponding inseparable transform can be determined based on the size of the corresponding transform block. In addition, when the tree type of the current block is a single tree, as described above, it can be configured to ensure that NSPT or LFNST is not applied to the chroma components. Of course, zeroing can also be applied in the same manner as when NSPT or LFNST is applied.
[0379] When a non-separable transform is applied to the luma component or the chroma component, the transform index of the non-separable transform can be signaled when the following conditions are met. The conditions for signaling the transform index described below can be considered as additional conditions to the first condition for the luma component and the second condition for the chroma component. In other words, the transform index can be signaled only when both the first and second conditions are met and at least one of the conditions described below is met. Alternatively, the conditions described below can be considered as independent conditions, regardless of the first and second conditions. In other words, when at least one of the conditions described below is met, the transform index can be signaled regardless of whether the first and second conditions are met.
[0380] 1) When a non-zero transform coefficient exists at a position other than the upper left sample position of at least one transform block among the color components (i.e., Y, Cb, Cr), the transform index may be signaled (i.e., when no non-zero transform coefficient exists at a position other than the upper left sample position of the transform blocks of all color components), the transform index may not be signaled.
[0381] As an example, when the tree type of the current block is a dual tree and the current block is a transform block of a chroma component, the corresponding transform index can be signaled only when a non-zero transform coefficient exists at a position other than the upper left sample position of at least one transform block among the Cb component and the Cr component.
[0382] Alternatively, when the tree type of the current block is a single tree, the corresponding transform index can be signaled only when there is a non-zero transform coefficient at a position other than the upper left sample position of at least one transform block among the luma component and the chroma component (the chroma component may be composed of a Cb component and a Cr component).
[0383] Alternatively, when the tree type of the current block is a single tree and a non-separable transform is applied only to the luma component, a transform index may be signaled when a non-zero transform coefficient exists at a position other than the upper left sample position of at least one transform block among the luma component and the chroma component and the first condition for the luma component and the second condition for the chroma component are satisfied. Even when no non-zero transform coefficient exists in the transform block for the luma component, the transform index may be signaled when the first and second conditions are satisfied.
[0384] 2) When at least one transform block among the color components (i.e., Y, Cb, Cr) of the current block is encoded using transform skipping (i.e., when the transform skip flag of the corresponding component is 1), the inseparable transform may not be applied to the transform blocks of all color components. When any color component is encoded using transform skipping, the transform index of the current block may not be signaled. The corresponding transform index may be derived as 0.
[0385] Alternatively, transform skipping may be applied to transform blocks of some of the color components of the current block, and transform skipping may not be applied to transform blocks of other components. In this case, the inseparable transform may not be applied to the transform blocks of the color components to which transform skipping is applied, and the inseparable transform may be applied to transform blocks of the remaining color components.
[0386] Based on whether a color component to which transform skipping is applied satisfies a predetermined condition, a non-separable transform may be applied to the color component to which transform skipping is not applied, and a transform index for the corresponding color component may be signaled. For example, a transform index may be signaled for the color component to which transform skipping is not applied only when the color component to which transform skipping is applied satisfies the predetermined condition. Here, the predetermined condition may include at least one of the first condition for the luma component or the second condition for the chroma components described above, or a third condition that a non-zero transform coefficient exists at a position other than the top-left sample position in the transform block of at least one color component. Alternatively, the predetermined condition may not be checked for color components to which transform skipping is applied.
[0387] Reference Figure 4 , the current block can be reconstructed based on the residual samples of the current block (S420).
[0388] The prediction samples of the current block may be derived based on the intra prediction mode of the current block. The reconstructed samples of the current block may be generated based on the prediction samples and the residual samples of the current block.
[0389] Figure 6 A schematic configuration of a decoding device (300) for performing an image decoding method according to the present disclosure is shown.
[0390] Reference Figure 6, the decoding device (300) according to the present disclosure may include a transform coefficient deriver (600), a residual sample deriver (610) and a reconstructed block generator (620). The transform coefficient deriver (600) may be configured in Figure 3 In the entropy decoder (310) of Figure 3 In the residual processor (320) of Figure 3 in the adder (340).
[0391] The transform coefficient deriver (600) can obtain residual information of the current block from the bitstream and decode it to derive the transform coefficient of the current block.
[0392] The residual sample deriver (610) may derivate the residual samples of the current block by performing at least one of dequantization or inverse transformation on the transform coefficients of the current block.
[0393] The residual sample deriver (610) can determine the transform kernel of the inverse transform of the current block by a predetermined transform kernel determination method, and derive the residual sample of the current block based on the transform kernel. Figure 4 The description is the same and its detailed description will be omitted here.
[0394] The reconstructed block generator (620) may reconstruct the current block based on the residual samples of the current block.
[0395] Figure 7 The following illustrates an image encoding method performed by an encoding device (200) according to an embodiment of the present disclosure.
[0396] Reference Figure 7 , the residual samples of the current block can be derived (S700).
[0397] The residual samples of the current block may be derived by subtracting the predicted samples from the original samples of the current block. Here, the predicted samples may be derived based on a predetermined intra prediction mode.
[0398] Reference Figure 7 , a transform coefficient of the current block may be derived by performing at least one of transform or quantization on the residual samples of the current block ( S710 ).
[0399] The transformation method according to the present disclosure can be understood as referring to Figure 4 The inverse process of the inverse transformation described. The method and reference for determining the transformation kernel of the transformation Figure 4 The detailed description will be omitted here.
[0400] For example, one or more transform sets may be defined / configured for the transform of the current block, and each transform set may include one or more transform core candidates. In this case, one of the multiple transform sets may be selected as the transform set for the current block. One of the multiple transform core candidates belonging to the transform set for the current block may be selected. This selection may be performed implicitly based on the context of the current block. Alternatively, an optimal transform set and / or transform core candidate for the current block may be selected, and its index may be signaled.
[0401] Alternatively, the transform core of the current block may be determined based on an MTS set. One of a plurality of MTS sets may be selected based on at least one of the size of the current block or the intra-frame prediction mode. The selected MTS set may include one or more transform core candidates. One of the one or more transform core candidates may be selected, and the transform core of the current block may be determined based on the selected transform core candidate. The selection of the transform core candidate may be performed using a transform core candidate index derived based on the context of the current block. Alternatively, an optimal transform core candidate for the current block may be selected, and a transform core candidate index indicating the selected transform core candidate may be signaled.
[0402] Alternatively, the transform kernel for the current block may be determined based on a non-separable primary transform (NSPT) kernel. When the current block size belongs to the first group of NSPT-applicable block sizes, a forward NSPT may be applied to the current block. When the current block size belongs to the second group, the forward NSPT may not be applied to the current block. When the current block size belongs to the second group, a forward separable primary transform (e.g., DCT-2) may be applied to the residual samples of the current block to derive transform coefficients. A forward LFNST may also be applied to all or part of the transform coefficients derived using the separable primary transform.
[0403] In addition, NSPT can be applied based on at least one of the tree type or component type of the current block. The symmetry between intra prediction modes or the symmetry between block shapes can be used to determine the NSPT kernel (or NSPT matrix) used for NSPT. When forward NSPT is applied to an M×N current block, the NSPT kernel can be expressed as r×MN. Here, r means the output length of NSPT or the number of transform coefficients generated by NSPT, and MN is the product of the width and height of the current block, which can mean the input length of NSPT or the number of residual samples to which NSPT is applied. The method for determining the size of the NSPT kernel is the same as that of reference numerals 1 and 2. Figure 4 Same as described.
[0404] The LFNST index and / or NSPT index used for the transform can be encoded as an integrated syntax, or the LFNST index and NSPT index can be encoded separately and inserted into the bitstream. Binarization of LFNST index and NSPT index and allocation and reference of CABAC context and initial value Figure 4 Same as described.
[0405] In addition, refer to Figure 4 A method of signaling a transform index is described, which may also be applied to a method of encoding a transform index.
[0406] Reference Figure 7 , a bitstream may be generated by encoding the transform coefficients of the current block ( S720 ).
[0407] Residual information about the transformation coefficient may be generated based on the transformation coefficient of the current block, and a bitstream may be generated by encoding the residual information.
[0408] Figure 8 A schematic configuration of an encoding device (200) for performing an image encoding method according to the present disclosure is shown.
[0409] Reference Figure 8 According to the present disclosure, the encoding device (200) may include a residual sample deriver (800), a transform coefficient deriver (810), and a transform coefficient encoder (820). The residual sample deriver (800) and the transform coefficient deriver (810) may be configured in Figure 2 In the residual processor (230) of Figure 2 in the entropy encoder (240).
[0410] The residual sample deriver (800) can derive the residual samples of the current block by subtracting the predicted samples from the original samples of the current block. Here, the predicted samples can be derived based on a predetermined intra prediction mode.
[0411] The transform coefficient deriver (810) may derive transform coefficients of the current block by performing at least one of transform or quantization on the residual samples of the current block. The transform coefficient deriver 810 may determine a transform kernel of the current block based on at least one of the above-described embodiments 1 to 3 and apply the transform kernel to the residual samples of the current block to derive transform coefficients.
[0412] The transform coefficient encoder (820) may encode the transform coefficients of the current block to generate a bitstream.
[0413] In the above embodiments, the method is described based on a flowchart as a series of steps or boxes, but the corresponding embodiments are not limited to the order of the steps. Some steps may occur simultaneously with other steps as described above or in a different order. In addition, it will be understood by those skilled in the art that the steps shown in the flowchart are not exclusive and other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of the present disclosure.
[0414] The above-mentioned method according to the embodiment of the present disclosure can be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure can be included in an apparatus that performs image processing, such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0415] In the present disclosure, when the embodiment is implemented as software, the above method can be implemented as a module (process, function, etc.) that performs the above functions. The module can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, a logic circuit and / or a data processing device. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or another storage device. In other words, the embodiments described herein can be executed by being implemented on a processor, a microprocessor, a controller or a chip. For example, the functional units shown in the various figures can be executed by being implemented on a computer, a processor, a microprocessor, a controller or a chip. In this case, the information (e.g., information about instructions) or the algorithm used for implementation can be stored in a digital storage medium.
[0416] In addition, the decoding device and the encoding device to which the embodiments of the present disclosure are applied may be included in a multimedia broadcast sending and receiving device, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video conversation device, a real-time communication device similar to video communication, a mobile streaming device, a storage medium, a camera, a device for providing a video on demand (VOD) service, an OTT (over the top) video device, a device for providing an Internet streaming service, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a video phone video device, a transportation terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, a ship terminal, etc.) and a medical video device, and may be used to process a video signal or a data signal. For example, an OTT (over the top) video device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0417] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the embodiment of the present disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BD), universal serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical media storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (for example, transmission via the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or can be sent via a wired / wireless communication network.
[0418] In addition, the embodiments of the present disclosure may be implemented by a computer program product through program code, and the program code may be executed on a computer through the embodiments of the present disclosure. The program code may be stored on a computer readable carrier.
[0419] Figure 9 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown.
[0420] Reference Figure 9 A content streaming system to which an embodiment of the present disclosure is applied may generally include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.
[0421] The encoding server generates a bitstream by compressing content input from a multimedia input device such as a smartphone, a camera, a camcorder, etc. into digital data and transmits it to the streaming server. As another example, when the multimedia input device such as a smartphone, a camera, a camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0422] A bitstream may be generated by an encoding method or a bitstream generating method to which an embodiment of the present disclosure is applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0423] The streaming server transmits multimedia data to a user device via a network server based on a user request, and the network server serves as an intermediary for informing the user of available services. When a user requests a desired service from the network server, the network server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server controls the commands and responses between the various devices in the content streaming system.
[0424] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a specific period of time.
[0425] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touch-screen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0426] The respective servers in the content streaming system may operate as distributed servers, and in this case, data received from the respective servers may be distributed and processed.
[0427] The claims set forth herein may be combined in various ways. For example, the technical features of the method claims of the present disclosure may be combined and implemented as a device, and the technical features of the device claims of the present disclosure may be combined and implemented as a method. Furthermore, the technical features of the method claims of the present disclosure and the technical features of the device claims of the present disclosure may be combined and implemented as a device, and the technical features of the method claims of the present disclosure and the technical features of the device claims of the present disclosure may be combined and implemented as a method.
Claims
1. A method for decoding an image, comprising the following steps: Obtain residual information from the bitstream; deriving transform coefficients of the current block based on the residual information; performing at least one of dequantization or inverse transformation on the transform coefficients of the current block to derive residual samples of the current block; as well as reconstructing the current block based on the residual samples of the current block, wherein the inverse transform is performed based on a backward non-separable primary transform NSPT, wherein a transform core for the NSPT is determined based on a transform index indicating any one of a plurality of transform core candidates belonging to a transform set, and The transform index is signaled based on at least one of a first condition for a luma component or a second condition for a chroma component.
2. The image decoding method according to claim 1, wherein: The first condition for the luma component indicates that there is no non-zero transform coefficient in a predetermined area within the luma component block of the current block, and The second condition for the chrominance component indicates that there is no non-zero transform coefficient in a predetermined area within the chrominance component block of the current block.
3. The image decoding method according to claim 2, wherein: The predetermined area within the luminance component block is a remaining area excluding the first area from the luminance component block, and The first region is a region consisting of the same number of samples as the number of transform coefficients input to the inseparable main transform.
4. The image decoding method according to claim 2, wherein: The predetermined area within the chroma component block is a remaining area excluding the first area from the chroma component block, and The first region is a region consisting of samples having the same number as the input length of the inseparable main transform corresponding to the size of the chrominance component block.
5. The image decoding method according to claim 1, wherein: When the tree type of the current block is a single tree and the non-separable main transform is applied to a luma component block of the current block but not to a chroma component block of the current block, the transform index is signaled based on a case where the first condition for the luma component and the second condition for the chroma component are satisfied. The image decoding method according to claim 1 , wherein: When the tree type of the current block is a single tree and the inseparable main transform is applied to a luma component block of the current block but not to a chroma component block of the current block, the transform index is signaled based on a case where the first condition for the luma component is satisfied, without checking whether the second condition for the chroma component is satisfied.
7. The image decoding method according to claim 1, wherein: When the tree type of the current block is a single tree and the inseparable main transform is applied to a luma component block and a chroma component block of the current block, the transform index is signaled based on a case where the first condition for the luma component and the second condition for the chroma component are satisfied.
8. The image decoding method according to claim 1, wherein: The transform index is signaled based on at least one of the following: whether there is a non-zero transform coefficient at a sample position other than an upper left sample position in at least one of the three color component blocks within the current block or whether the at least one of the three color component blocks is a transform skip block.
9. A method for encoding an image, comprising the following steps: Derive the residual samples of the current block; performing at least one of transform or quantization on the residual samples of the current block to derive transform coefficients of the current block; as well as encoding the transform coefficients of the current block, wherein the transformation is performed based on a forward non-separable primary transform NSPT, The transform kernel for the NSPT is determined to be any one of a plurality of transform kernel candidates belonging to a transform set, wherein a transform index indicating the any one of the plurality of transform kernel candidates is encoded, and The transform index is encoded based on at least one of a first condition for a luma component or a second condition for a chroma component. 10 . A computer-readable storage medium storing a bit stream generated by the image encoding method according to claim 9 .
11. A method for transmitting data, the method comprising the following steps: Obtaining a bitstream for image information, wherein the bitstream is generated by deriving residual samples of a current block, performing at least one of transform or quantization on the residual samples of the current block to derive transform coefficients, and encoding the transform coefficients of the current block; and sending said data comprising said bitstream, wherein the transformation is performed based on a forward non-separable primary transform NSPT, The transform kernel for the NSPT is determined to be any one of a plurality of transform kernel candidates belonging to a transform set, wherein a transform index indicating the any one of the plurality of transform kernel candidates is encoded, and The transform index is encoded based on at least one of a first condition for a luma component or a second condition for a chroma component.