Image encoding / decoding method and apparatus, and recording medium storing bitstreams
By employing a set of transform kernels determined by block dimensions and intra prediction modes, the method addresses inefficiencies in image encoding and decoding, particularly for high-resolution images, improving transform and coding efficiency.
Patent Information
- Application Number
- JP2025521044
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-10-28
AI Technical Summary
Existing image compression techniques struggle to efficiently handle high-resolution and high-quality images, particularly in constructing and signaling transform kernels for current blocks during image encoding and decoding processes.
The method and apparatus provide a set of transform kernels for current blocks, determining horizontal and vertical kernels based on block dimensions and intra prediction modes, and signaling an index for the transform kernel, incorporating non-trigonometric functions to enhance transform performance.
This approach improves transform efficiency and coding efficiency by utilizing various transform kernel candidates and effectively signaling the transform kernel index, enhancing image encoding and decoding processes.
Smart Images

Figure 2025535761000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream. [Background technology]
[0002] 2. Description of the Related Art In recent years, the demand for high-resolution, high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has increased in various application fields, and as a result, highly efficient image compression techniques have been discussed.
[0003] There are various image compression techniques, such as inter-prediction techniques that predict pixel values contained in a current picture from pictures before or after the current picture, intra-prediction techniques that predict pixel values contained in a current picture using pixel information within the current picture, and entropy coding techniques that assign short codes to values that occur frequently and long codes to values that occur less frequently. Using such image compression techniques, image data can be effectively compressed and transmitted or stored. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure seeks to provide a method and apparatus for constructing a predetermined set of transformations for the (inverse) transformation of a current block.
[0005] The present disclosure seeks to provide a method and apparatus for determining one or more candidate transform kernels for a current block.
[0006] The present disclosure seeks to provide a method and apparatus for signaling an index for a transform kernel of a current block. [Means for solving the problem]
[0007] The image decoding method and apparatus according to the present disclosure can derive transform coefficients of a current block from a bitstream, perform at least one of inverse quantization and inverse transformation on the transform coefficients of the current block to derive residual samples of the current block, and reconstruct the current block based on the residual samples of the current block.
[0008] In the image decoding method and apparatus according to the present disclosure, deriving the residual sample may include determining a transform set for a first-order inverse transform of the current block, and deriving a horizontal transform kernel and a vertical transform kernel of the current block from the transform set based on a width and a height of the current block, respectively. Here, the transform set may be determined based on at least one of an intra prediction mode of the current block or a predefined mapping table.
[0009] In the image decoding method and apparatus according to the present disclosure, the mapping table may define a mapping relationship between predefined intra prediction modes and transform sets.
[0010] In the image decoding method and apparatus according to the present disclosure, the transformation set may be mapped to an intra prediction mode that has a symmetric relationship with the intra prediction mode of the current block.
[0011] In the image decoding method and apparatus according to the present disclosure, the transform set may include a plurality of transform kernel candidates, and the plurality of transform kernel candidates may include length-based transform kernel candidates that are applied equally in the horizontal and vertical directions.
[0012] In the image decoding method and apparatus according to the present disclosure, the horizontal transform kernel may be set to a transform kernel candidate having the same length as the width of the current block, and the vertical transform kernel may be set to a transform kernel candidate having the same length as the height of the current block.
[0013] In the image decoding method and apparatus according to the present disclosure, when the current block is an MxN block, a horizontal transform kernel having a length of M for the current block may be the same as a vertical transform kernel having a length of M for an NxM block.
[0014] In the image decoding method and apparatus according to the present disclosure, the number of available transform sets may be one, two, three, four, or seven.
[0015] The image encoding method and apparatus according to the present disclosure may derive residual samples of a current block, derive transform coefficients of the current block by performing at least one of transform and quantization on the residual samples of the current block, and encode the transform coefficients of the current block. Here, deriving the transform coefficients may include determining a transform set for a first-order inverse transform of the current block, and deriving horizontal and vertical transform kernels of the current block from the transform set based on a width and a height of the current block. Here, the transform set may be determined based on at least one of an intra prediction mode of the current block or a predefined mapping table.
[0016] A computer-readable digital storage medium is provided having encoded video / image information stored thereon that enables an image decoding method to be performed by a decoding device according to the present disclosure.
[0017] A computer-readable digital storage medium is provided having stored thereon video / image information generated by the image encoding method of the present disclosure.
[0018] A method and apparatus for transmitting video / image information generated by the image encoding method according to the present disclosure is provided. [Effects of the Invention]
[0019] According to the present disclosure, the performance of a transform can be improved by constructing a transform set that includes a variety of transform kernel candidates.
[0020] The present disclosure can improve the performance of the transform by further utilizing non-trigonometric function based transform kernels in addition to trigonometric function based transform kernels.
[0021] The present disclosure can improve coding efficiency of transform-related information by effectively signaling an index for the transform kernel of the current block. [Brief explanation of the drawings]
[0022] [Figure 1] FIG. 1 illustrates a video / image coding system according to the present disclosure. [Figure 2] 1 is a schematic block diagram of an encoding device to which embodiments of the present disclosure can be applied, in which video / image signals are encoded. [Figure 3] 1 is a schematic block diagram of a decoding device to which embodiments of the present disclosure can be applied, in which video / image signals are decoded. [Figure 4] 1 illustrates an image decoding method performed by a decoding device (300) according to an embodiment of the present disclosure. [Figure 5] FIG. 10 is a diagram illustrating intra-prediction modes and their prediction directions according to the present disclosure. [Figure 6]FIG. 3 is a diagram showing a schematic configuration of a decoding device (300) that performs the image decoding method according to the present disclosure. [Figure 7] 1 illustrates an image encoding method performed by an encoding device (200) according to one embodiment of the present disclosure. [Figure 8] FIG. 1 is a diagram showing a schematic configuration of an encoding device (200) that performs an image encoding method according to the present disclosure. [Figure 9] FIG. 1 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0023] While the present disclosure may be modified in various ways and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present disclosure to the specific embodiments, and it should be understood that the present disclosure includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present disclosure. In the description of each figure, similar reference numerals are used to refer to similar components.
[0024] Terms such as "first," "second," etc. may be used to describe various components, but these components should not be limited by such terms. These terms are used merely to distinguish one component from another. For example, a first component could be termed a second component, and similarly, a second component could be termed a first component, without departing from the scope of the present disclosure. The term "and / or" includes a combination of multiple associated listed items or any item of multiple associated listed items.
[0025] When a component is referred to as being "coupled" or "connected" to another component, it should be understood that the component may be directly coupled or connected to the other component, and that there may be additional components in between. On the other hand, when a component is referred to as being "directly coupled" or "directly connected" to another component, it should be understood that there are no additional components in between.
[0026] The terms used in this application are merely for the purpose of describing particular embodiments and are not intended to limit the present disclosure. The singular terms also include the plural terms unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0027] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the versatile video coding (VVC) standard. The methods / embodiments disclosed herein may also be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0028] This specification presents various embodiments relating to video / image coding, and unless otherwise stated, the above embodiments may be performed in combination with each other.
[0029] In this specification, video may refer to a collection of a series of images over time. A picture generally refers to a unit that represents an image at a specific time period, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). One picture may be composed of one or more slices / tiles. A tile is a rectangular area composed of multiple CTUs in a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs with a height equal to the height of the picture and a width specified by a picture parameter set syntax requirement. A tile row is a rectangular area of CTUs with a height specified by a picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively by CTU raster scanning, while tiles within a picture may be arranged consecutively by tile raster scanning. A slice may contain an integer number of complete tiles or an integer number of consecutive complete CTU rows within the tiles of a picture that may be contained exclusively in a single NAL unit, while a picture may be partitioned into two or more sub-pictures, which may be rectangular regions of one or more slices in a picture.
[0030] A picture element, pixel, or pel can refer to the smallest unit that makes up a picture (or image). A "sample" can also be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and may indicate only a pixel / pixel value of a luminance (luma) component, or may indicate only a pixel / pixel value of a color difference (chroma) component.
[0031] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an MxN block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0032] As used herein, "A or B" can mean "A only," "B only," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B, or C" can mean "A only," "B only," "C only," or "any combination of A, B, and C."
[0033] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Thus, "A / B" can mean "A only," "B only," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0034] As used herein, "at least one of A and B" can mean "A only," "B only," or "both A and B." Furthermore, as used herein, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as being the same as "at least one of A and B."
[0035] Furthermore, in this specification, "at least one of A, B, and C" can mean "A only," "B only," "C only," or "any combination of A, B, and C." Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C."
[0036] Furthermore, parentheses used in this specification may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra prediction," and "intra prediction" may be suggested as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction."
[0037] In this specification, technical features individually described in the same drawing may be embodied individually or simultaneously.
[0038] FIG. 1 is a diagram illustrating a video / image coding system according to the present disclosure.
[0039] Referring to FIG. 1, a video / image coding system may include a first device (a source device) and a second device (a receiving device).
[0040] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.
[0041] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include a computer, tablet, smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which the associated data is generated.
[0042] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0043] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to a receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file according to a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0044] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, which correspond to the operations of the encoding device.
[0045] The renderer can render the decoded video / images, and the rendered video / images can be displayed on a display unit.
[0046] FIG. 2 is a schematic block diagram of an encoding device to which the embodiments of the present disclosure can be applied, where encoding of video / image signals is performed.
[0047] Referring to FIG. 2, the encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image divider 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filterer 260 may be configured by one or more hardware components (e.g., an encoding device chipset or processor) depending on the embodiment. Furthermore, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0048] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided into coding tree units (CTUs) or largest coding units (LCUs) according to a QTBTTT (Quad-tree, Binary-tree, Ternary-tree) structure.
[0049] For example, one coding unit may be divided into multiple coding units having deeper depths based on a quadtree structure, a binary tree structure, and / or a tertiary structure. In this case, for example, the quadtree structure may be applied first, and then the binary tree structure and / or the tertiary structure may be applied later. Alternatively, the binary tree structure may be applied before the quadtree structure. The coding procedure according to the present specification may be performed based on a final coding unit that is not further divided. In this case, based on coding efficiency according to image characteristics, the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0050] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0051] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component, or may represent only a pixel / pixel value of a chroma component. A sample may be used in terms corresponding to one picture (or image), pixel, or pel.
[0052] The encoding apparatus 200 may subtract a prediction signal (prediction block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, a unit in the encoding apparatus 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) may be referred to as a subtraction unit 231.
[0053] The prediction unit 220 may perform prediction on a current block (hereinafter referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit 220 may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit 220 may generate various information related to prediction, such as prediction mode information, as will be described later in the description of each prediction mode, and transmit the information related to prediction to the entropy encoding unit 240. The entropy encoding unit 240 may encode the information related to prediction and output it in the form of a bitstream.
[0054] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional modes may include at least one of DC mode and planar mode. The directional modes may include 33 directional modes or 65 directional modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0055] The inter prediction unit 221 may derive a prediction block for a current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be called collocated reference blocks, collocated control units (colCUs), etc., and the reference picture including the temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit 221 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidates are used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of a skip mode or a merge mode, the inter prediction unit 221 may use motion information of neighboring blocks as motion information for the current block. In the case of the skip mode, unlike in the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of the surrounding block as a motion vector predictor and signaling the motion vector difference.
[0056] The prediction unit 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as a combined inter and intra prediction (CIIP) mode. The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for coding content images / videos, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, a sample value within the picture may be signaled based on information about a palette table and a palette index. The predicted signal generated by the prediction unit 220 may be used to generate a reconstructed signal or a residual signal.
[0057] The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), and a Conditionally Non-Linear Transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to square pixel blocks of the same size, or to non-square blocks of variable sizes.
[0058] The quantization unit 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 rearranges the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and generates information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0059] The entropy encoding unit 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit 240 can encode information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients.
[0060] Encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded using the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 may be transmitted to a transmitting unit (not shown) and / or stored to a storing unit (not shown) configured as an internal / external element of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.
[0061] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, the inverse quantization unit 234 and the inverse transform unit 235 may apply inverse quantization and inverse transform to the quantized transform coefficients to reconstruct a residual signal (residual block or residual sample). The adder 250 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to a prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the current block, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, or may be used for inter prediction of the next picture after filtering, as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during picture encoding and / or reconstruction.
[0062] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240. The information related to filtering may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0063] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter prediction unit 221. This allows the encoding apparatus to avoid prediction mismatch between the encoding apparatus 200 and the decoding apparatus when inter prediction is applied, and also improves coding efficiency.
[0064] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0065] FIG. 3 is a schematic block diagram of a decoding device to which the embodiments of the present disclosure can be applied, in which video / image signals are decoded.
[0066] 3, the decoding device 300 may include an entropy decoding unit (entropy decoder 310), a residual processor (residual processor 320), a predictor (predictor 330), an adder (adder 340), a filter (filter 350), and a memory (memory 360). The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer (dequantizer 321) and an inverse transformer (inverse transformer 322).
[0067] The entropy decoding unit 310, residual processing unit 320, prediction unit 330, addition unit 340, and filtering unit 350 may be configured as a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. Also, the memory 360 may include a decoded picture buffer (DPB) and may be configured as a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0068] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the process by which the video / image information was processed by the encoding apparatus of FIG. 2. For example, the decoding apparatus 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using the processing units applied by the encoding apparatus. Accordingly, the processing unit for decoding may be a coding unit, which may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a tertiary tree structure. One or more transform units may be derived from the coding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 may be reproduced by a playback device.
[0069] The decoding device 300 may receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal may be decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream and derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The decoding device may decode pictures further based on the information on the parameter sets and / or the general constraint information. Signal / received information and / or syntax elements described later in this specification may be decoded by the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded, decoding information on neighboring and current blocks, or information on symbols / bins decoded in previous steps, predicts the occurrence probability of the bins based on the determined context model, and generates symbols corresponding to the values of each syntax element by performing arithmetic decoding of the bins. In this case, after determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbols / bins for the context model of the next symbol / bin.Information related to prediction among the information decoded by the entropy decoding unit 310 may be provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to a residual processing unit 320. The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). In addition, information related to filtering among the information decoded by the entropy decoding unit 310 may be provided to a filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding apparatus may be further configured as an internal / external element of the decoding apparatus 300, or the receiving unit may be a component of the entropy decoding unit 310.
[0070] Meanwhile, the decoding apparatus according to the present specification may be referred to as a video / image / picture decoding apparatus, and the decoding apparatus may be divided into an information decoding apparatus (video / image / picture information decoding apparatus) and a sample decoding apparatus (video / image / picture sample decoding apparatus). The information decoding apparatus may include the entropy decoding unit 310, and the sample decoding apparatus may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.
[0071] The inverse quantization unit 321 can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 can inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0072] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0073] The prediction unit 320 may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit 320 may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0074] The prediction unit 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit 320 may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as a combined inter and intra prediction (CIIP) mode. The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information about a palette table and a palette index may be included in the video / image information and signaled.
[0075] The intra prediction unit 331 may predict a current block by referring to samples in a current picture. The referenced samples may be located in the neighborhood of the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit 331 may determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.
[0076] The inter prediction unit 332 may derive a prediction block for a current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0077] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a prediction signal (prediction block, prediction sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when the skip mode is applied, the prediction block may be used as the reconstructed block.
[0078] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in a current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture. Meanwhile, luma mapping with chroma scaling (LMCS) may be applied during picture decoding.
[0079] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0080] The (modified) reconstructed picture stored in the DPB of the memory 360 may be used as a reference picture in the inter predictor 332. The memory 360 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.
[0081] In this specification, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 may also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0082] FIG. 4 is a diagram illustrating an image decoding method performed by a decoding device 300 according to an embodiment of the present disclosure.
[0083] 4, transform coefficients of a current block can be derived from a bitstream (S400). That is, the bitstream may include residual information of the current block, and the transform coefficients of the current block can be derived by decoding the residual information.
[0084] Referring to FIG. 4, residual samples of the current block can be derived by performing at least one of dequantization and inverse transform on the transform coefficients of the current block (S410).
[0085] When an Adaptive Multiple Transform Selection (MTS) is applied, the inverse transform may be performed based on at least one of DCT-2, DST-7, or DCT-8, which may be referred to as a transform type, transform kernel, or transform core.
[0086] In the present disclosure, the inverse transform may refer to a separable transform. However, without being limited thereto, the inverse transform may also refer to a non-separable transform, and may be a concept including both separable and non-separable transforms. Furthermore, while the inverse transform in the present disclosure refers to a primary transform, it is not limited thereto and may be transformed into an identical / similar form and applied to a secondary transform.
[0087] For example, as a method for inverse transformation, only DCT-2 and a non-separable transform may be used, or a non-separable transform may be further used in addition to at least one of DCT-2, DST-7, and DCT-8, or a non-separable transform may be used instead of one or more transform kernels of DCT-2, DST-7, and DCT-8.
[0088] As a more specific example, when transform kernel candidates for a separable transform include (DCT-2, DCT-2), (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8), a non-separable transform may replace or be added to one or more of the five transform kernel candidates. Here, a notation such as (Transform 1, Transform 2) indicates that transform 1 is applied in the horizontal direction and transform 2 is applied in the vertical direction. When a non-separable transform replaces some of the transform kernel candidates, the remaining transform kernel candidates except for (DCT-2, DCT-2) and (DST-7, DST-7) may be replaced with a non-separable transform. However, the above transform kernel candidates are merely examples, and other types of DCTs and / or DSTs may be included, and a transform skip may also be included as a transform kernel candidate.
[0089] A non-separable transform can refer to a transform or inverse transform based on a non-separable transform matrix. That is, unlike a separable transform, which separates vertical and horizontal transforms and performs horizontal and vertical transforms independently, a non-separable transform can perform horizontal and vertical transforms at the same time.
[0090] For example, when a non-separable transform is performed on a 4x4 block, the input data X to the non-separable transform is as follows:
[0091] [Formula 1]
number
[0092] When the input data X is expressed in vector form, the vector X' may be expressed as follows:
[0093] [Formula 2]
number
[0094] In this case, the non-separable transformation may be performed as shown in Equation 3 below.
[0095] [Formula 3]
number
[0096] In Equation 3, F represents a transform coefficient vector, T represents a 16x16 non-separable transform matrix, and · represents matrix-vector multiplication.
[0097] A 16x1 transform coefficient vector F may be derived by Equation 3 above, and F may be rearranged into 4x4 blocks according to a predetermined scan order, which may be a horizontal scan, a vertical scan, a diagonal scan, a z scan, a raster scan, or a predefined scan.
[0098] The non-separable transform set and / or transform kernel for the non-separable transform may be configured differently based on at least one of the prediction mode (e.g., intra mode, inter mode, etc.), the width, height, or number of pixels of the current block, the position of the sub-block within the current block, explicitly signaled syntax elements, statistical characteristics of surrounding samples, whether or not a quadratic transform is used, or the quantization parameter (QP).
[0099] Specifically, in an intra mode, predefined intra prediction modes may be grouped to correspond to n non-separable transform sets, and each non-separable transform set may include k transform kernel candidates, where n and k may be any constants that follow rules (conditions) that are defined identically for the encoding device and the decoding device.
[0100] The number of non-separable transform sets and / or the number of transform kernel candidates included in the non-separable transform sets may be configured to vary depending on the width and / or height of the current block. For example, for a 4x4 block, n1 non-separable transform sets and k1 transform kernel candidates may be configured. For a 4x8 block, n2 non-separable transform sets and k2 transform kernel candidates may be configured. Furthermore, the number of non-separable transform sets and the number of transform kernel candidates included in each non-separable transform set may be configured to vary depending on the product of the width and height of the current block. For example, if the product of the width and height of the current block is equal to or greater than 256, n3 non-separable transform sets and k3 transform kernel candidates may be configured; otherwise, n4 non-separable transform sets and k4 transform kernel candidates may be configured. That is, since the degree of change in the statistical characteristics of the residual signal varies depending on the block size, the number of non-separable transform sets and the number of transform kernel candidates may be configured to vary to reflect this.
[0101] When a current block is divided into multiple sub-blocks, the statistical characteristics of the residual signals may differ for each sub-block, so the number of non-separable transform sets and transform kernel candidates may be configured to be different. For example, when a 4x8 or 8x4 block is divided into two 4x4 sub-blocks and a non-separable transform is applied to each sub-block, n5 non-separable transform sets and k5 transform kernel candidates may be configured for the top-left 4x4 sub-block, and n6 non-separable transform sets and k6 transform kernel candidates may be configured for the other 4x4 sub-blocks.
[0102] The number of non-separable transform sets and transform kernel candidates may be configured to differ from one another based on an explicitly signaled syntax element. Information indicating one of a plurality of non-separable transform configurations may be used as the syntax element. For example, if three types of non-separable transform configurations are supported (i.e., n7 non-separable transform sets and k7 transform kernel candidates, n8 non-separable transform sets and k8 transform kernel candidates, and n9 non-separable transform sets and k9 transform kernel candidates), the syntax element may have a value of 0, 1, or 2, and the non-separable transform configuration to be applied to the current block may be determined based on the value of the signaled syntax element.
[0103] The number of non-separable transformation sets and transformation kernel candidates may be configured differently depending on whether and / or what kind of secondary transformation is applied. For example, when no secondary transformation is applied, n 10 A set of non-separable transformations and k 10 A non-separable transformation configuration containing n candidate transformation kernels can be applied. 11 A set of non-separable transformations and k 11 A non-separable transform configuration including transform kernel candidates can be applied.
[0104] Different non-separable transform configurations may be applied according to the quantization parameter (QP) and / or range of QP values. For example, when the QP value has a small value, n 12 A set of non-separable transformations and k 12On the other hand, when the QP value is large, n 13 A set of non-separable transformations and k 13 A non-separable transform configuration including candidate transform kernels can be applied. If the QP value is equal to or less than a threshold (e.g., 32), the QP value can be classified as having a small value; otherwise, the QP value can be classified as having a large value. Alternatively, the QP value range can be divided into three or more ranges, and a different non-separable transform configuration can be applied to each range.
[0105] For relatively large blocks, instead of using a non-separable transform corresponding to the width and height of the block, the block can be divided into multiple sub-blocks and a non-separable transform corresponding to the width and height of the sub-block can be used. For example, when performing a non-separable transform on a 4x8 block, the 4x8 block can be divided into two 4x4 sub-blocks and a 4x4 block-based non-separable transform can be used for each 4x4 sub-block. Or, for an 8x16 block, it can be divided into two 8x8 sub-blocks and an 8x8 block-based non-separable transform can be used.
[0106] The non-separable transform set may be determined based on the intra prediction mode of the current block and a mapping table. The mapping table may define a mapping relationship between predefined intra prediction modes and the non-separable transform set. The predefined intra prediction modes may include two non-directional modes and 65 directional modes. Generally, a non-separable transform has a larger transform kernel size than a separable transform. This means that the computational complexity required for the transform process is high and a large memory is required for storing the transform kernel. Meanwhile, a separable transform can only consider statistical characteristics existing in the horizontal and / or vertical directions, while a non-separable transform can simultaneously consider statistical characteristics in two-dimensional space including the horizontal and vertical directions, thereby providing better compression efficiency. Since the statistical characteristics and diversity of residuals differ depending on the directionality of the intra prediction mode, a non-separable transform may be absolutely necessary. Alternatively, there may be an intra prediction mode that can sufficiently grasp the characteristics of the residuals through separable transform alone. Therefore, by predefining which transform to use according to the intra prediction mode in the encoding device and the decoding device, a transform process can be designed with optimized complexity and memory requirements. The non-directional modes may include a planar mode numbered 0 and a DC mode numbered 1, and the directional modes may include intra prediction modes numbered 2 to 66. However, this is merely an example, and the present disclosure may also be applied to cases where the number of predefined intra prediction modes is different.
[0107] By applying wide angle intra prediction (WAIP), the predefined intra prediction modes may further include intra prediction modes from -14 to -1 and intra prediction modes from 67 to 80.
[0108] FIG. 5 exemplarily illustrates intra prediction modes and their prediction directions according to the present disclosure. Referring to FIG. 5, modes -14 to -1 and modes 2 to 33 and modes 35 to 80 are symmetrical in terms of prediction direction around mode 34. For example, modes 10 and 58 are symmetrical around the direction corresponding to mode 34, and mode -1 is symmetrical to mode 67. Therefore, for vertical direction modes that are symmetrical to horizontal direction modes around mode 34, input data can be transposed before use. Transposing input data means that rows in MxN two-dimensional block input data become columns and columns become rows, forming NxM data.
[0109] For example, when a 4x4 block is used, 16 data pieces constituting the 4x4 block can be appropriately arranged to form a 16x1 one-dimensional vector for non-separable transformation. In this case, the one-dimensional vector can be arranged in row-major order or column-major order. The residual samples resulting from the non-separable transformation can be arranged in the above order to form a two-dimensional block.
[0110] If the data arrangement order for constructing the 16x1 input vector is row-major for modes -14 to -1 and modes 2 to 33, the input vector may be constructed in column-major order for modes 35 to 80.
[0111] Although the 34th mode can be considered neither a horizontal nor a vertical directional mode, this disclosure classifies it as belonging to the horizontal directional mode. That is, for modes -14 to -1 and 2 to 33, the input data sorting method for the horizontal directional mode, i.e., row-priority order, is used, and the input data can be transposed and used for the vertical directional mode, which is symmetrical around the 34th mode.
[0112] Non-square blocks cannot utilize symmetry in square blocks (i.e., the symmetry between the P mode and the (68-P) mode in an NxN block (2<=P<=33) or the symmetry between the Q mode and the (66-Q) mode (-14<=Q<=-1)). Therefore, in addition to the symmetry based only on the intra prediction mode, symmetry between block types that are in a mutually preceding relationship, i.e., the symmetry between the KxL block and the LxK block, can also be utilized. Specifically, a symmetrical relationship exists between the KxL block predicted with the P mode and the LxK block predicted with the (68-P) mode. Alternatively, a symmetrical relationship exists between the KxL block predicted with the Q mode and the LxK block predicted with the (66-Q) mode.
[0113] Since the KxL block having mode 2 and the LxK block having mode 66 are considered to be symmetrical to each other, the same transform kernel can be applied to the KxL block and the LxK block. If a non-separable transform set is mapped to the intra prediction mode of the KxL block, in order to apply a non-separable transform to the LxK block, the non-separable transform set can be derived using a mapping table corresponding to the KxL block based on mode (68-P) instead of mode P applied to the LxK block. Alternatively, the non-separable transform set can be derived using a mapping table corresponding to the KxL block based on mode (66-Q) instead of mode Q applied to the LxK block.
[0114] For example, to apply a non-separable transform to an LxK block, a non-separable transform set can be selected based on mode 2 instead of mode 66. For a KxL block, input data can be read in a predetermined order (e.g., row-major order or column-major order) to construct a one-dimensional vector, and then the non-separable transform can be applied. For an LxK block, input data can be read in a transposed order to construct a one-dimensional vector, and then the non-separable transform can be applied. That is, if the KxL block is read in row-major order, the LxK block can be read in column-major order. Conversely, if the KxL block is read in column-major order, the LxK block can be read in row-major order.
[0115] Also, when the 34th mode is applied to a KxL block, a non-separable transform set is determined based on the 34th mode, and the input data is read in a predetermined order to construct a one-dimensional vector, which can then be used to perform the non-separable transform. Similarly, when the 34th mode is applied to an LxK block, a non-separable transform set is determined based on the 34th mode, but the input data is read in a transposed order to construct a one-dimensional vector, which can then be used to perform the non-separable transform.
[0116] While this disclosure has described a method for determining a non-separable transform set and a method for configuring input data based on a KxL block, the non-separable transform can be performed by using the same symmetry described above for a KxL block based on an LxK block. Alternatively, a block whose width is greater than its height may be restricted to be used as a reference block. Alternatively, a non-square block may be restricted from utilizing symmetry. In this case, a non-square block may use a different number of non-separable transform sets and / or transform kernel candidates than a square block, and a different mapping table may be used to select a non-separable transform set.
[0117] An example of a mapping table for selecting a non-separable transformation set is as follows:
[0118] [Table 1]
[0119] Table 1 shows an example of allocating non-separable transform sets for each intra prediction mode when there are five non-separable transform sets. The value of predModeIntra indicates the value of the intra prediction mode taking WAIP into consideration, and TrSetIdx is an index indicating a specific non-separable transform set. It can be seen from Table 1 that the same non-separable transform set is applied to modes located in symmetric directions depending on the intra prediction mode. Table 1 is merely an example of using five non-separable transform sets and does not limit the total number of non-separable transform sets for non-separable transforms.
[0120] Alternatively, as shown in Table 2, non-separable transforms may not be applied to WAIP for compression performance reasons.
[0121] [Table 2]
[0122] Alternatively, as shown in Table 3, a separate non-separable transform set may not be configured for WAIP, and non-separable transform sets corresponding to adjacent intra prediction modes may be shared.
[0123] [Table 3]
[0124] The non-separable transform set may include a plurality of transform kernel candidates, and any one of the plurality of transform kernel candidates may be selectively used. To this end, an index signaled by a bitstream may be used. Alternatively, any one of the plurality of transform kernel candidates may be implicitly determined based on context information of the current block. Here, the context information may indicate the size of the current block or whether a non-separable transform is applied to a neighboring block. Here, the size of the current block may be defined as the width, height, maximum / minimum value of the width and height, the sum of the width and height, or the product of the width and height.
[0125] A method for determining a transformation kernel for the inverse transformation of a current block will now be described in detail.
[0126] Example 1
[0127] As mentioned above, inverse transforms can be classified into separable transforms and non-separable transforms. A separable transform refers to performing a transform in the horizontal and vertical directions on a two-dimensional block, while a non-separable transform refers to performing a single transform on samples that constitute the entire two-dimensional block or a partial region. A separable transform can be expressed as a pair of a horizontal transform kernel and a vertical transform kernel, and will be expressed as (horizontal transform kernel, vertical transform kernel) in this disclosure.
[0128] Multiple transform sets may be defined for the inverse transform of the current block, and each transform set may include one or more candidate transform kernels.
[0129] For example, any one of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), or (DCT-8, DCT-8) may be applied as a separate transform, and these four transform kernel candidates may be considered as one transform set. Also, (DCT-2, DCT-2) may be considered as one transform set. A transform skip that does not apply a transform may also be considered as one transform set, and (DCT-2, DCT-2) and the transform skip may be considered as one transform set. In the present disclosure, a transform kernel may represent one transform (e.g., DCT-2, DST-7) or a pair of two transforms (e.g., (DCT-2, DCT-2)).
[0130] Another example of a transform set may be the non-separable transform set described above. In the present disclosure, a non-separable transform applied as a primary transform may be referred to as an NSPT (Non-Separable Primary Transform). In an NSPT, multiple non-separable transform sets may be configured, and each non-separable transform set may include one or more transform kernels as transform kernel candidates. In the case of an NSPT, one of the multiple non-separable transform sets is selected depending on the intra prediction mode, and the multiple non-separable transform sets for an NSPT may be referred to as an NSPT set list. This has been described above, and a detailed description thereof will be omitted here.
[0131] A group of one or more transform sets available for the current block may be configured from a plurality of predefined transform sets. The group of one or more transform sets may be configured for a predetermined area unit to which the current block belongs, and hereinafter referred to as a collection. Here, the predetermined area unit may be at least one of a picture, a slice, a coding tree unit row (CTU row), or a coding tree unit (CTU).
[0132] For example, a transformation set consisting of (DCT-2, DCT-2) is called S1, and a transformation set consisting of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8) is called S2. In addition, the above-mentioned NSPT set list may include N non-separable transformation sets, and the N non-separable transformation sets are called S 3,1 , S 3,2 , ..., S 3,N Here, N may be 35, but is not limited to this.
[0133] S depending on the intra prediction mode of the current block 3,13 is selected as the non-separable transformation set for NSPT, the transformation kernels applicable to the current block are S1, S2, or S 3,13 In this case, the collection in which the block is currently available is defined as {S1, S2, S 3,13} can be written as
[0134] As described above, a collection according to the present disclosure is a group of one or more transform sets available to a current block, and the collection may be configured to vary depending on the context of the current block. Here, the context may include at least one of a shape, a size, or an intra-prediction mode. If a total of K contexts are defined, K collections may be generated, and each collection may include C i (i=1, 2, ..., N). For example, if the block sizes to which NSPT can be applied are 4x4, 8x8, 16x16, and 32x32, and one of a total of 35 non-separable transform sets is selected depending on the intra prediction mode, if different transform kernels are applied for each block size, a total of 4x35=140 contexts may be defined.
[0135] A collection may be constructed based on the context of the current block, and a process of selecting one of a plurality of transformation sets belonging to the collection and selecting one of a plurality of transformation kernel candidates belonging to the selected transformation set may be performed. Here, the selection of the transformation set and the transformation kernel candidate may be performed implicitly based on the context of the current block or based on an explicitly signaled index.
[0136] Alternatively, the process of selecting one of the plurality of transform sets belonging to the collection and the process of selecting one of the plurality of transform kernel candidates belonging to the selected transform set may be performed separately. For example, an index for selecting a transform set may be first signaled, and one of the plurality of transform sets belonging to the collection may be selected based on the index. Then, an index indicating one of the plurality of transform kernel candidates belonging to the transform set may be signaled, and one of the transform kernel candidates may be selected from the transform set based on the signaled index. A transform kernel for the current block may be determined based on the selected transform kernel candidate. Alternatively, the selection of one of the transform sets from the collection may be implicitly performed based on the context of the current block, and the selection of one of the transform kernel candidates from the selected transform set may be performed based on the signaled index. Alternatively, the selection of one of the transform sets from the collection may be performed based on the signaled index, and the selection of one of the transform kernel candidates from the selected transform set may be implicitly performed based on the context of the current block. Alternatively, the selection of any one transform set from the collection may be implicitly based on the context of the current block, and the selection of any one transform kernel candidate from the selected transform set may also be implicitly based on the context of the current block.
[0137] Of course, if the number of transform sets belonging to the collection is one, an index for selecting the transform set does not need to be signaled. Similarly, if the number of transform kernel candidates belonging to the selected transform set is one, an index for indicating the transform kernel candidate does not need to be signaled.
[0138] Alternatively, an index indicating any one of all transform kernel candidates currently belonging to the collection may be signaled. In this case, the process of selecting any one transform set from the collection may be omitted. In this case, all transform sets belonging to the collection may be rearranged in consideration of priority. For example, when assigning small-length binary codes to small-value indices, such as truncated unary codes, it may be advantageous to assign small-value indices to transform kernel candidates that are relatively advantageous in improving coding performance. When rearranging all transform kernel candidates belonging to a collection according to priority, different shuffling may be applied to each collection. Also, instead of rearranging all transform kernel candidates belonging to a collection, only some of them may be selectively rearranged.
[0139] Example 2
[0140] The transform kernel for the inverse transform of the current block may be determined based on multiple transform selection (MTS).
[0141] The MTS according to the present disclosure can use at least one of DST-7, DCT-8, DCT-5, DST-4, DST-1, or IDT (identity transform) as a transform kernel, and may further include a DCT-2 transform kernel.
[0142] In the present disclosure, multiple MTS sets for MTS may be defined. One of the multiple MTS sets may be determined based on the size of the current block and / or the intra prediction mode. For example, when determining one of the MTS sets, 16 transform block sizes may be considered, and for directional modes, symmetry between the transform block shape and the intra prediction mode may be considered. In the case of Wide Angle Intra Prediction (WAIP) modes (i.e., -1 to -14 (or -15), 67 to 80 (or 81)), the MTS set corresponding to mode 2 may be applied to modes -1 to -14 (or -15), and the MTS set corresponding to mode 66 may be applied to modes 67 to 80 (or 81). A separate MTS set may be assigned to a Matrix-based Intra Prediction (MIP) mode.
[0143] For example, the MTS set according to the transform block size and the intra prediction mode may be assigned / defined as shown in Table 4 below.
[0144] [Table 4]
[0145] Table 4 shows the allocation of MTS sets according to 16 transform block sizes and intra prediction modes. The number of predefined MTS sets is 80, and an index indicating one of the 80 MTS sets may range from 0 to 79 as shown in Table 4.
[0146] [Table 5-1] [Table 5-2]
[0147] Table 5 shows the transform kernel candidates included in each MTS set described in Table 4. Each MTS set may be composed of six transform kernel candidates. A transform kernel candidate index has a value of 0 to 5 and can indicate one of the six transform kernel candidates. Here, each transform kernel candidate may be a combination of a horizontal transform kernel and a vertical transform kernel for a separation transform, and 25 transform kernel candidates having indices of 0 to 24 may be defined.
[0148] [Table 6]
[0149] Table 6 shows an example of the 25 transform kernel candidates described in Table 5. Specifically, the horizontal transform and vertical transform of the transform kernel candidate are denoted as (horizontal transform, vertical transform). For each transform kernel candidate index, the horizontal / vertical transform when the intra prediction mode is less than 35 may be opposite to the horizontal / vertical transform when the intra prediction mode is 35 or greater. When the intra prediction mode value is 35 or greater, a mode symmetrical with respect to mode 34 may be derived, and an MTS set may be selected from Table 4 based on the derived mode. Symmetry of the block shape may also be taken into consideration. When an original transform block has a WxH size, the original transform block may be symmetrically considered to have an HxW size, and an MTS set may be selected from Table 4. Here, the value of the intra prediction mode may be the value of the modified intra prediction mode. That is, for the mode values for WAIP, -14 (or -15) to -1 are modified to mode 2, 67 to 80 (or 81) are modified to mode 66, and the original intra prediction mode values for the remaining modes can be set as modified intra prediction mode values. In this case, since the extended modes for WAIP are also configured symmetrically around mode 34, symmetry around mode 34 can be used for all directional modes except for planar mode and DC mode.
[0150] For example, when a 16x32 block is predicted as the 54th mode, the 14th mode (=68-54) is induced as a mode symmetric to the 54th mode, and the block size may be considered as 32x16. In this case, an MTS set having an index of 72 may be selected as defined in Table 4.
[0151] When the MIP mode is applied, the MTS set assigned to the MIP mode may be selected based on the size of the current block without considering the symmetry of the block type. Alternatively, when the MIP mode is applied, the MTS set assigned to the MIP mode may be selected based on the symmetric block size while considering the symmetry of the block type. For example, when the MIP mode is applied to an 8x16 block, the 8x16 block may be regarded as a 16x8 block symmetric to the 8x16 block, and the MTS set having an index of 49 may be selected as defined in Table 4. Alternatively, when the MIP mode is applied, the intra prediction mode may be regarded as the planar mode. In this case, the MTS set assigned to the MIP mode may be selected based on the size of the current block without considering the symmetry of the block type. Alternatively, the MTS set assigned to the MIP mode may be selected based on the symmetric block size while considering the symmetry of the block type.
[0152] In the case of the MIP mode, a flag indicating whether the MIP mode is applied as a transpose mode may be used. When the MIP mode is applied to an MxN current block and the flag indicates the application of the transpose mode, the intra prediction mode may be regarded as the planar mode, and the MxN current block may be regarded as an NxM block. That is, an MTS set corresponding to the NxM block size and the planar mode may be selected from Table 4. As described in Table 6, when the intra prediction mode value is 35 or greater, the horizontal transform and the vertical transform are interchanged. However, since the intra prediction mode of the current block is regarded as the planar mode, the horizontal transform and the vertical transform of the transform kernel candidate do not need to be interchanged. Alternatively, when the MIP mode is applied to an MxN current block and the flag indicates the application of the transpose mode, the intra prediction mode may not be regarded as the planar mode, and the MxN current block may be regarded as an NxM block. That is, an MTS set corresponding to the NxM block size and the MIP mode may be selected from Table 4.
[0153] In Table 5, a transform kernel candidate selected by a transform kernel candidate index may be set as the transform kernel for the current block. Alternatively, at least one of the horizontal transform or vertical transform of the selected transform kernel candidate may be changed to another transform kernel depending on the size of the current block. For example, if the transform kernel candidate index is 3 and both the width and height of the current block are 16 or less, at least one of the horizontal transform or vertical transform of the transform kernel candidate corresponding to the transform kernel candidate index of 3 may be changed to another transform kernel. In this case, the horizontal transform and the vertical transform may be changed independently of each other. If the difference (or the absolute value of the difference) between the intra prediction mode value and the horizontal mode value of the current block is equal to or less than a predetermined threshold, the vertical transform of the selected transform kernel candidate may be changed to IDT (identity transform). If the difference (or the absolute value of the difference) between the intra prediction mode value and the vertical mode value of the current block is equal to or less than a predetermined threshold, the horizontal transform of the selected transform kernel candidate may be changed to IDT (identity transform). Here, the threshold may be determined as shown in Table 7 below based on the width and height of the current block.
[0154] [Table 7]
[0155] Table 7 defines thresholds according to the size of the transform block for changing the horizontal transform and / or vertical transform of the transform kernel candidate selected by the transform kernel candidate index to another transform kernel.
[0156] Six transform kernel candidates constituting one MTS set may be distinguished by transform kernel candidate indices 0 to 5 as defined in Table 5. The transform kernel candidate indices may be signaled by a bitstream. A flag indicating whether an MTS set is available / applied (MTS enabled flag or MTS flag) may be signaled, and if the flag indicates that an MTS set is available / applied, the transform kernel candidate index may be signaled. The MTS flag may be configured as one bin, and one or more contexts for CABAC-based entropy coding (hereinafter referred to as CABAC contexts) may be assigned to the bin. For example, different CABAC contexts may be assigned to non-MIP mode and MIP mode, respectively.
[0157] The number of transform kernel candidates available to the current block may be set differently depending on the context of the current block. For example, the sum of absolute values of all or some of the transform coefficients in the current block may be considered as the context of the current block. The sum of absolute values of the transform coefficients is referred to as AbsSum. If AbsSum is less than or equal to T1, only one transform kernel candidate corresponding to a transform kernel candidate index of 0 may be available. If AbsSum is greater than T1 and less than or equal to T2, four transform kernel candidates corresponding to transform kernel candidate indexes of 0 to 3 may be available. If AbsSum is greater than T2, six transform kernel candidates corresponding to transform kernel candidate indexes of 0 to 5 may be available. Here, T1 may be 6 and T2 may be 32, but this is merely an example.
[0158] When AbsSum is less than or equal to T1, the number of transform kernel candidates available for the current block is one, and therefore, the transform kernel candidate corresponding to the transform kernel candidate index of 0 may be set as the transform kernel for the current block without signaling the transform kernel candidate index. When AbsSum is greater than T1 and less than or equal to T2, four transform kernel candidates are available, and therefore, one of the four transform kernel candidates may be selected based on the transform kernel candidate index having two bins. That is, the transform kernel candidate indexes 0 to 3 may be signaled as 00, 01, 10, and 11, respectively. For the two bins, the most significant bit (MSB) may be signaled first, and the least significant bit (LSB) may be signaled later. A different CABAC context may be assigned to each bin. For example, a CABAC context other than the CABAC context assigned for the MTS flag may be assigned to each of the two bins. Alternatively, no CABAC context may be assigned to two bins, and bypass coding may be applied. When AbsSum is greater than T2, the transform kernel candidate indexes have values between 0 and 5, so the transform kernel candidate indexes cannot be expressed using only two bins. In this case, two or more bins may be assigned to express the transform kernel candidate indexes, as in truncated binary coding. A CABAC context may be assigned to each bin assigned by the truncated binary coding scheme, or bypass coding may be applied without assigning a CABAC context. Alternatively, a CABAC context may be assigned to some of the bins (e.g., the first bin or the first and second bins), and bypass coding may be applied to the remaining bins.
[0159] Example 3
[0160] A transformation kernel for a current block may be determined based on a transformation set including one or more transformation kernel candidates, and the transformation kernel for the current block may be derived from any one of the one or more transformation kernel candidates belonging to the transformation set.
[0161] The process of determining a transform kernel for the current block may include at least one of 1) determining a transform set for the current block, or 2) selecting one transform kernel candidate from the transform set for the current block. The process of determining the transform set may be a process of selecting one of a plurality of transform sets predefined identically for the encoding device and the decoding device. Alternatively, the process of determining the transform set may be a process of constructing one or more transform sets available for the current block from a plurality of transform sets predefined identically for the encoding device and the decoding device, and selecting one of the constructed transform sets. Alternatively, the process of determining the transform set may be a process of constructing one transform set based on a transform kernel candidate available for the current block from a plurality of transform kernel candidates predefined identically for the encoding device and the decoding device.
[0162] When a transform set of a current block includes multiple transform kernel candidates, a process of selecting one of the multiple transform kernel candidates may be performed for the current block. However, when the transform set of the current block includes one transform kernel candidate (i.e., when there is one transform kernel candidate available for the current block), the transform kernel of the current block may be set to the selected transform kernel candidate.
[0163] The transform set according to the present disclosure may refer to the (non-separable) transform set in the first embodiment or the MTS set in the second embodiment. Alternatively, the transform set may be defined separately from the (non-separable) transform set in the first embodiment or the MTS set in the second embodiment. In this case, the transform set may include one or more specific transform kernels as candidate transform kernels. One specific transform kernel may be defined as a pair of a transform kernel for a horizontal transform and a transform kernel for a vertical transform, or as one transform kernel that is applied equally to both horizontal and vertical transforms. The specific transform kernel will be described in detail below.
[0164] The specific transform kernel according to the present disclosure may be a transform kernel predefined in the same way for both the encoding device and the decoding device, or may further include a transform kernel derived based on the predefined transform kernel, or may refer to a transform kernel having a predetermined index in the (non-separable) transform set of the first embodiment or the MTS set of the second embodiment.
[0165] The specific transform kernel may be defined as a combination of trigonometric function-based transform kernels (e.g., DCT-2, DST-7, DCT-8, DCT-5, DST-4, and DST-1). Alternatively, the specific transform kernel may be defined as a combination of non-trigonometric function-based transform kernels. Examples of non-trigonometric function-based transform kernels include KLT, SOT, orthogonal transform kernel, and non-orthogonal transform kernel. The KLT may be a transform kernel trained with training feature data, i.e., a training-based transform kernel. Alternatively, the specific transform kernel may be defined as a combination of a trigonometric function-based transform kernel and a non-trigonometric function-based transform kernel.
[0166] For example, a specific transformation kernel is represented as (T_h, T_v). T_h is a horizontal transformation kernel and may be a training-based transformation kernel such as KLT. T_v is a vertical transformation kernel and may be a trigonometric function-based transformation kernel such as DCT-2. Alternatively, T_h may be KLT and T_v may be DST-7. Alternatively, when DST-7 and DCT-2 are allowed as T_h and KLT1 and KLT2 corresponding to KLT are allowed as T_v, four specific transformation kernels may be defined as combinations of allowable transformation kernels.
[0167] The specific transform kernel or a transform set based on the specific transform kernel may be used in a manner that replaces the (non-separable) transform set of Example 1 and / or the MTS set of Example 2. Alternatively, the specific transform kernel or a transform set based on the specific transform kernel may be added as a transform set independent of the (non-separable) transform set of Example 1 and / or the MTS set of Example 2.
[0168] A flag indicating whether a transform set based on a specific transform kernel is to be applied may be defined. When the flag is a first value, a transform set based on a specific transform kernel may be applied. When the flag is a second value, a (non-separable) transform set of Example 1 or an MTS set of Example 2 may be applied. The flag may be signaled before a syntax element specifying at least one transform kernel candidate in the transform set. For example, when the flag is a first value, a transform set based on a specific transform kernel may be applied, and an index indicating one of multiple specific transform kernels belonging to the transform set may be further signaled. On the other hand, when the flag is a second value, a (non-separable) transform set of Example 1 or an MTS set of Example 2 may be applied, and an index indicating one of multiple transform kernel candidates (i.e., a transform kernel candidate index) may be further signaled.
[0169] The particular transformation kernel may be applied only to the luminance component of the current block, or may be applied to both the luminance and chrominance components of the current block.
[0170] A specific transform kernel according to the present disclosure may be defined as a length-based transform kernel. That is, a specific transform kernel may be one or more transform kernels having a predetermined length. Hereinafter, for convenience of explanation, a transform kernel having a length of K will be referred to as a length-K transform kernel. Here, K may be at least one of an integer of 4, 8, 16, 32, 64, or greater. Here, the length may refer to the length of one side (width and / or height) that a transform block may have. In the present disclosure, the width that a transform block may have may be referred to as an allowable width(s), and the height that a transform block may have may be referred to as an allowable height(s). When a separation transform is applied to a current block of size MxN, a length-M transform kernel having a length equal to the width of the current block may be applied horizontally, and a length-N transform kernel having a length equal to the height of the current block may be applied vertically.
[0171] A separate transform can be applied to MxN blocks of all sizes by assigning transform kernels according to length. For example, if the current block is 4xN, a length-4 transform kernel having the same length as the width of the current block in the horizontal direction may be applied. If the current block is Nx4, a length-4 transform kernel having the same length as the height of the current block in the vertical direction may be applied. Here, N may be an integer of 4, 8, 16, 32, 64, or more.
[0172] The length-based transform kernels may be configured to be different in the horizontal and vertical directions. That is, a length-4 horizontal transform kernel and a length-4 vertical transform kernel may be different from each other. If the number of allowable widths of the transform blocks is P and the number of allowable widths of the transform blocks is Q, one transform set may be configured with (P+Q) length-based transform kernels. For example, if the allowable widths and allowable heights of the transform blocks are 4, 8, 16, and 32, respectively, a length-4 transform kernel, a length-8 transform kernel, a length-16 transform kernel, and a length-32 transform kernel may be available in the horizontal and vertical directions. In this case, the P and Q values are each 4, and one transform set may be configured based on a total of eight length-based transform kernels. In this case, the eight length-based transform kernels may be a length-4 horizontal transform kernel, a length-8 horizontal transform kernel, a length-16 horizontal transform kernel, a length-32 horizontal transform kernel, a length-4 vertical transform kernel, a length-8 vertical transform kernel, a length-16 vertical transform kernel, and a length-32 vertical transform kernel.
[0173] Even if transform kernels of the same length are applied in the same direction, the transform kernel applied to the current block may differ depending on the size of the current block. Here, the size of the current block may be defined as any one of the width, height, sum of the width and height, product of the width and height, or maximum / minimum value of the width and height of the current block. A length-M horizontal transform kernel may be assigned to each allowable height of the transform block. Similarly, a length-N vertical transform kernel may be assigned to each allowable width of the transform block. Here, if M and N can each have values of 4, 8, 16, and 32, different length-M horizontal transform kernels may be assigned to Mx4, Mx8, Mx16, and Mx32 blocks, respectively, and different length-N vertical transform kernels may be assigned to 4xN, 8xN, 16xn, and 32xN blocks, respectively. For example, a length-4 horizontal transform kernel applied to a 4x8 block may be different from a length-4 horizontal transform kernel applied to a 4x16 block.
[0174] Alternatively, the allowable widths and allowable heights of the transform blocks may be divided into a plurality of groups, and the same length-based transform kernel may be assigned to each group.
[0175] For example, the allowable widths and allowable heights of a transform block may be divided into two groups based on a predetermined threshold. For an MxN block, if the value of N is less than or equal to a first threshold, a first length-M horizontal transform kernel may be assigned, and if the value of N is greater than the first threshold, a second length-M horizontal transform kernel may be assigned. Here, the first length-M horizontal transform kernel and the second length-M horizontal transform kernel may have the same length but be different transform kernels. The first threshold may be an integer of 8, 16, or greater. Similarly, for an MxN block, if the value of M is less than or equal to a second threshold, a first length-N vertical transform kernel may be assigned, and if the value of M is greater than the second threshold, a second length-N vertical transform kernel may be assigned. Here, the first length-N vertical transform kernel and the second length-N vertical transform kernel may have the same length but be different transform kernels. The second threshold may be an integer of 8, 16, or greater.
[0176] Alternatively, the allowable widths and heights of the transform blocks may be divided into three or more groups based on at least two thresholds. Assume that the allowable widths and heights of the transform blocks are 4, 8, 16, and 32, and are divided into three groups, i.e., {4, 8}, {16}, and {32}. For an MxN current block, if N is 4 or 8, the first length-M horizontal transform kernel may be assigned. If N is 16, the second length-M horizontal transform kernel may be assigned. If N is 32, the third length-M horizontal transform kernel may be assigned. Similarly, if M is 4 or 8, the first length-N vertical transform kernel may be assigned. If M is 16, the second length-N vertical transform kernel may be assigned. If M is 32, the third length-N vertical transform kernel may be assigned. In this way, one of the first to third length-M horizontal transform kernels may be applied depending on the height of the current block, and one of the first to third length-N vertical transform kernels may be applied depending on the width of the current block. Nine specific transform kernels may be configured by combining the first to third length-M horizontal transform kernels and the first to third length-N vertical transform kernels, and one transform set for the current block may be configured with the nine specific transform kernels.
[0177] Alternatively, length-specific transform kernels may not be configured separately for the horizontal and vertical directions. That is, a transform kernel having a length corresponding to the width or height may be applied regardless of the horizontal or vertical direction. For example, a length-4 transform kernel applied to the horizontal direction of a 4x8 block may be the same as a length-4 transform kernel applied to the vertical direction of an 8x4 block.
[0178] Regardless of the horizontal and vertical directions, when a length-M transform kernel having the same length as a side of length M is applied, only one transform kernel may be required for each allowable width and / or allowable height of the transform block. For example, when the allowable widths or heights of the transform block are 4, 8, 16, and 32, one transform set may be configured based on four length-based transform kernels. Here, the four length-based transform kernels may be a length-4 transform kernel, a length-8 transform kernel, a length-16 transform kernel, and a length-32 transform kernel. When the current block is an MxN block, a length-M transform kernel may be applied in the horizontal direction, and a length-N transform kernel may be applied in the vertical direction.
[0179] The application of the length-based transform kernel may be limited depending on the context of the current block, or the number of length-based transform kernels available to the current block may differ. For convenience of explanation, it is assumed that the length-based transform kernel is a non-trigonometric function-based transform kernel such as the KLT described above, but is not limited thereto.
[0180] A non-trigonometric function-based transform kernel may be applied to a direction of an edge having a specific length, and a trigonometric function-based transform kernel may be applied to other lengths. Here, the specific length may be determined based on a length that is predefined identically for the encoding device and the decoding device. A non-trigonometric function-based transform kernel may be applied to one of the horizontal and vertical directions, and a trigonometric function-based transform kernel may be applied to the other direction.
[0181] For example, if the length of one side of the current block is greater than or equal to 8, a non-trigonometric function-based transformation kernel may be applied to that side; otherwise, a trigonometric function-based transformation kernel may be applied to that side. Alternatively, if the length of one side of the current block is less than 8, a non-trigonometric function-based transformation kernel may be applied to that side; otherwise, a trigonometric function-based transformation kernel may be applied to that side. Alternatively, if the length of one side of the current block is less than or equal to 16, a non-trigonometric function-based transformation kernel may be applied to that side; otherwise, a trigonometric function-based transformation kernel may be applied to that side. Alternatively, if the length of one side of the current block is greater than 16, a non-trigonometric function-based transformation kernel may be applied to that side; otherwise, a trigonometric function-based transformation kernel may be applied to that side. In this way, in the case where a non-trigonometric function-based transformation kernel is restricted not to be applied to an edge whose length is p, the following two methods may be considered for a block whose width or height is p and the other is not p.
[0182] (1) A first method in which a non-trigonometric function-based transformation kernel is not applied to a direction having a length p in the width or height, and a non-trigonometric function-based transformation kernel is applied to a direction not having a length p.
[0183] (2) A second method in which a non-trigonometric function-based transformation kernel is not applied to the horizontal and vertical directions when either the width or the height has a length of p.
[0184] For example, length-4 non-trigonometric function based transform kernels may be restricted from being applied.
[0185] According to the first method, for a block whose width or height is 4, a length-4 non-trigonometric function-based transform kernel may not be applied to the direction of a side whose length is 4, but a length-4 trigonometric function-based transform kernel may be applied. On the other hand, a non-trigonometric function-based transform kernel corresponding to the length may be applied to the direction of a side whose length is not 4. That is, a length-4 non-trigonometric function-based transform kernel may not be applied to a 4x4 block. For a 4x8, 4x16, 4x32, 8x4, 16x4, or 32x4 block, a non-trigonometric function-based transform kernel corresponding to the length may be applied only to a side whose length is not 4. For blocks of the remaining sizes (e.g., 8x8, 8x16, 8x32, 16x8, 16x16, 16x32, 32x8, 32x16, 32x32, etc.), a non-trigonometric function-based transform kernel corresponding to the length may be applied to the horizontal and vertical directions.
[0186] Alternatively, according to the second method, a non-trigonometric function-based transformation kernel corresponding to the length may be applied only when both the width and height of the current block are greater than 4. In other words, when either the width or height has a length of 4, a length-4 non-trigonometric function-based transformation kernel may not be applied to the horizontal and vertical directions.
[0187] Alternatively, length-4 and length-8 non-trigonometric function based transform kernels may be restricted from being applied.
[0188] According to the first method, for 4x4, 4x8, 8x4, and 8x8 blocks, length-4 and length-8 non-trigonometric function-based transform kernels may not be applied, but length-4 and length-8 trigonometric function-based transform kernels may be applied. For 4x16, 4x32, 8x16, 8x32, 16x4, 16x8, 32x4, and 32x8 blocks, non-trigonometric function-based transform kernels corresponding to the lengths may be applied only to edges whose lengths are not 4 or 8. For blocks of the remaining sizes (e.g., 16x16, 16x32, 32x16, 32x32, etc.), non-trigonometric function-based transform kernels corresponding to the lengths may be applied in the horizontal and vertical directions.
[0189] According to the second method, only when both the width and height of the current block are greater than 8 (e.g., 16x16, 16x32, 32x16, or 32x32 blocks), the non-trigonometric function-based transformation kernel corresponding to the length may be applied. In other words, when either the width or height is 4 or 8, the length-4 and length-8 non-trigonometric function-based transformation kernels may not be applied to the horizontal and vertical directions.
[0190] Alternatively, length-32 non-trigonometric function based transform kernels may be restricted to not be applied.
[0191] According to the first method, a length-32 non-trigonometric function-based transformation kernel may not be applied to a 32x32 block, but a length-32 trigonometric function-based transformation kernel may be applied. For 4x32, 8x32, 16x32, 32x4, 32x8, and 32x16 blocks, a non-trigonometric function-based transformation kernel corresponding to the length may be applied only to edges whose length is not 32. For blocks of the remaining sizes (e.g., 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, and 16x16), a non-trigonometric function-based transformation kernel corresponding to the length may be applied in the horizontal and vertical directions.
[0192] According to the second method, only when both the width and height of the current block are smaller than 32, a non-trigonometric function-based transformation kernel corresponding to the length may be applied. That is, for 4x32, 8x32, 16x32, 32x4, 32x8, 32x16, or 32x32 blocks, a length-32 non-trigonometric function-based transformation kernel may not be applied in the horizontal and vertical directions. For blocks of the remaining sizes (e.g., 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, 16x16), a non-trigonometric function-based transformation kernel corresponding to the length may be applied in the horizontal and vertical directions.
[0193] The above-described embodiments may be adaptively performed based on the slice type of the current block. For example, the above examples may be applied when the slice type is an I-slice, but may not be applied when the slice type is not an I-slice.
[0194] Furthermore, in the above-described embodiments, the application of a non-trigonometric function-based transform kernel to a specific length is restricted. However, this is merely an example. For example, a first type of trigonometric function-based transform kernel may be restricted from being applied to a specific length, and in this case, a second type of trigonometric function-based transform kernel may be applied. That is, in the above-described embodiments, the non-trigonometric function-based transform kernel and the trigonometric function-based transform kernel may be understood as a first type of trigonometric function-based transform kernel and a second type of trigonometric function-based transform kernel, respectively. The second type may be a different transform type from the first type.
[0195] Example 4
[0196] A transform set may be determined based on the context of the current block. The transform set may be determined based on the intra-prediction mode of the current block and may be composed of one or more transform kernel candidates. For example, the transform set of the present disclosure may be the (non-separable) transform set of Example 1 or the MTS set of Example 2. The transform set of the present disclosure may include one or more length-based transform kernels having a predetermined length as the transform kernel candidates, as in Example 3.
[0197] In the present disclosure, a mapping relationship between predefined intra prediction modes and transform sets may be defined.
[0198] The predefined intra prediction modes may include at least one of a planar mode, a DC mode, a directional mode, a mode for WAIP (wide angle modes), a matrix-based intra prediction (MIP) mode, a decoder-side intra mode derivation (DIMD) mode, or a template-based intra mode derivation (TIMD) mode.
[0199] Each mode is assigned a mode value to identify it: 0 and 1 for planar and DC modes, 2 through 66 for directional modes, and values below 0 and above 66 for wide angle modes.
[0200] A separate transform set may be mapped for each of the predefined intra prediction modes (or for each mode value). Alternatively, the predefined intra prediction modes may be divided into at least two groups according to a predetermined rule, and a transform set may be mapped for each group. Here, the predetermined rule may be any one of rules (1) to (7) described below, or a combination of at least two of them. The predetermined rule for grouping the predefined intra prediction modes will now be described.
[0201] (1) Planar mode and DC mode can be grouped together.
[0202] (2) Directional modes may be grouped together by grouping adjacent modes.
[0203] (3) Wide-angle modes can be grouped together. Alternatively, modes with values less than 0 can be grouped together in one group, and modes with values greater than 66 can be grouped together in another group. Alternatively, wide-angle modes can be divided into three or more groups, with adjacent modes grouped together, like directional modes. For example, wide-angle modes less than 0 are adjacent to mode 2 in the angle, so wide-angle modes less than 0 and mode 2 (or mode 2 and at least one directional mode adjacent to mode 2) can be grouped together (Group A). Wide-angle modes greater than 66 are adjacent to mode 66 in the angle, so wide-angle modes greater than 66 and mode 66 (or mode 66 and at least one directional mode adjacent to mode 66) can be grouped together (Group B). Groups A and B can also be merged into a single group.
[0204] (4) MIP modes may be grouped together with planar modes, or may form a separate group for MIP modes only, or may be considered as planar modes when mapping between intra prediction modes and transform sets.
[0205] (5) In the case of a DIMD mode, a separate group may be formed only for the DIMD mode, or may be included in a group to which the intra prediction mode induced by the DIMD mode belongs. Alternatively, if multiple intra prediction modes are induced by the DIMD mode, the group may be included in a group to which one of the multiple intra prediction modes (e.g., the first mode, the mode with the smallest value) belongs.
[0206] (6) In the case of the TIMD mode, a separate group may be formed only for the TIMD mode, or the TIMD mode may be included in a group to which the intra prediction mode induced by the TIMD mode belongs. Alternatively, if multiple intra prediction modes are induced by the TIMD mode, the TIMD mode may be included in a group to which one of the multiple intra prediction modes (e.g., the first mode, the mode with the smallest value) belongs. The intra prediction mode values induced by the TIMD mode may be given based on 131 modes. In this case, the mode values based on 131 modes may be converted to mode values based on 67 modes. Any one of values 0 to 130 may be mapped / converted to any one of values 0 to 66 according to a mapping function (MAP131TO67) according to the following Equation 4. The TIMD mode may be included in a group to which the changed mode value belongs.
[0207] [Formula 4]
number
[0208] (7) Intra template matching mode can be included in the group to which the planar mode belongs. The intra template matching mode may be considered as the planar mode when mapping between the intra prediction mode and the transform set.
[0209] According to the above-mentioned predetermined rule, the predefined intra prediction modes may be divided into a plurality of groups, and a transform set may be assigned to each group. The following is an example of a mapping table that defines the transform sets assigned to each intra prediction mode (or each group).
[0210] [Table 8]
[0211] [Table 9]
[0212] [Table 10]
[0213] [Table 11]
[0214] [Table 12]
[0215] The transform set may be determined based on symmetry between intra prediction modes and / or block types. As described with reference to FIG. 5, intra prediction modes, particularly directional modes, may be symmetrical around the 34th mode. Due to the symmetry between intra prediction modes based on the 34th mode, for an MxN block, the mode symmetrical to the pth mode may be the (68-p) mode, and the mode symmetrical to the qth mode, which is a wide-angle mode, may be the (66-q) mode. Furthermore, due to the symmetry between block types, the block symmetrical to the MxN block may be an NxM block.
[0216] If the mode corresponding to the rth mode is the sth mode, then intra prediction performed in the rth mode for an MxN block and intra prediction performed in the sth mode for an NxM block may be symmetrical to each other. Alternatively, application of a length-M horizontal transform kernel and a length-N vertical transform kernel to an MxN block may be symmetrical to application of a length-M vertical transform kernel and a length-N horizontal transform kernel to an NxM block. If the same specific transform kernel is to be applied to the two cases, the length-M horizontal transform kernel for the MxN block and the length-M vertical transform kernel for the NxM block may be set to be the same for the two cases, and the length-N vertical transform kernel for the MxN block and the length-N horizontal transform kernel for the NxM block may be set to be the same.
[0217] For example, if the intra prediction mode for an MxN block is a directional mode and a pth mode greater than 34, a transform set mapped to a (68-p)th mode for an NxM block may be used. In this case, a length-M vertical transform kernel and a length-N horizontal transform kernel, which are transform kernel candidates belonging to the transform set, may be used as the length-M horizontal transform kernel and the length-N vertical transform kernel of the MxN block, respectively. When the above-described symmetry is utilized, the same transform set is used between mutually symmetric modes, and the mapping between the intra prediction mode and the transform set can be set symmetrically as follows:
[0218] [Table 13]
[0219] [Table 14]
[0220] [Table 15]
[0221] In the present disclosure, a transform set may include one or more transform kernel candidates, and the candidates selected from among these may include a horizontal transform kernel and a vertical transform kernel. The transform kernel candidates may be assigned to respective lengths of the sides of the transform block. In this case, the transform kernel for the MxN block predicted with the rth mode may be derived based on the symmetry between the intra prediction modes and / or block types described above. That is, a transform set may be determined based on the sth mode, which is symmetric to the rth mode, and a transform kernel for the current block may be derived from the transform kernel candidates belonging to the transform set. In this case, a length-N horizontal transform kernel from the transform kernel candidates selected for an NxM block symmetric to the MxN block may be applied as a length-N vertical transform kernel for the MxN block. Similarly, a length-M vertical transform kernel from the transform kernel candidates selected for an NxM block symmetric to the MxN block may be applied as a length-M horizontal transform kernel for the MxN block. However, as described in Example 3, when a transform set is composed of length-based transform kernels that are applied equally in the horizontal and vertical directions, the transform kernels are not changed while the horizontal transform kernel is changed to a vertical transform kernel.
[0222] In the case of MIP mode, a flag indicating whether prediction is performed in a transposed form may be defined, and when the flag is 0, the flag may be symmetrical to when the flag is 1. When the flag is 0, a length-M horizontal transform kernel applied in the horizontal direction of an MxN block may be set to be the same as a length-M vertical transform kernel applied in the vertical direction of an NxM block when the flag is 1.
[0223] Alternatively, if the flag is 1, the horizontal and vertical transform kernels applied to the MxN block may be set to the horizontal and vertical transform kernels of the NxM block, respectively. For example, if DST-7 and DCT-5 are applied as the horizontal and vertical transform kernels for the NxM block, respectively, the horizontal and vertical transform kernels applied to the MxN block may be set to DST-7 and DCT-5, respectively. Although the transform types are the same, the lengths of the transform kernels may be different. That is, a length-N DST-7 may be applied as the horizontal transform kernel for the NxM block, while a length-M DST-7 may be applied as the horizontal transform kernel for the MxN block.
[0224] When a non-trigonometric function-based transform kernel (e.g., a training-based transform kernel) is applied instead of a trigonometric function-based transform kernel for the MIP mode, the only difference is that a non-trigonometric function-based transform kernel is applied, but a transform kernel of a different length may be applied, and different non-trigonometric function-based transform kernels may be applied in different directions to MxN blocks and NxM blocks. In the MIP mode, the horizontal transform kernel and vertical transform kernel specified in the transform set can be used as is, regardless of whether transposition is applied or not.
[0225] The transformation kernel of the current block may be determined based on any one of the above-described Examples 1 to 4. Alternatively, the transformation kernel of the current block may be determined based on a combination of at least two of the above-described Examples 1 to 4, as long as the inventions according to the above-described Examples 1 to 4 do not conflict with each other.
[0226] Referring to FIG. 4, the current block can be reconstructed based on the residual samples of the current block (S420).
[0227] A prediction sample of the current block may be derived based on the intra prediction mode of the current block, and a reconstructed sample of the current block may be generated based on the prediction sample and residual sample of the current block.
[0228] FIG. 6 shows a schematic configuration of a decoding device 300 that performs the image decoding method according to the present disclosure.
[0229] 6, a decoding apparatus 300 according to the present disclosure may include a transform coefficient derivation unit 600, a residual sample derivation unit 610, and a reconstructed block generation unit 620. The transform coefficient derivation unit 600 may be configured in the entropy decoding unit 310 of FIG. 3, the residual sample derivation unit 610 may be configured in the residual processing unit 320 of FIG. 3, and the reconstructed block generation unit 620 may be configured in the adder 340 of FIG. 3.
[0230] The transform coefficient deriving unit 600 can obtain residual information of the current block from the bitstream and decode it to derive transform coefficients of the current block.
[0231] The residual sample derivation unit 610 may derive residual samples of the current block by performing at least one of inverse quantization and inverse transformation on the transform coefficients of the current block.
[0232] The residual sample deriving unit 610 determines a transform kernel for inverse transform of the current block using a predetermined transform kernel determination method, and derives a residual sample of the current block based on the determined transform kernel. The transform kernel determination method has been described with reference to FIG. 4, and a detailed description thereof will be omitted here.
[0233] The reconstructed block generator 620 can reconstruct the current block based on the residual samples of the current block.
[0234] FIG. 7 is a diagram illustrating an image encoding method performed by an encoding device 200 according to an embodiment of the present disclosure.
[0235] Referring to FIG. 7, residual samples of the current block can be derived (S700).
[0236] Residual samples of the current block may be derived by subtracting prediction samples from original samples of the current block, where the prediction samples may be derived based on a predetermined intra prediction mode.
[0237] Referring to FIG. 7, transform coefficients of the current block can be derived by performing at least one of transformation and quantization on residual samples of the current block (S710).
[0238] The method for determining the transformation kernel for the transformation is the same as that described with reference to Fig. 4, and detailed description thereof will be omitted here. That is, the transformation kernel for the transformation can be determined based on at least one of the above-described first to fourth embodiments.
[0239] For example, as in Example 1, one or more transform sets for transforming the current block may be defined / configured, and each transform set may include one or more transform kernel candidates. In this case, one of the multiple transform sets may be selected as the transform set for the current block. One of the multiple transform kernel candidates belonging to the transform set for the current block may be selected. The selection may be made implicitly based on the context of the current block. Alternatively, the optimal transform set and / or transform kernel candidate for the current block may be selected, and an index indicating this may be signaled.
[0240] Alternatively, as in Example 2, the transform kernel for the current block may be determined based on an MTS set. One of a plurality of MTS sets may be selected based on at least one of the size of the current block or the intra prediction mode. The selected MTS set may include one or more transform kernel candidates. One of the one or more transform kernel candidates may be selected, and the transform kernel for the current block may be determined based on the selected transform kernel candidate. The selection of the transform kernel candidate may be performed using a transform kernel candidate index derived based on the context of the current block. Alternatively, the optimal transform kernel candidate for the current block may be selected, and a transform kernel candidate index indicating the selected transform kernel candidate may be signaled.
[0241] Alternatively, as in the third embodiment, the transformation kernel for the current block may be determined based on a transformation set consisting of one or more specific transformation kernels.
[0242] Alternatively, as in Example 4, the transform kernel of the current block may be determined based on a mapping table that defines a mapping relationship between intra prediction modes and transform sets. Here, the mapping table may be defined in consideration of symmetry between intra prediction modes and / or block types.
[0243] Alternatively, the transformation kernel for the current block may be determined based on a combination of at least two of the first to fourth embodiments.
[0244] Referring to FIG. 7, the transform coefficients of the current block can be coded to generate a bitstream (S720).
[0245] Residual information regarding the transform coefficients may be generated based on the transform coefficients of the current block, and the residual information may be coded to generate a bitstream.
[0246] FIG. 8 shows a schematic configuration of an encoding device 200 that performs the image encoding method according to the present disclosure.
[0247] 8, the encoding apparatus 200 according to the present disclosure may include a residual sample derivation unit 800, a transform coefficient derivation unit 810, and a transform coefficient encoding unit 820. The residual sample derivation unit 800 and the transform coefficient derivation unit 810 may be configured in the residual processing unit 230 of FIG. 2, and the transform coefficient encoding unit 820 may be configured in the entropy encoding unit 240 of FIG. 2.
[0248] The residual sample deriving unit 800 may derive residual samples of the current block by subtracting predicted samples from original samples of the current block, where the predicted samples may be derived based on a predetermined intra prediction mode.
[0249] The transform coefficient deriving unit 810 may derive transform coefficients of the current block by performing at least one of transforming and quantizing on residual samples of the current block. The transform coefficient deriving unit 810 may determine a transform kernel for the current block based on any one of the first to fourth embodiments or a combination of at least two of the first to fourth embodiments, and derive transform coefficients by applying the transform kernel to residual samples of the current block.
[0250] The transform coefficient encoding unit 820 may encode the transform coefficients of the current block to generate a bitstream.
[0251] In the above-described embodiments, the method is described based on a flowchart with a series of steps or blocks, but the embodiment is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of the embodiments of this document.
[0252] The methods according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in an image processing device such as a TV, a computer, a smartphone, a set-top box, or a display device.
[0253] When embodiments in this document are embodied as software, the methods described above may be embodied as modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor by various known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be embodied and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures may be embodied and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for the implementation may be stored on a digital storage medium.
[0254] In addition, the decoding device and encoding device to which the embodiments of the present specification are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a custom video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, an image telephone video device, a vehicle terminal (e.g., a vehicle terminal (including an autonomous vehicle), an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process video signals or data signals. For example, over-the-top (OTT) video (over-the-top) video devices may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0255] In addition, a processing method to which the embodiments of the present specification are applied may be produced in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of the present specification may also be stored in a computer-readable recording medium. The computer-readable recording medium may include any type of storage device or distributed storage device in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium may also include media embodied in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0256] Furthermore, the embodiments of the present specification may be embodied as a computer program product using program code, which may be executed by a computer according to the embodiments of the present specification. The program code may be stored on a computer-readable carrier.
[0257] FIG. 9 illustrates an example of a content streaming system to which the embodiments of the present disclosure can be applied.
[0258] Referring to FIG. 9, a content streaming system to which the embodiments of the present specification are applied may broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0259] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0260] The bitstream may be generated by an encoding method or a bitstream generation method to which the embodiments of this specification are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0261] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0262] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0263] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, and head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signage.
[0264] Each server in the content streaming system may be operated as a distributed server, in which case data received by each server may be processed in a distributed manner.
[0265] The claims described herein may be combined in various ways. For example, the technical features of the method claims herein may be combined and embodied as an apparatus, and the technical features of the apparatus claims herein may be combined and embodied as a method. Furthermore, the technical features of the method claims herein and the technical features of the apparatus claims herein may be combined and embodied as an apparatus, and the technical features of the method claims herein and the technical features of the apparatus claims herein may be combined and embodied as a method.
Claims
1. deriving transform coefficients of a current block from a bitstream; performing at least one of inverse quantization and inverse transformation on the transform coefficients of the current block to derive residual samples of the current block; reconstructing the current block based on residual samples of the current block; The step of deriving a residual sample comprises: determining a transform set for a primary inverse transform of the current block, the transform set being determined based on at least one of an intra prediction mode of the current block or a predefined mapping table; deriving horizontal and vertical transform kernels for the current block from the transform set based on the width and height of the current block, respectively.
2. The image decoding method of claim 1 , wherein the mapping table defines a mapping relationship between predefined intra-prediction modes and transform sets.
3. The image decoding method of claim 1 , wherein the transform set is mapped to an intra prediction mode that has a symmetric relationship with the intra prediction mode of the current block.
4. the transformation set includes a plurality of candidate transformation kernels; The image decoding method of claim 1 , wherein the plurality of transform kernel candidates comprises length-based transform kernel candidates that are applied equally in horizontal and vertical directions.
5. The horizontal transformation kernel is set to a transformation kernel candidate having the same length as the width of the current block; The image decoding method of claim 4 , wherein the vertical transform kernel is set to a transform kernel candidate having the same length as a height of the current block.
6. If the current block is an MxN block, 2. The image decoding method of claim 1, wherein a horizontal transform kernel having a length of M for the current block is the same as a vertical transform kernel having a length of M for an NxM block.
7. 2. The image decoding method of claim 1, wherein the number of available transform sets is 1, 2, 3, 4, or 7.
8. deriving residual samples of a current block; performing at least one of transforming and quantizing on residual samples of the current block to derive transform coefficients of the current block; encoding the transform coefficients of the current block; The step of deriving the transformation coefficients comprises: determining a transform set for a primary inverse transform of the current block, the transform set being determined based on at least one of an intra prediction mode of the current block or a predefined mapping table; deriving horizontal and vertical transform kernels for the current block from the transform set based on the width and height of the current block, respectively.
9. A computer-readable storage medium storing a bitstream generated by the image encoding method of claim 8.
10. obtaining a bitstream for image information, the bitstream being generated by deriving residual samples of a current block, performing at least one of transforming and quantizing on the residual samples of the current block to derive transform coefficients of the current block, and encoding the transform coefficients of the current block; transmitting data including the bitstream; the transformation is performed based on predetermined transformation kernels including a horizontal transformation kernel and a vertical transformation kernel; the predetermined transform kernel is determined from a set of transforms for a linear inverse transform of the current block based on a size of the current block; The data transmission method, wherein the transformation set is determined based on at least one of an intra prediction mode of the current block or a predefined mapping table.