Image encoding / decoding method and device, and recording medium storing bitstream
Patent Information
- Application Number
- PCT/KR2025/003040
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2025-03-07
- Publication Date
- 2025-10-02
AI Technical Summary
Existing image compression technologies face challenges in efficiently handling high-resolution and high-quality images, particularly in adapting to the diverse characteristics of residual signals, leading to suboptimal compression performance.
The implementation of an adaptive transform unit (TU) structure in video encoding and decoding, which allows for selective utilization of non-separable transform methods based on the characteristics of the residual signal, enabling better matching of the transform unit to the signal characteristics.
This approach enhances compression performance by allowing for more efficient transform methods and kernels, improving the adaptability and efficiency of image encoding and decoding processes.
Smart Images

Figure KR2025003040_02102025_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device, and recording medium storing bitstream
[0001] The present invention relates to a video encoding / decoding method and device, and a recording medium storing a bitstream.
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed.
[0003] There are various technologies for image compression, such as inter prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and entropy encoding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance, and these technologies can be used to effectively compress and transmit or store image data.
[0004] The present disclosure provides a method and apparatus for determining a transform unit for encoding a residual signal.
[0005] The present disclosure provides a method and device for signaling the structure of an adaptive transformation unit.
[0006] The present disclosure provides a signaling method and device for a conversion method according to an adaptive TU structure.
[0007] The video decoding method and device according to the present disclosure can derive transform coefficients of a transform unit from a bitstream and generate residual samples of the transform unit based on inverse transform of the transform coefficients. Here, the inverse transform can be performed based on whether the transform unit has an adaptive TU structure.
[0008] In the video decoding method and device according to the present disclosure, whether the conversion unit has the adaptive TU structure can be determined based on a flag related to whether the adaptive TU structure is applied.
[0009] In the image decoding method and device according to the present disclosure, based on the transform unit having the adaptive TU structure, the inverse transform can be performed based on a non-separable transform method.
[0010] In the image decoding method and device according to the present disclosure, the non-separable transformation method applied to the transformation unit can be determined based on a transformation index indicating any one of the non-separable transformation methods available to the transformation unit.
[0011] In the image decoding method and device according to the present disclosure, the inverse transformation can be performed based on a separation transformation method based on the fact that the transformation unit does not have the adaptive TU structure.
[0012] In the video decoding method and device according to the present disclosure, the separation transformation method applied to the transformation unit can be determined based on a transformation index indicating any one of the separation transformation methods available to the transformation unit.
[0013] In the video decoding method and device according to the present disclosure, an index indicating whether a non-separable transform method is applied to the transform unit can be signaled based on whether the transform unit does not have the adaptive TU structure, and a transform method for the transform unit can be determined based on the index.
[0014] In the image decoding method and device according to the present disclosure, based on the transform unit having the adaptive TU structure, a transform method applied to the transform unit can be determined from a first non-separable transform set.
[0015] In the image decoding method and device according to the present disclosure, based on the fact that the transform unit does not have the adaptive TU structure, the transform method applied to the transform unit can be determined from a second non-separable transform set.
[0016] In the image decoding method and device according to the present disclosure, the number of non-separable transformation methods belonging to the first non-separable transformation set may be different from the number of non-separable transformation methods belonging to the second non-separable transformation set.
[0017] In the video decoding method and device according to the present disclosure, the transformation set for non-separable transformation of the transformation unit can be determined based on at least one of a division mode of a coding unit corresponding to the transformation unit, a division direction of the coding unit, the number of division lines for dividing the transformation unit, or an interval of division lines for dividing the transformation unit.
[0018] The video encoding method and device according to the present disclosure can derive residual samples of a transform unit, derive transform coefficients of the transform unit based on a transform for the residual samples, and encode residual information regarding the transform coefficients. Here, the transform can be performed based on whether the transform unit has an adaptive TU structure.
[0019] A computer-readable digital storage medium is provided having encoded video / image information stored thereon, which causes a decoding device according to the present disclosure to perform a video decoding method.
[0020] A computer-readable digital storage medium storing video / image information generated by a video encoding method according to the present disclosure is provided.
[0021] A method and device for transmitting video / image information generated by a video encoding method according to the present disclosure are provided.
[0022] According to the present disclosure, compression performance of a residual signal can be improved by utilizing an adaptive TU structure that can include a plurality of coding units or prediction units.
[0023] According to the present disclosure, the adaptive TU structure can be selectively utilized through signaling of the adaptive TU structure, and a transformation unit that better matches the characteristics of the residual signal can be set.
[0024] According to the present disclosure, compression performance of a residual signal can be improved by using a more efficient transform method and / or transform kernel in a transform unit having an adaptive TU structure.
[0025] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0026] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0027] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0028] FIG. 4 illustrates a decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0029] FIG. 5 illustrates various segmentation modes based on geometric segmentation as an embodiment according to the present disclosure.
[0030] FIG. 6 is an example according to the present disclosure, showing various types of division modes according to the spacing of division lines.
[0031] FIG. 7 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.
[0032] FIG. 8 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0033] FIG. 9 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.
[0034] FIG. 10 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0035] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0036] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0037] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0038] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0039] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the versatile video coding (VVC) standard. In addition, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or the next generation of video / image coding standards (e.g., H.267 or H.268).
[0040] This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0041] In this specification, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs that has a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively according to the CTU raster scan, while tiles within a picture may be arranged consecutively according to the tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture, which may be exclusively contained within a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.
[0042] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.
[0043] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0044] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0045] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0046] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0047] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0048] Additionally, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction."
[0049] Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously.
[0050] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0051] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device).
[0052] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting device. The receiving device may include a receiving device, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0053] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0054] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0055] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0056] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0057] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0058] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0059] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0060] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0061] For example, a single coding unit may be split into multiple coding units with deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0062] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0063] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0064] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).
[0065] The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information related to prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0066] The intra prediction unit (222) can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of a DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail in the prediction direction. However, this is only an example, and a greater or lesser number of directional modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0067] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0068] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal.
[0069] The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square size and the same size, or can be applied to a block of a non-square variable size.
[0070] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0071] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately from quantized transform coefficients.
[0072] Encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).
[0073] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0074] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0075] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0076] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks in the current picture and transfer them to the intra prediction unit (222).
[0077] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0078] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0079] The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0080] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit of decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device.
[0081] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310).
[0082] Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).
[0083] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0084] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0085] The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra-prediction or inter-prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter-prediction mode.
[0086] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be included and signaled in the video / image information.
[0087] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0088] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.
[0089] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the restoration block.
[0090] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0091] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0092] The (corrected) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) in the current picture and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (332) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (331).
[0093] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.
[0094] The present disclosure may be applied in relation to at least one of a block division process or a residual signal encoding process. Hereinafter, with reference to FIG. 4, a structure in which one transform unit (TU) can include a plurality of prediction units (hereinafter referred to as an adaptive TU structure) is proposed, and a method for improving compression performance by utilizing the same is described. In the present disclosure, a transform unit may mean a coding unit for encoding a residual signal, and a prediction unit may mean a coding unit in which prediction information (e.g., a prediction mode indicating an intra mode or an inter mode, an intra prediction mode, motion information) is determined. The coding unit, the transform unit, and the prediction unit in the present disclosure may also be referred to as a coding block, a transform block, and a prediction block, respectively.
[0095] FIG. 4 illustrates an image decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0096] Referring to Fig. 4, transform coefficients of the transform unit can be derived (S400).
[0097] Residual information of a transform unit can be obtained from a bitstream. The residual information can be decoded to derive transform coefficients of the transform unit.
[0098] The above transformation unit (or the size of the transformation unit) may be derived based on an adaptive TU structure. Specifically, a first coding unit may be split into a plurality of second coding units based on a predetermined split mode. Here, the second coding units may be prediction units for which prediction information is determined. It may be determined whether an adaptive TU structure is applied to a first coding unit including a plurality of second coding units. If it is determined that the adaptive TU structure is applied to the first coding unit, the first coding unit may be set as the transformation unit. That is, the first coding unit including the plurality of second coding units may be treated as one transformation unit. In this case, the first coding unit may not be split further into a plurality of transformation units having smaller sizes. On the other hand, if it is determined that the adaptive TU structure is not applied to the first coding unit, each of the plurality of second coding units may be set as the transformation unit. That is, each of the plurality of second coding units may be treated as one transformation unit.
[0099] The residual information of a transform unit may include position information regarding a last significant coefficient within the transform unit. The last significant coefficient may refer to a non-zero transform coefficient that is located last in a predetermined scan order among non-zero transform coefficient(s) belonging to the transform unit. The position of the last significant coefficient may be specified based on the position information of the last significant coefficient, and the transform coefficients of the transform unit may be derived based on the position information of the last significant coefficient. If it is determined that an adaptive TU structure is applied to the first coding unit, one position information regarding the last significant coefficient may be signaled for the first coding unit. On the other hand, if it is determined that an adaptive TU structure is not applied to the first coding unit, position information regarding the last significant coefficient may be signaled for each of a plurality of second coding units.
[0100] In this way, when applying an adaptive TU structure, the amount of bits required to signal position information about the last significant coefficient can be reduced, and the compression performance of the residual signal can be improved by performing the transformation on a transformation unit having a relatively large size.
[0101] The division mode according to the present disclosure is predefined in the same manner for the encoding device and the decoding device, and may include at least one of quad division and multi-type division. The multi-type division may include at least one of binary division and ternary division.
[0102] Quad partitioning may be a mode that divides a block into four blocks. For example, a block may be divided into four blocks based on one horizontal line and one vertical line. Here, the horizontal line and the vertical line may pass through the center position within the block. However, this is not limited to this, and at least one of the horizontal line or the vertical line may not pass through the center position within the block. Alternatively, a block may be divided into four blocks based on three horizontal lines or three vertical lines. The four blocks resulting from the quad partitioning may all have the same size. Alternatively, one of the four blocks resulting from the quad partitioning may have a different size from the others. Binary partitioning may be a mode that divides a block into two blocks. Binary partitioning may be divided into horizontal binary partitioning, which divides the block into two horizontally, and vertical binary partitioning, which divides the block into two vertically. Ternary partitioning may be a mode that divides a block into three blocks. Ternary division can be divided into horizontal ternary division, which divides into three in the horizontal direction, and vertical ternary division, which divides into three in the vertical direction.
[0103] For example, an adaptive TU structure may be applied to a coding unit to which quad splitting is applied. Specifically, it may be determined whether quad splitting is applied to a first coding unit. Whether quad splitting is applied may be determined based on a quad split flag (qt_split_flag) indicating whether quad splitting is applied. If the quad split flag indicates that quad splitting is applied (qt_split_flag=1), the first coding unit may be split into four second coding units. In this case, it may be determined whether an adaptive TU structure is applied to the first coding unit to which quad splitting is applied. If it is determined that the adaptive TU structure is applied to the first coding unit, the first coding unit including four second coding units may be set as a transformation unit. In this case, if each of the four second coding units corresponds to a prediction unit, one transformation unit may include four prediction units. If the quad split flag indicates that quad splitting is not applied (qt_split_flag=0), the first coding unit corresponds to a leaf node of the quad split, and it can be determined whether multi-type splitting is applied to the first coding unit. Whether multi-type splitting is applied can be determined based on the multi-type split flag (mtt_split_flag) indicating whether multi-type splitting is applied. If the multi-type split flag indicates that multi-type splitting is applied (mtt_split_flag=1), the splitting direction and whether binary splitting is applied can be determined. The splitting direction can be determined based on the splitting direction flag (mtt_split_vertical_flag) indicating either a horizontal or vertical splitting direction. Whether binary splitting is applied can be determined based on the binary split flag (mtt_split_bt_flag) indicating whether binary splitting is applied.Based on the split direction flag and the binary split flag, the first coding unit may be split into two or three second coding units. For example, when both mtt_split_vertical_flag and mtt_split_bt_flag are 1, the first coding unit may be split into two second coding units in the vertical direction. When mtt_split_vertical_flag is 1 and mtt_split_bt_flag is 0, the first coding unit may be split into three coding units in the vertical direction. When mtt_split_vertical_flag is 0 and mtt_split_bt_flag is 1, the first coding unit may be split into two coding units in the horizontal direction. If both mtt_split_vertical_flag and mtt_split_bt_flag are 0, the first coding unit can be split into three second coding units in the horizontal direction. If the multi-type split flag indicates that multi-type splitting is not applied (mtt_split_flag=0), the first coding unit corresponds to a leaf node of the multi-type split, and no further quad splitting or multi-type splitting may be applied to the first coding unit.
[0104] For example, an adaptive TU structure may be applied to a coding unit to which binary splitting is applied. Specifically, it may be determined whether quad splitting is applied to a first coding unit. Whether quad splitting is applied may be determined based on a quad split flag (qt_split_flag) indicating whether quad splitting is applied. If the quad split flag indicates that quad splitting is applied (qt_split_flag=1), the first coding unit may be split into four second coding units. If the quad split flag indicates that quad splitting is not applied (qt_split_flag=0), the first coding unit corresponds to a leaf node of the quad splitting, and it may be determined whether multi-type splitting is applied to the first coding unit. Whether multi-type splitting is applied may be determined based on a multi-type split flag (mtt_split_flag) indicating whether multi-type splitting is applied. When the multi-type split flag indicates that multi-type splitting is applied (mtt_split_flag=1), the splitting direction and whether binary splitting is applied can be determined. The splitting direction can be determined based on a splitting direction flag (mtt_split_vertical_flag) that indicates either a horizontal or vertical splitting direction. Whether binary splitting is applied can be determined based on a binary splitting flag (mtt_split_bt_flag) that indicates whether binary splitting is applied. Based on the splitting direction flag and the binary splitting flag, the first coding unit can be split into two or three second coding units. For example, when both mtt_split_vertical_flag and mtt_split_bt_flag are 1, the first coding unit can be split into two second coding units in the vertical direction.When mtt_split_vertical_flag is 1 and mtt_split_bt_flag is 0, the first coding unit may be split into three coding units in the vertical direction. When mtt_split_vertical_flag is 0 and mtt_split_bt_flag is 1, the first coding unit may be split into two coding units in the horizontal direction. When both mtt_split_vertical_flag and mtt_split_bt_flag are 0, the first coding unit may be split into three second coding units in the horizontal direction. At this time, it may be determined whether an adaptive TU structure is applied to the first coding unit to which binary splitting is applied. That is, if the binary split flag for the first coding unit indicates that binary splitting is applied (mtt_split_bt_flag=1), it may be determined whether an adaptive TU structure is applied to the first coding unit. If it is determined that the adaptive TU structure is applied to the first coding unit, the first coding unit including two second coding units may be set as a transformation unit. In this case, if each of the two second coding units corresponds to a prediction unit, one transformation unit may include two prediction units. If the multi-type split flag indicates that multi-type splitting is not applied (mtt_split_flag=0), the first coding unit corresponds to a leaf node of the multi-type splitting, and no further quad splitting and multi-type splitting may be applied to the first coding unit.
[0105] For example, an adaptive TU structure may be applied to a coding unit to which ternary splitting is applied. Specifically, it may be determined whether quad splitting is applied to a first coding unit. Whether quad splitting is applied may be determined based on a quad split flag (qt_split_flag) indicating whether quad splitting is applied. If the quad split flag indicates that quad splitting is applied (qt_split_flag=1), the first coding unit may be split into four second coding units. If the quad split flag indicates that quad splitting is not applied (qt_split_flag=0), the first coding unit corresponds to a leaf node of the quad splitting, and it may be determined whether multi-type splitting is applied to the first coding unit. Whether multi-type splitting is applied may be determined based on a multi-type split flag (mtt_split_flag) indicating whether multi-type splitting is applied. When the multi-type split flag indicates that multi-type splitting is applied (mtt_split_flag=1), the splitting direction and whether binary splitting is applied can be determined. The splitting direction can be determined based on a splitting direction flag (mtt_split_vertical_flag) that indicates either a horizontal or vertical splitting direction. Whether binary splitting is applied can be determined based on a binary splitting flag (mtt_split_bt_flag) that indicates whether binary splitting is applied. Based on the splitting direction flag and the binary splitting flag, the first coding unit can be split into two or three second coding units. For example, when both mtt_split_vertical_flag and mtt_split_bt_flag are 1, the first coding unit can be split into two second coding units in the vertical direction.When mtt_split_vertical_flag is 1 and mtt_split_bt_flag is 0, the first coding unit may be split into three coding units in the vertical direction. When mtt_split_vertical_flag is 0 and mtt_split_bt_flag is 1, the first coding unit may be split into two coding units in the horizontal direction. When both mtt_split_vertical_flag and mtt_split_bt_flag are 0, the first coding unit may be split into three second coding units in the horizontal direction. At this time, it may be determined whether an adaptive TU structure is applied to the first coding unit to which ternary splitting is applied. That is, if the binary split flag for the first coding unit indicates that ternary splitting is applied (mtt_split_bt_flag=0), it may be determined whether an adaptive TU structure is applied to the first coding unit. If it is determined that the adaptive TU structure is applied to the first coding unit, the first coding unit including three second coding units may be set as a transform unit. In this case, if each of the three second coding units corresponds to a prediction unit, one transform unit may include three prediction units. If the multi-type split flag indicates that multi-type splitting is not applied (mtt_split_flag=0), the first coding unit corresponds to a leaf node of the multi-type splitting, and no further quad splitting and multi-type splitting may be applied to the first coding unit.
[0106] For example, it may be determined whether a first coding unit is split into a plurality of second coding units. Whether the first coding unit is split into a plurality of second coding units may be determined based on a split flag (split_cu_flag) indicating whether the coding unit is split. If the split flag indicates that the first coding unit is split into a plurality of second coding units (split_cu_flag=1), the first coding unit may be split into a plurality of second coding units based on a predetermined split mode. In this way, if the first coding unit is split (or split_cu_flag=1), it may be determined whether an adaptive TU structure is applied to the first coding unit. If it is determined that the adaptive TU structure is applied to the first coding unit, the first coding unit including the plurality of second coding units may be set as a transformation unit. In this case, if each of the plurality of second coding units corresponds to a prediction unit, one transformation unit may include the plurality of prediction units. Here, the plurality of second coding units may be generated by splitting the first coding unit based on quad splitting, binary splitting, or ternary splitting. The adaptive TU structure is not limited to being applicable only to the first coding unit to which a specific splitting mode is applied. If the first coding unit is split based on a splitting mode that is identically pre-defined for the encoding device and the decoding device, it is possible to determine whether to apply the adaptive TU structure to the first coding unit regardless of the splitting mode applied to the first coding unit. If the split flag indicates that the first coding unit is not split into a plurality of second coding units (split_cu_flag=0), the first coding unit may not be split into coding units of a smaller size, and it may not be determined whether to apply the adaptive TU structure to such first coding unit.
[0107] Alternatively, the partitioning mode to which the adaptive TU structure is applicable may be different from the partitioning mode pre-defined equally in the encoding device and the decoding device for the aforementioned block partitioning.
[0108] For example, the pre-defined partitioning modes may include quad partitioning, binary partitioning, and ternary partitioning, but the partitioning modes to which the adaptive TU structure is applicable may include at least one of quad partitioning, binary partitioning, and ternary partitioning. For example, it may be determined whether the adaptive TU structure is applied to the first coding unit only when quad partitioning is applied to the first coding unit. Alternatively, it may be determined whether the adaptive TU structure is applied to the first coding unit only when multi-type partitioning is applied to the first coding unit. Alternatively, it may be determined whether the adaptive TU structure is applied to the first coding unit only when binary partitioning is applied to the first coding unit. Alternatively, it may be determined whether the adaptive TU structure is applied to the first coding unit only when ternary partitioning is applied to the first coding unit. Alternatively, it may be determined whether the adaptive TU structure is applied to the first coding unit only when quad partitioning and binary partitioning are applied to the first coding unit. Alternatively, it may be determined whether the adaptive TU structure is applied to the first coding unit only when quad splitting and ternary splitting are applied to the first coding unit.
[0109] The segmentation mode to which the adaptive TU structure can be applied can be predefined identically for the encoding device and the decoding device separately from the segmentation mode predefined for block segmentation. Alternatively, the segmentation mode to which the adaptive TU structure can be applied can be defined in the high-level syntax of the bitstream. Here, the high-level syntax can include at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH).
[0110] Whether the aforementioned adaptive TU structure is applied can be determined based on at least one of a flag (TU_mPU_flag) related to whether the adaptive TU structure is applied, a split flag indicating whether the first coding unit is split, split information for block splitting of the first coding unit, or the size of the first coding unit. Hereinafter, methods for determining whether the adaptive TU structure is applied in methods 1 to 5 will be described in detail.
[0111] It goes without saying that whether an adaptive TU structure is applied may be determined based on any one of the methods 1 to 5 described below, or may be determined based on a combination of at least two of the methods 1 to 5.
[0112] Method 1
[0113] Whether an adaptive TU structure is applied can be determined based on a flag (TU_mPU_flag) related to whether an adaptive TU structure is applied.
[0114] For example, it may be determined that an adaptive TU structure is applied to the first coding unit based on TU_mPU_flag being 1, and it may be determined that an adaptive TU structure is not applied to the first coding unit based on TU_mPU_flag being 0.
[0115] Method 2
[0116] Whether an adaptive TU structure is applied can be determined based on at least one of a split flag (split_cu_flag) or TU_mPU_flag indicating whether the first coding unit is split.
[0117] For example, based on split_cu_flag being 1, it may be determined that the adaptive TU structure is applied to the first coding unit. On the other hand, based on split_cu_flag being 0, it may be determined that the adaptive TU structure is not applied to the first coding unit. Alternatively, based on both split_cu_flag and TU_mPU_flag being 1, it may be determined that the adaptive TU structure is applied to the first coding unit. On the other hand, based on at least one of split_cu_flag or TU_mPU_flag being 0, it may be determined that the adaptive TU structure is not applied to the first coding unit.
[0118] Method 3
[0119] Whether an adaptive TU structure is applied can be determined based on at least one of the partitioning information for block partitioning of the first coding unit or the TU_mPU_flag. Here, the partitioning information can include at least one of a quad partitioning flag, a multi-type partitioning flag, a binary partitioning flag, or a partitioning direction flag.
[0120] For example, based on qt_split_flag being 1, it may be determined that the adaptive TU structure is applied to the first coding unit. Based on qt_split_flag being 0, it may be determined that the adaptive TU structure is not applied to the first coding unit. Alternatively, based on mtt_split_flag being 1, it may be determined that the adaptive TU structure is applied to the first coding unit. Based on mtt_split_flag being 0, it may be determined that the adaptive TU structure is not applied to the first coding unit. Alternatively, based on mtt_split_bt_flag being 1, it may be determined that the adaptive TU structure is applied to the first coding unit. Based on mtt_split_bt_flag being 0, it may be determined that the adaptive TU structure is not applied to the first coding unit.
[0121] Alternatively, based on both qt_split_flag and TU_mPU_flag being 1, it may be determined that the adaptive TU structure is applied to the first coding unit. Based on at least one of qt_split_flag or TU_mPU_flag being 0, it may be determined that the adaptive TU structure is not applied to the first coding unit. Alternatively, based on both mtt_split_flag and TU_mPU_flag being 1, it may be determined that the adaptive TU structure is applied to the first coding unit. Based on at least one of mtt_split_flag or TU_mPU_flag being 0, it may be determined that the adaptive TU structure is not applied to the first coding unit. Alternatively, based on both mtt_split_bt_flag and TU_mPU_flag being 1, it may be determined that the adaptive TU structure is applied to the first coding unit. Based on at least one of mtt_split_bt_flag or TU_mPU_flag being 0, it may be determined that the adaptive TU structure is not applied to the first coding unit.
[0122] Method 4
[0123] Whether an adaptive TU structure is applied can be determined based on the size of the first coding unit.
[0124] For example, if the size of the first coding unit is less than or equal to a predetermined threshold, an adaptive TU structure may be applied to the first coding unit. Otherwise, the adaptive TU structure may not be applied to the first coding unit. Here, the size may be defined as at least one of width, height, the product of width and height, the ratio of width and height, or the maximum / minimum values of width and height.
[0125] The above threshold may refer to a maximum block size to which the adaptive TU structure can be applied. The threshold may be a value predefined equally for the encoding device and the decoding device. Alternatively, size information regarding the maximum block size to which the adaptive TU structure can be applied may be signaled through the bitstream. In this case, the threshold may be derived based on the signaled size information. The size information may be signaled at at least one level among SPS, PPS, PH, and SH.
[0126] For example, the size information may include at least one of information about a maximum block width to which the adaptive TU structure can be applied (TU_mPU_width_max) and information about a maximum block height to which the adaptive TU structure can be applied (TU_mPU_height_max). The maximum block width and the maximum block height to which the adaptive TU structure can be applied may be derived based on TU_mPU_width_max and TU_mPU_height_max, respectively. If the width of the first coding unit is less than or equal to the maximum block width and the height of the first coding unit is less than or equal to the maximum block height, the adaptive TU structure may be applied to the first coding unit. Otherwise, the adaptive TU structure may not be applied to the first coding unit.
[0127] Alternatively, the size information may be defined as information regarding the product of the maximum block width and the maximum block height to which the adaptive TU structure can be applied (or the maximum number of samples of a block to which the adaptive TU structure can be applied). In this case, if the product of the width and the height of the first coding unit is less than or equal to a threshold derived based on the size information, the adaptive TU structure may be applied to the first coding unit. Otherwise, the adaptive TU structure may not be applied to the first coding unit.
[0128] The above threshold value may be defined differently depending on the partitioning mode applied to the first coding unit. Alternatively, size information for deriving the threshold value may be signaled for each partitioning mode to which the adaptive TU structure is applicable.
[0129] For example, if the partitioning mode applied to the first coding unit is quad partitioning, an adaptive TU structure may be applied to the first coding unit based on the width of the first coding unit being less than or equal to a first threshold, and the adaptive TU structure may not be applied to the first coding unit based on the width of the first coding unit being greater than the first threshold. On the other hand, if the partitioning mode applied to the first coding unit is binary partitioning, an adaptive TU structure may be applied to the first coding unit based on the width of the first coding unit being less than or equal to a second threshold, and the adaptive TU structure may not be applied to the first coding unit based on the width of the first coding unit being greater than the second threshold. Here, the first threshold and the second threshold may be any one of 16, 32, 64, or 128, respectively. The first threshold may be defined to have a value greater than the second threshold.
[0130] Method 5
[0131] Whether an adaptive TU structure is applied can be determined based on the size of the first coding unit.
[0132] For example, if the size of the first coding unit is greater than or equal to a predetermined threshold, an adaptive TU structure may be applied to the first coding unit. Otherwise, the adaptive TU structure may not be applied to the first coding unit. Here, the size may be defined as at least one of width, height, the product of width and height, the ratio of width and height, or the maximum / minimum values of width and height.
[0133] The above threshold may refer to a minimum block size to which the adaptive TU structure can be applied. The threshold may be a value predefined equally for the encoding device and the decoding device. Alternatively, size information regarding the minimum block size to which the adaptive TU structure can be applied may be signaled through the bitstream. In this case, the threshold may be derived based on the signaled size information. The size information may be signaled at at least one level among SPS, PPS, PH, and SH.
[0134] For example, the size information may include at least one of information about a minimum block width to which the adaptive TU structure is applicable (TU_mPU_width_min) and information about a minimum block height to which the adaptive TU structure is applicable (TU_mPU_height_min). Based on TU_mPU_width_min and TU_mPU_height_min, the minimum block width and the minimum block height to which the adaptive TU structure is applicable may be derived, respectively. If the width of the first coding unit is greater than or equal to the minimum block width and the height of the first coding unit is greater than or equal to the minimum block height, the adaptive TU structure may be applied to the first coding unit. Otherwise, the adaptive TU structure may not be applied to the first coding unit.
[0135] Alternatively, the size information may be defined as information regarding the product of the minimum block width and the minimum block height to which the adaptive TU structure can be applied (or the minimum number of samples of a block to which the adaptive TU structure can be applied). In this case, if the product of the width and the height of the first coding unit is greater than or equal to a threshold derived based on the size information, the adaptive TU structure may be applied to the first coding unit. Otherwise, the adaptive TU structure may not be applied to the first coding unit.
[0136] The above threshold value may be defined differently depending on the partitioning mode applied to the first coding unit. Alternatively, size information for deriving the threshold value may be signaled for each partitioning mode to which the adaptive TU structure is applicable.
[0137] For example, if the partitioning mode applied to the first coding unit is quad partitioning, an adaptive TU structure may be applied to the first coding unit based on the width of the first coding unit being greater than or equal to a first threshold, and the adaptive TU structure may not be applied to the first coding unit based on the width of the first coding unit being less than the first threshold. On the other hand, if the partitioning mode applied to the first coding unit is binary partitioning, an adaptive TU structure may be applied to the first coding unit based on the width of the first coding unit being greater than or equal to a second threshold, and the adaptive TU structure may not be applied to the first coding unit based on the width of the first coding unit being less than the second threshold. Here, the first threshold and the second threshold may be any one of 4, 8, 16, or 32, respectively. The first threshold may be defined to have a value greater than the second threshold.
[0138] The aforementioned TU_mPU_flag can be explicitly signaled through the bitstream as information on whether or not the adaptive TU structure is applied. For example, a TU_mPU_flag of 1 can indicate that the transform unit is split into a plurality of prediction units based on a predetermined partitioning mode, and a TU_mPU_flag of 0 can indicate that the transform unit is not split into a plurality of prediction units. Alternatively, a TU_mPU_flag of 1 can indicate that a first coding unit split into a plurality of second coding units is treated as a transform unit. On the other hand, a TU_mPU_flag of 0 can indicate that the first coding unit split into a plurality of second coding units based on a predetermined partitioning mode is not treated as a transform unit. Alternatively, a TU_mPU_flag of 1 can indicate that the transform unit has the same size as the first coding unit split into a plurality of second coding units (or prediction units). On the other hand, a TU_mPU_flag of 0 may indicate that the restriction that the transform unit has the same size as the first coding unit does not apply. That is, when TU_mPU_flag is 0, the transform unit may have the same size as the first coding unit, or may have different sizes.
[0139] The TU_mPU_flag according to the present disclosure may be encoded and signaled based on at least one of a split flag indicating whether the first coding unit is split, split information for block splitting of the first coding unit, or the size of the first coding unit. If the TU_mPU_flag is not signaled, the value of the TU_mPU_flag may be derived as 0. Hereinafter, the method of signaling the TU_mPU_flag in methods 1 to 6 will be described in detail.
[0140] TU_mPU_flag may be signaled based on any one of the methods 1 to 6 described below, or may be signaled based on a combination of at least two of the methods 1 to 6.
[0141] Method 1
[0142] TU_mPU_flag can be signaled based on a split flag (split_cu_flag) indicating whether the first coding unit is split.
[0143] For example, based on split_cu_flag being 1, TU_mPU_flag may be signaled for the first coding unit. On the other hand, based on split_cu_flag being 0, TU_mPU_flag may not be signaled for the first coding unit.
[0144] Method 2
[0145] TU_mPU_flag may be signaled based on partitioning information for block partitioning of the first coding unit. Here, the partitioning information may include at least one of a quad partitioning flag, a multi-type partitioning flag, a binary partitioning flag, or a partitioning direction flag.
[0146] For example, based on qt_split_flag being 1, TU_mPU_flag may be signaled for the first coding unit. On the other hand, based on qt_split_flag being 0, TU_mPU_flag may not be signaled for the first coding unit. Alternatively, based on mtt_split_flag being 1, TU_mPU_flag may be signaled for the first coding unit. On the other hand, based on mtt_split_flag being 0, TU_mPU_flag may not be signaled for the first coding unit. Alternatively, based on mtt_split_bt_flag being 1, TU_mPU_flag may be signaled for the first coding unit. On the other hand, based on mtt_split_bt_flag being 0, TU_mPU_flag may not be signaled for the first coding unit.
[0147] Method 3
[0148] TU_mPU_flag can be signaled based on the size of the first coding unit.
[0149] For example, if the size of the first coding unit is less than or equal to a predetermined threshold, the TU_mPU_flag may be signaled for the first coding unit. Otherwise, the TU_mPU_flag may not be signaled for the first coding unit. Here, the size may be defined as at least one of width, height, product of width and height, ratio of width and height, or maximum / minimum values of width and height.
[0150] The above threshold may refer to a maximum block size to which the adaptive TU structure can be applied. The threshold may be a value predefined equally for the encoding device and the decoding device. Alternatively, size information regarding the maximum block size to which the adaptive TU structure can be applied may be signaled through the bitstream. In this case, the threshold may be derived based on the signaled size information. The size information may be signaled at at least one level among SPS, PPS, PH, and SH.
[0151] For example, the size information may include at least one of information about a maximum block width to which the adaptive TU structure can be applied (TU_mPU_width_max) and information about a maximum block height to which the adaptive TU structure can be applied (TU_mPU_height_max). The maximum block width and the maximum block height to which the adaptive TU structure can be applied may be derived based on TU_mPU_width_max and TU_mPU_height_max, respectively. If the width of the first coding unit is less than or equal to the maximum block width and the height of the first coding unit is less than or equal to the maximum block height, TU_mPU_flag may be signaled for the first coding unit. Otherwise, TU_mPU_flag may not be signaled for the first coding unit.
[0152] Alternatively, the size information may be defined as information about the product of the maximum block width and the maximum block height to which the adaptive TU structure can be applied (or, the maximum number of samples of a block to which the adaptive TU structure can be applied). In this case, if the product of the width and the height of the first coding unit is less than or equal to a threshold derived based on the size information, the TU_mPU_flag may be signaled for the first coding unit. Otherwise, the TU_mPU_flag may not be signaled for the first coding unit.
[0153] The above threshold value may be defined differently depending on the partitioning mode applied to the first coding unit. Alternatively, size information for deriving the threshold value may be signaled for each partitioning mode to which the adaptive TU structure is applicable.
[0154] For example, if the partitioning mode applied to the first coding unit is quad partitioning, the TU_mPU_flag may be signaled for the first coding unit based on the width of the first coding unit being less than or equal to a first threshold, and the TU_mPU_flag may not be signaled for the first coding unit based on the width of the first coding unit being greater than the first threshold. On the other hand, if the partitioning mode applied to the first coding unit is binary partitioning, the TU_mPU_flag may be signaled for the first coding unit based on the width of the first coding unit being less than or equal to a second threshold, and the TU_mPU_flag may not be signaled for the first coding unit based on the width of the first coding unit being greater than the second threshold. Here, the first threshold and the second threshold may each be any one of 16, 32, 64, or 128. The first threshold can be defined to have a value greater than the second threshold.
[0155] Method 4
[0156] TU_mPU_flag can be signaled based on the size of the first coding unit.
[0157] For example, if the size of the first coding unit is greater than or equal to a predetermined threshold, the TU_mPU_flag may be signaled for the first coding unit. Otherwise, the TU_mPU_flag may not be signaled for the first coding unit. Here, the size may be defined as at least one of width, height, product of width and height, ratio of width and height, or maximum / minimum values of width and height.
[0158] The above threshold may refer to a minimum block size to which the adaptive TU structure can be applied. The threshold may be a value predefined equally for the encoding device and the decoding device. Alternatively, size information regarding the minimum block size to which the adaptive TU structure can be applied may be signaled through the bitstream. In this case, the threshold may be derived based on the signaled size information. The size information may be signaled at at least one level among SPS, PPS, PH, and SH.
[0159] For example, the size information may include at least one of information about a minimum block width to which the adaptive TU structure is applicable (TU_mPU_width_min) and information about a minimum block height to which the adaptive TU structure is applicable (TU_mPU_height_min). The minimum block width and the minimum block height to which the adaptive TU structure is applicable may be derived based on TU_mPU_width_min and TU_mPU_height_min, respectively. If the width of the first coding unit is greater than or equal to the minimum block width and the height of the first coding unit is greater than or equal to the minimum block height, TU_mPU_flag may be signaled for the first coding unit. Otherwise, TU_mPU_flag may not be signaled for the first coding unit.
[0160] Alternatively, the size information may be defined as information regarding the product of the minimum block width and the minimum block height to which the adaptive TU structure is applicable (or the minimum number of samples of a block to which the adaptive TU structure is applicable). In this case, if the product of the width and the height of the first coding unit is greater than or equal to a threshold derived based on the size information, the TU_mPU_flag may be signaled for the first coding unit. Otherwise, the TU_mPU_flag may not be signaled for the first coding unit.
[0161] The above threshold value may be defined differently depending on the partitioning mode applied to the first coding unit. Alternatively, size information for deriving the threshold value may be signaled for each partitioning mode to which the adaptive TU structure is applicable.
[0162] For example, if the partitioning mode applied to the first coding unit is quad partitioning, the TU_mPU_flag may be signaled for the first coding unit based on the width of the first coding unit being greater than or equal to a first threshold, and the TU_mPU_flag may not be signaled for the first coding unit based on the width of the first coding unit being less than the first threshold. On the other hand, if the partitioning mode applied to the first coding unit is binary partitioning, the TU_mPU_flag may be signaled for the first coding unit based on the width of the first coding unit being greater than or equal to a second threshold, and the TU_mPU_flag may not be signaled for the first coding unit based on the width of the first coding unit being less than the second threshold. Here, the first threshold and the second threshold may be any one of 4, 8, 16, or 32, respectively. The first threshold can be defined to have a value greater than the second threshold.
[0163] Method 5
[0164] TU_mPU_flag can be encoded and signaled based on a variable (allow TU_mPU) regarding whether or not adaptive TU structures are allowed.
[0165] For example, if allow TU_mPU is True (or allow TU_mPU is 1), TU_mPU_flag may be signaled for the first coding unit. If allow TU_mPU is False (or allow TU_mPU is 0), TU_mPU_flag may not be signaled for the first coding unit.
[0166] The allow TU_mPU of the present disclosure may be set to True if a predetermined condition is satisfied, and to False otherwise. Here, the predetermined condition may include at least one of Conditions 1 and 2 described below.
[0167] Condition 1: Adaptive TU structure is allowed for the sequence.
[0168] Condition 2: Adaptive TU structure is allowed for the current picture.
[0169] Condition 3: Adaptive TU structure is allowed for the current slice.
[0170] Condition 4: Adaptive TU structure is allowed for the split mode applied to the first coding unit.
[0171] Whether an adaptive TU structure is allowed for a sequence to which the first coding unit belongs can be determined based on a flag, which is a syntax signaled in a sequence parameter set and indicates whether an adaptive TU structure is allowed. Whether an adaptive TU structure is allowed for a current picture to which the first coding unit belongs can be determined based on a flag, which is a syntax signaled in at least one of a picture parameter set or a picture header and indicates whether an adaptive TU structure is allowed. Whether an adaptive TU structure is allowed for a current slice to which the first coding unit belongs can be determined based on a flag, which is a syntax signaled in a slice header and indicates whether an adaptive TU structure is allowed. When a partitioning mode applied to the first coding unit belongs to a partitioning mode to which an adaptive TU structure is applicable, it can be determined that an adaptive TU structure is allowed for the corresponding partitioning mode. As previously discussed, the segmentation mode to which the adaptive TU structure can be applied can be defined in the high level syntax of the bitstream.
[0172] Table 1 below is an example syntax table showing how to signal TU_mPU_flag.
[0173] coding_tree( x0, y0, cbWidth, cbHeight, qgOnY, qgOnC, cbSubdiv, cqtDepth, mttDepth, depthOffset, partIdx, treeTypeCurr, modeTypeCurr ) {Descriptorif( ( allowSplitBtVer | | allowSplitBtHor | | allowSplitTtVer | | allowSplitTtHor | | allowSplitQt ) && ( x0 + cbWidth <= pps_pic_width_in_luma_samples ) && ( y0 + cbHeight <= pps_pic_height_in_luma_samples ) )split_cu_flagae(v)...if( split_cu_flag ) {if( ( allowSplitBtVer | | allowSplitBtHor | | allowSplitTtVer | | allowSplitTtHor ) && allowSplitQt )split_qt_flagae(v)if (allowTU_mPU&& cbWidth <= TU_mPU_width_max && cbHeight <= TU_mPU_height_max )TU_mPU_flagae(v)if( !split_qt_flag ) {if( ( allowSplitBtHor | | allowSplitTtHor ) && ( allowSplitBtVer | | allowSplitTtVer ) )mtt_split_cu_vertical_flagae(v)if( ( allowSplitBtVer && allowSplitTtVer && mtt_split_cu_vertical_flag ) | | ( allowSplitBtHor && allowSplitTtHor && !mtt_split_cu_vertical_flag ) )mtt_split_cu_binary_flagae(v)if (allowTU_mPU&& cbWidth <= TU_mPU_width_max && cbHeight <= TU_mPU_height_max )TU_mPU_flagae(v)}if( ModeTypeCondition = = 1 )modeType = MODE_TYPE_INTRAelse if( ModeTypeCondition = = 2 ) {non_inter_flagae(v)modeType = non_inter_flag ? MODE_TYPE_INTRA : MODE_TYPE_INTER} elsemodeType = modeTypeCurr.
[0174] According to Table 1, TU_mPU_flag can be signaled based on the variable (allow TU_mPU) being True, cbWidth being less than or equal to TU_mPU_width_max, and cbHeight being less than or equal to TU_mPU_height_max. Here, cbWidth and cbHeight can mean the width and height of the coding unit, respectively, for which tree-structure-based block segmentation is performed.
[0175] Additionally, according to Table 1, split_qt_flag may be signaled based on split_cu_flag being 1, and TU_mPU_flag may be signaled after split_qt_flag is signaled. However, this is only an example, and TU_mPU_flag may be signaled before split_qt_flag is signaled.
[0176] Additionally, according to Table 1, mtt_split_cu_binary_flag may be signaled based on split_qt_flag being 0, and TU_mPU_flag may be signaled after mtt_split_cu_binary_flag is signaled. However, this is just an example, and TU_mPU_flag may be signaled before mtt_split_cu_binary_flag is signaled.
[0177] Method 6
[0178] The TU_mPU_flag for the first coding unit may be signaled based on the TU_mPU_flag of the upper coding unit. Here, the upper coding unit is a coding unit to which the first coding unit belongs, and may be a coding unit having a split depth smaller than the split depth of the first coding unit. If the split depth of the first coding unit is N, the split depth of the upper coding unit may be (N-1).
[0179] For example, based on the TU_mPU_flag of the upper coding unit being 1, the TU_mPU_flag may not be signaled for the first coding unit. On the other hand, based on the TU_mPU_flag of the upper coding unit being 0, the TU_mPU_flag may be signaled for the first coding unit.
[0180] This is to prevent the first coding unit from being reset as a transform unit when the TU_mPU_flag of the upper coding unit is 1, since the upper coding unit is set as a transform unit. On the other hand, when the TU_mPU_flag is 0, this means that the adaptive TU structure is determined not to be applied to the upper coding unit, so it is possible to determine whether the adaptive TU structure is applied by signaling the TU_mPU_flag for the first coding unit generated by dividing the upper coding unit.
[0181] A single coding tree unit (CTU) can be divided into coding units, which are units for encoding a video signal, through tree-based block division. In this tree-based block division structure, the TU_mPU_flag may be recursively signaled along with division information for the coding unit.
[0182] For example, a first coding unit may be split into four second coding units based on a quad split. Here, if the first coding unit has a split depth of N, the second coding blocks may have a split depth of (N+1). A TU_mPU_flag may be signaled for the first coding unit, and it may be determined whether an adaptive TU structure is applied to the first coding unit based on the TU_mPU_flag. If the TU_mPU_flag for the first coding unit is 1, the first coding unit including the four second coding units may be set as a transform unit. On the other hand, if the TU_mPU_flag for the first coding unit is 0, the first coding unit may not be set as a transform unit.
[0183] If the second coding unit corresponds to a node of a quad partition structure, the second coding unit may be partitioned into four third coding units based on the quad partition. Here, the third coding units may have a partition depth of (N+2). In this case, a TU_mPU_flag may be signaled for the second coding unit, and it may be determined whether an adaptive TU structure is applied to the second coding unit based on the TU_mPU_flag. However, the TU_mPU_flag for the second coding unit may be restricted to be signaled when the TU_mPU_flag for the first coding unit is 0.
[0184] Alternatively, if the second coding unit corresponds to a leaf node of a quad partitioning structure, the second coding unit may be partitioned into a plurality of third coding units based on a multi-type partitioning. The second coding unit may be partitioned into two third coding units based on a binary partitioning, or may be partitioned into three third coding units based on a ternary partitioning. Here, the third coding units may have a partitioning depth of (N+2). In this case, a TU_mPU_flag may be signaled for the second coding unit, and it may be determined whether an adaptive TU structure is applied to the second coding unit based on the TU_mPU_flag. However, the TU_mPU_flag for the second coding unit may be restricted to be signaled when the TU_mPU_flag for the first coding unit is 0.
[0185] Referring to FIG. 4, residual samples of the conversion unit can be derived based on the conversion coefficients of the conversion unit (S410).
[0186] Residual samples of a transform unit can be derived based on at least one of inverse quantization or inverse transformation of transform coefficients. Specifically, inverse quantized transform coefficients can be derived based on inverse quantization of transform coefficients. Residual samples can be derived based on inverse transformation of inverse quantized transform coefficients.
[0187] The above inverse transformation can be understood as the reverse process of the transformation performed in the encoding device. The transformation according to the present disclosure is a compression technique that transforms the residual signal (i.e., residual block or residual samples) based on basis vectors called transformation kernels in order to efficiently encode statistical similarity (spatial correlation) existing in the residual signal derived through intra prediction or inter prediction. In the encoding device, the residual signal is input and transform coefficients are output, and in the decoding device, the transform coefficients are input and the residual signal is output.
[0188] The above inverse transformation may be performed based on at least one of a separable transformation and a non-separable transformation. The separable transformation may be a process of first performing a transformation in either a horizontal or vertical direction for a two-dimensional transformation unit and then performing a transformation in the other direction on the result. The non-separable transformation may be a process of performing a transformation once on transformation coefficients constituting the entire or a portion of a two-dimensional transformation unit.
[0189] The present disclosure describes a method for signaling a transform method for efficiently decoding a transform unit having an adaptive TU structure.
[0190] A decoding device may define a plurality of transform methods for inverse transform, and at least one of the plurality of transform methods may be selectively used. The plurality of transform methods may include at least one of a non-separable transform-based transform method (hereinafter referred to as a non-separable transform method) or a separable transform-based transform method (hereinafter referred to as a separable transform method). First to Nth non-separable transform method(s) may be defined as the non-separable transform methods, and first to Mth separable transform method(s) may be defined as the separable transform methods. Here, N and M may be integers greater than or equal to 1. N and M may be the same value or different values. The non-separable transform may be referred to as a non-separable primary transform (NSPT) or a non-separable secondary transform (NSST).
[0191] The transform method according to the present disclosure can be determined based on whether a transform unit has an adaptive TU structure. As described above, it can be determined whether an adaptive TU structure is applied to a first coding unit including a plurality of second coding units. If it is determined that the adaptive TU structure is applied to the first coding unit, the first coding unit including the plurality of second coding units is set as one transform unit, and thus the transform unit can have the adaptive TU structure. On the other hand, if it is determined that the adaptive TU structure is not applied to the first coding unit, each of the plurality of second coding units can be set as one transform unit, and thus the transform unit may not have the adaptive TU structure. That is, whether a transform unit has an adaptive TU structure can be determined based on whether an adaptive TU structure is applied to the first coding unit, and the method for determining whether an adaptive TU structure is applied is as described above.
[0192] Example 1
[0193] Based on whether the transform unit has an adaptive TU structure, a variable (Transform_idx) indicating a transform method to be applied to the transform unit can be derived. The transform method for the transform unit can be determined based on the derived variable (Transform_idx). Specifically, when the transform unit has an adaptive TU structure, the variable (Transform_idx) can be set to a predefined value. Here, the predefined value is a value defined identically in an encoding device and a decoding device, and can be a value indicating a non-separable transform method. When the transform unit has an adaptive TU structure, the non-separable transform method can be applied without signaling a transform index (Tr_idx) to be described later. That is, an inverse transform can be performed based on a transform kernel for non-separable transform for the transform unit.
[0194] If the transform unit does not have an adaptive TU structure, the variable (Transform_idx) can be derived based on the transform index (Tr_idx) signaled through the bitstream. The variable (Transform_idx) can be set to the value of the signaled transform index (Tr_idx). The transform index (Tr_idx) can be obtained from the bitstream based on the fact that the transform unit does not have an adaptive TU structure.
[0195] The above transformation index (Tr_idx) may indicate one or more separable transformation methods other than the non-separable transformation method. In this case, the separable transformation method indicated by the variable (Transform_idx) may be applied. That is, an inverse transformation may be performed based on a transformation kernel for separable transformation for the transformation unit.
[0196] Alternatively, the above transformation index (Tr_idx) may indicate any one of a plurality of transformation methods, including a non-separable transformation method and a separable transformation method. In this case, either the non-separable transformation method or the separable transformation method may be adaptively applied based on the variable (Transform_idx). That is, an inverse transformation may be performed based on a transformation kernel for a non-separable or separable transformation for the transformation unit.
[0197] For example, based on the value of TU_mPU_flag being 1, it can be determined that the transform unit has an adaptive TU structure. In this case, the variable (Transform_idx) can be set to a predefined value. A non-separable transform method can be applied to the transform unit without signaling the transform index (Tr_idx). On the other hand, based on the value of TU_mPU_flag being 0, it can be determined that the transform unit does not have an adaptive TU structure. In this case, the variable (Transform_idx) can be derived based on the transform index (Tr_idx) signaled through the bitstream. A non-separable or separable transform method indicated by the variable (Transform_idx) can be applied to the transform unit. The transform index can be obtained from the bitstream based on the value of TU_mPU_flag being 0.
[0198] Example 2
[0199] Based on whether the transform unit has an adaptive TU structure, a variable (NSTransform_idx) indicating whether a non-separable transform method is applied can be derived. A transform method for the transform unit can be determined based on the derived variable (NSTransform_idx).
[0200] If the value of the variable (NSTransform_idx) is 1, this may indicate that the non-separable transformation method is applied, and if the value of the variable (NSTransform_idx) is 0, this may indicate that the non-separable transformation method is not applied. Conversely, if the value of the variable (NSTransform_idx) is 1, this may indicate that the non-separable transformation method is not applied, and if the value of the variable (NSTransform_idx) is 0, this may indicate that the non-separable transformation method is applied. For the convenience of explanation, it is assumed below that the variable (NSTransform_idx) is defined as the former.
[0201] Specifically, if the transform unit has an adaptive TU structure, the variable (NSTransform_idx) may be set to 1. On the other hand, if the transform unit does not have an adaptive TU structure, the variable (NSTransform_idx) may be derived based on an index (NSTr_idx) signaled through a bitstream. Here, the index (NSTr_idx) may indicate whether a non-separable transform method is applied. For example, according to the definition of the variable (NSTransform_idx) described above, if the value of the index (NSTr_idx) is 1, this may indicate that the non-separable transform method is applied, and if the value of the index (NSTr_idx) is 0, this may indicate that the non-separable transform method is not applied. Based on the value of the index (NSTr_idx) being 1, the value of the variable (NSTransform_idx) can be derived as 1, and based on the value of the index (NSTr_idx) being 0, the value of the variable (NSTransform_idx) can be derived as 0.
[0202] For example, based on the value of TU_mPU_flag being 1, it can be determined that the transform unit has an adaptive TU structure. In this case, the variable (NSTransform_idx) can be set to 1. On the other hand, based on the value of TU_mPU_flag being 0, it can be determined that the transform unit does not have an adaptive TU structure. In this case, the variable (NSTransform_idx) can be derived to 0 or 1 based on an index (NSTr_idx) signaled through the bitstream.
[0203] Based on the value of the variable (NSTransform_idx) being 1, a non-separable transformation method can be applied to the transformation unit. Meanwhile, two or more non-separable transformation methods may be defined as non-separable transformation methods available to the transformation unit. In this case, one of the two or more non-separable transformation methods can be specified based on a transformation index (Tr_idx) signaled through a bitstream. Here, the transformation index (Tr_idx) can indicate one of the two or more pre-defined non-separable transformation methods. The specified non-separable transformation method can be applied to the transformation unit. However, if one non-separable transformation method is defined as a non-separable transformation method available to the transformation unit, signaling of the transformation index (Tr_idx) may be omitted.
[0204] Based on the value of the variable (NSTransform_idx) being 0, a non-separable transformation method may not be applied to the transform unit. Instead, a separable transformation method may be applied to the transform unit. Meanwhile, two or more separable transformation methods may be defined as separable transformation methods available to the transform unit. In this case, one of the two or more separable transformation methods may be specified based on a transformation index (Tr_idx) signaled through a bitstream. Here, the transformation index (Tr_idx) may indicate one of the two or more pre-defined separable transformation methods. The specified separable transformation method may be applied to the transform unit. However, if one separable transformation method is defined as a separable transformation method available to the transform unit, signaling of the transformation index (Tr_idx) may be omitted.
[0205] Whether the transformation index (Tr_idx) is an index for a non-separable transformation method can be determined based on at least one of a variable (NSTransform_idx) or whether the transformation unit has an adaptive TU structure (e.g., TU_mPU_flag).
[0206] For example, if the value of the variable (NSTransform_idx) indicates that a non-separable transformation method is applied (i.e., NSTransform_idx=1), the signaled transformation index (Tr_idx) may correspond to a transformation index indicating one of two or more non-separable transformation methods. On the other hand, if the value of the variable (NSTransform_idx) indicates that a separable transformation method is applied (i.e., NSTransform_idx=0), the signaled transformation index (Tr_idx) may correspond to a transformation index indicating one of two or more separable transformation methods.
[0207] Alternatively, the transformation index (Tr_idx) for the aforementioned non-separable transformation method and the transformation index (Tr_idx) for the separable transformation method may be defined as separate syntax elements. Hereinafter, the transformation index (Tr_idx) for the non-separable transformation method will be referred to as the non-separable transformation index, and the transformation index (Tr_idx) for the separable transformation method will be referred to as the separable transformation index.
[0208] Either a non-separable transform index or a separable transform index can be optionally signaled based on at least one of a variable (NSTransform_idx) or whether the transform unit has an adaptive TU structure (e.g., TU_mPU_flag).
[0209] For example, if the value of the variable (NSTransform_idx) indicates that a non-separable transformation method is applied (i.e., NSTransform_idx=1), the non-separable transformation index may be signaled. Conversely, if the value of the variable (NSTransform_idx) indicates that a separable transformation method is applied (i.e., NSTransform_idx=0), the separable transformation index may be signaled.
[0210] For transform units with an adaptive TU structure, the statistical characteristics of the residual signal can be well represented through a non-separable transform. Therefore, when either a non-separable transform or a separable transform can be selectively utilized, as described above, by implicitly signaling the application of a non-separable transform, the amount of information that must be transmitted to determine the variable (NSTransform_idx) can be minimized per TU, thereby improving encoding efficiency.
[0211] The present disclosure relates to a method for signaling a transform method for efficiently decoding a transform unit having an adaptive TU structure. In particular, the inverse transform according to the present disclosure may be composed of a secondary inverse transform and a primary inverse transform, and the method according to the present disclosure may be applied to the secondary inverse transform. Here, the primary inverse transform may be performed based on a separable transform. However, the present disclosure is not limited thereto, and the primary inverse transform may also be performed based on a non-separable transform. The secondary inverse transform may be performed based on either a non-separable transform or a separable transform.
[0212] In an encoding device, input data is a residual signal (or residual samples), and output data is transform coefficients. Specifically, primary transform coefficients can be generated based on a primary transform on the residual signal, and secondary transform coefficients can be generated based on a secondary transform on the primary transform coefficients. Quantized transform coefficients can be derived based on quantization on the secondary transform coefficients.
[0213] In the decoding device, the transform coefficients become input data and the residual signal becomes output data. Specifically, the primary transform coefficients can be generated based on the secondary inverse transform on the inverse quantized transform coefficients, and the residual signal can be generated based on the primary inverse transform on the primary transform coefficients. The inverse quantized transform coefficients can correspond to the secondary transform coefficients in the encoding device.
[0214] A decoding device may define a plurality of transform methods for secondary inverse transform, and at least one of the plurality of transform methods may be selectively used. Here, the plurality of transform methods may include at least one of a non-separable transform method or a separable transform method. As the non-separable transform method, the first to Nth non-separable transform method(s) may be defined, and as the separable transform method, the first to Mth separable transform method(s) may be defined. Here, N and M may be integers greater than or equal to 1. N and M may be the same value or different values.
[0215] The transformation method according to the present disclosure can be determined based on whether the transformation unit has an adaptive TU structure, and the method for determining whether the transformation unit has an adaptive TU structure is as described above.
[0216] Example 1
[0217] Based on whether the transform unit has an adaptive TU structure, a variable (STransform_idx) indicating a transform method for the secondary inverse transform of the transform unit can be derived. The transform method of the transform unit can be determined based on the derived variable (STransform_idx).
[0218] Specifically, when the transform unit has an adaptive TU structure, the variable (STransform_idx) can be set to a predefined value. Here, the predefined value is a value defined identically in the encoding device and the decoding device, and can be a value indicating a non-separable transform method. When the transform unit has an adaptive TU structure, the non-separable transform method can be applied without signaling the transform index (STr_idx) described later. That is, a secondary inverse transform can be performed based on a transform kernel for non-separable transform for the transform unit.
[0219] If the transform unit does not have an adaptive TU structure, the variable (STransform_idx) can be derived based on the transform index (STr_idx) signaled through the bitstream. The variable (STransform_idx) can be set to the value of the signaled transform index (STr_idx). The transform index (STr_idx) can be obtained from the bitstream based on the fact that the transform unit does not have an adaptive TU structure.
[0220] The above transformation index (STr_idx) may indicate one or more separable transformation methods other than the non-separable transformation method. In this case, the separable transformation method indicated by the variable (STransform_idx) may be applied. That is, a secondary inverse transformation may be performed based on a transformation kernel for the separable transformation for the transformation unit.
[0221] Alternatively, the above transformation index (STr_idx) may indicate any one of a plurality of transformation methods, including a non-separable transformation method and a separable transformation method. In this case, either the non-separable transformation method or the separable transformation method may be adaptively applied based on the variable (STransform_idx). That is, a secondary inverse transformation may be performed based on a transformation kernel for a non-separable or separable transformation for the transformation unit.
[0222] For example, based on the value of TU_mPU_flag being 1, it can be determined that the transform unit has an adaptive TU structure. In this case, the variable (STransform_idx) can be set to a predefined value. A non-separable transform method can be applied to the transform unit without signaling the transform index (STr_idx).
[0223] On the other hand, based on the value of TU_mPU_flag being 0, it can be determined that the transform unit does not have an adaptive TU structure. In this case, the variable (STransform_idx) can be derived based on the transform index (STr_idx) signaled through the bitstream. A non-separable or separable transform method indicated by the variable (STransform_idx) can be applied to the transform unit. The transform index (STr_idx) can be obtained from the bitstream based on the value of TU_mPU_flag being 0.
[0224] Example 2
[0225] Based on whether the transform unit has an adaptive TU structure, a variable (STransform_idx) indicating a transform method to be applied to the transform unit among the non-separable transform methods available to the transform unit can be derived. The transform method of the transform unit can be determined based on the derived variable (STransform_idx).
[0226] Depending on whether the transform unit has an adaptive TU structure, the sets of non-separable transformation methods available to the transform unit may differ. However, if the number of non-separable transformation methods available to the transform unit is 1 depending on whether the transform unit has an adaptive TU structure, the transformation method for the transform unit may be set to the available non-separable transformation method, and the process of deriving the variable (STransform_idx) may be omitted.
[0227] Specifically, when the transform unit has an adaptive TU structure, the transform method applied to the transform unit can be determined from a first non-separable transform set. Here, the first non-separable transform set can be defined as a set of non-separable transform methods available for the transform unit having the adaptive TU structure. The first non-separable transform set can include at least one of the first to Nth non-separable transform methods described above, and N can be an integer greater than or equal to 1.
[0228] On the other hand, if the transform unit does not have an adaptive TU structure, the transform method applied to the transform unit can be determined from a second non-separable transform set. Here, the second non-separable transform set can be defined as a set of non-separable transform methods available for the transform unit that does not have an adaptive TU structure. The second non-separable transform set can include at least one of the first to Nth non-separable transform methods, and N can be an integer greater than or equal to 1.
[0229] For example, based on the value of TU_mPU_flag being 1, it can be determined that the transform unit has an adaptive TU structure. In this case, the transform method applied to the transform unit can be determined from the first non-separable transform set. On the other hand, based on the value of TU_mPU_flag being 0, it can be determined that the transform unit does not have an adaptive TU structure. In this case, the transform method applied to the transform unit can be determined from the second non-separable transform set.
[0230] When the first non-separable transform set includes two or more non-separable transform methods, one of the two or more non-separable transform methods can be selected based on a variable (STransform_idx). The variable (STransform_idx) can be derived based on a first transform index (STr_idx1) signaled through a bitstream. For example, the variable (STransform_idx) can be set to the value of the signaled first transform index (STr_idx1). The first transform index (STr_idx1) can indicate one of the two or more non-separable transform methods included in the first non-separable transform set. The first transform index (STr_idx1) can be signaled based on whether the transform unit has an adaptive TU structure (or whether the value of TU_mPU_flag is 1). If the first non-separable transformation set includes one non-separable transformation method, signaling of the first transformation index (STr_idx1) may be omitted.
[0231] If the second non-separable transform set includes two or more non-separable transform methods, one of the two or more non-separable transform methods may be selected based on a variable (STransform_idx). The variable (STransform_idx) may be derived based on a second transform index (STr_idx2) signaled through a bitstream. For example, the variable (STransform_idx) may be set to the value of the signaled second transform index (STr_idx2). The second transform index (STr_idx2) may indicate one of the two or more non-separable transform methods included in the second non-separable transform set. The second transform index (STr_idx2) may be signaled based on the transform unit not having an adaptive TU structure (or the value of TU_mPU_flag being 0). If the second non-separable transformation set includes one non-separable transformation method, signaling of the second transformation index (STr_idx2) may be omitted.
[0232] The number of non-separable transformation methods belonging to the first non-separable transformation group may be different from the number of non-separable transformation methods belonging to the second non-separable transformation group. At least one of the non-separable transformation methods belonging to the first non-separable transformation group may not be included in the second non-separable transformation group. The number of non-separable transformation methods belonging to the first non-separable transformation group may be less than the number of non-separable transformation methods belonging to the second non-separable transformation group. However, the present invention is not limited thereto, and the number of transform kernels corresponding to the non-separable transformation method of the first non-separable transformation group may be greater than the number of transform kernels corresponding to the non-separable transformation method of the second non-separable transformation group. Alternatively, the first and second non-separable transformation groups may each include the same number of non-separable transformation methods.
[0233] For example, for a transform unit having an adaptive TU structure (or when the value of TU_mPU_flag is 1), the number of non-separable transform methods available to the transform unit may be 2. In this case, the first transform index (STr_idx1) may be represented by a flag having a value of 0 or 1. If the transform unit does not have an adaptive TU structure (or when the value of TU_mPU_flag is 0), the number of non-separable transform methods available to the transform unit may be 4. In this case, the value of the second transform index (STr_idx2) may have a value of 0, 1, 2, or 3.
[0234] The present disclosure relates to a method for selecting a transform kernel for a non-separable transform (NST) of a transform unit having an adaptive TU structure. Here, the non-separable transform may refer to a non-separable primary transform or a non-separable secondary transform. The method for selecting a transform kernel described below may be applied to a transform unit to which the aforementioned non-separable transform method is applied.
[0235] Non-separable transform (NST) is a technique for improving compression efficiency by using a transform kernel that reflects the statistical characteristics of the input signal well. The transform set can be determined based on the factors that determine the statistical characteristics of the input signal. From an encoding perspective, when a non-separable transform is used as a primary transform, the input signal of the non-separable transform can be regarded as a residual signal obtained by subtracting the prediction signal from the original signal. When a non-separable transform is used in a secondary transform applied after the primary transform, the input signal of the non-separable transform can be regarded as the primary transform coefficients generated through the primary transform. From a decoding perspective, when a non-separable transform is used in a primary inverse transform, the input signal of the non-separable transform can be regarded as the primary transform coefficients (or dequantized primary transform coefficients) transmitted through the bitstream. When a non-separable transform is used in a secondary inverse transform, the input signal of the non-separable transform can be regarded as the secondary transform coefficients (or dequantized secondary transform coefficients) transmitted through the bitstream.
[0236] For a transform unit with an adaptive TU structure, a transform set for non-separable transform can be determined by considering the following characteristics. Any one of a plurality of transform kernels belonging to the determined transform set can be selected, and an inverse transform can be performed based on the selected transform kernel.
[0237] (Characteristic 1) A set of transformations for non-separable transformations can be determined based on the partitioning mode applied to the coding unit corresponding to the transformation unit.
[0238] Here, the segmentation mode may belong to a segmentation mode to which an adaptive TU structure can be applied. The segmentation mode to which the adaptive TU structure can be applied may include at least one of quad segmentation, binary segmentation, ternary segmentation, and geometric segmentation. The geometric segmentation may be a segmentation mode that segments based on one or more segmentation lines having a predetermined angle. Various types of geometric segmentation may be defined based on at least one of the angle of the segmentation line, the number of segmentation lines, or the spacing between the segmentation lines (or, split density). The main reason why segmentation lines with various angles can be applied to the geometric segmentation is that the non-separable transform used in conjunction with the corresponding segmentation mode has relatively clear characteristics of the input signal generated based on the corresponding segmentation mode (i.e., the statistical characteristics of the residual image or the primary transform coefficients generated through the given split angle).
[0239] For example, FIG. 5 illustrates various segmentation modes based on geometric segmentation, as an embodiment of the present disclosure. As illustrated in FIG. 5, the segmentation line may have an angle of 45 degrees or 135 degrees. However, this is merely an example, and segmentation lines with different angles may be utilized. Furthermore, segmentation may be performed based on two or more segmentation lines positioned at a predetermined interval.
[0240] (Characteristic 2) A transform set for non-separable transform can be determined based on the splitting direction of a plurality of coding units included in a transform unit having an adaptive TU structure.
[0241] Even when the same partitioning mode is applied, the statistical characteristics of the residual signal can be reflected depending on the partitioning direction of multiple coding units included in one transform unit, and thus the transform set can be determined based on the partitioning direction. For example, when a first coding unit corresponding to a transform unit is partitioned into two second coding units based on binary partitioning, the transform set can be determined based on whether the partitioning direction of the binary partitioning is horizontal.
[0242] (Characteristic 3) A transform set for non-separable transform can be determined based on at least one of the number or spacing of division lines that divide a transform unit having an adaptive TU structure.
[0243] Since the statistical characteristics of the residual signal may vary depending on at least one of the number or spacing of the dividing lines that divide the transform unit, a transform set may be determined based on at least one of the number or spacing of the dividing lines. For example, FIG. 6 illustrates various types of dividing modes according to the spacing of the dividing lines, as an embodiment according to the present disclosure. The left drawing in FIG. 6 illustrates a dividing mode based on one dividing line having an angle of 45 degrees. The middle drawing in FIG. 6 illustrates a dividing mode based on three dividing lines having angles of 45 degrees. The right drawing in FIG. 6 illustrates a dividing mode based on five dividing lines having angles of 45 degrees. As illustrated in FIG. 6, all three cases are divided at an angle of 45 degrees, but since the statistical characteristics of the residual block may differ depending on the number and spacing of the dividing lines, different transform sets may be determined for the three cases.
[0244] The transformation set may be determined based on any one of the aforementioned characteristics 1 to 3. Alternatively, the transformation set may be determined based on a combination of at least two of the aforementioned characteristics 1 to 3. Furthermore, the transformation method in the aforementioned disclosure may be understood as being replaced with transformation, transformation process, or transformation type.
[0245] FIG. 7 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.
[0246] Referring to FIG. 7, the decoding device (300) may include a transform coefficient derivation unit (700) and a residual sample generation unit (710). The transform coefficient derivation unit (700) and the residual sample generation unit (710) may be provided in the residual processing unit (320) of FIG. 3.
[0247] The transform coefficient derivation unit (700) can derive transform coefficients of a transform unit based on residual information of the transform unit. Here, the transform unit may be derived based on an adaptive TU structure, and for this purpose, a transform unit determination unit (not shown) may be provided in the residual processing unit (320).
[0248] A transform unit decision unit (not shown) can determine whether an adaptive TU structure is applied to a coding unit to which a predetermined division mode is applied, and can derive a transform unit based on whether the adaptive TU structure is applied. This is as described with reference to FIG. 4, and any duplicate description will be omitted here.
[0249] Information (TU_mPU_flag) regarding whether or not the adaptive TU structure is applied can be explicitly signaled through the bitstream, and the signaling method is as described with reference to FIG. 4. The entropy decoding unit (310) can obtain TU_mPU_flag from the bitstream based on the signaling method described above.
[0250] The residual sample generation unit (710) can derive residual samples of the transform unit based on the transform coefficients of the transform unit. The residual sample generation unit (710) can derive residual samples based on at least one of inverse quantization or inverse transformation of the transform coefficients.
[0251] The residual sample generation unit (710) can determine a transformation method for inverse transformation, as described with reference to FIG. 4. If it is determined that a non-separable transformation method is to be applied to the transformation unit, the residual sample generation unit (710) can select a transformation kernel for non-separable transformation, as described with reference to FIG. 4. FIG. 8 illustrates an image encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0252] Referring to FIG. 8, residual samples of the conversion unit can be derived (S800).
[0253] A transform unit, which is a unit for encoding residual samples, can be determined, and residual samples for the transform unit can be derived.
[0254] The above transform unit (or the size of the transform unit) may be determined based on an adaptive TU structure. That is, a first coding unit may be split into a plurality of second coding units based on a predetermined split mode. It may be determined whether an adaptive TU structure is applied to the first coding unit. The transform unit may be determined based on the determination of whether an adaptive TU structure is applied to the first coding unit. A method for determining a transform unit based on an adaptive TU structure has been described with reference to FIG. 4, and a duplicate description thereof will be omitted herein.
[0255] A flag (TU_mPU_flag) related to whether an adaptive TU structure is applied may be encoded in the bitstream. Whether an adaptive TU structure is applied may be determined based on at least one of a flag (TU_mPU_flag) related to whether an adaptive TU structure is applied, a split flag (split_cu_flag) indicating whether a first coding unit is split, split information for block splitting of the first coding unit, or the size of the first coding unit, as described with reference to FIG. 4.
[0256] TU_mPU_flag according to the present disclosure may be encoded and signaled based on at least one of a split flag (split_cu_flag) indicating whether the first coding unit is split, split information for block splitting of the first coding unit, or the size of the first coding unit. If TU_mPU_flag is not signaled, the value of TU_mPU_flag may be derived as 0. The signaling method of TU_mPU_flag is as described with reference to FIG. 4.
[0257] Referring to FIG. 8, the transform coefficients of the transform unit can be derived based on the residual samples of the transform unit (S810).
[0258] The transform coefficients of the transform unit can be derived based on at least one of transform or quantization of residual samples of the transform unit.
[0259] The above transformation can be performed based on at least one of a separable transform and a non-separable transform.
[0260] The present disclosure describes a method for signaling a transform method for efficiently encoding a transform unit having an adaptive TU structure.
[0261] An encoding device may define a plurality of transformation methods for transformation, and at least one of the plurality of transformation methods may be selectively used. The plurality of transformation methods may include at least one of a non-separable transformation method or a separable transformation method. First to Nth non-separable transformation method(s) may be defined as the non-separable transformation method(s), and first to Mth separable transformation method(s) may be defined as the separable transformation method(s). Here, N and M may be integers greater than or equal to 1. N and M may be the same value or different values. The non-separable transformation may be referred to as a non-separable primary transformation (NSPT) or a non-separable secondary transformation (NSST).
[0262] The transformation method according to the present disclosure can be determined based on whether a transformation unit has an adaptive TU structure. Whether a transformation unit has an adaptive TU structure can be determined based on whether an adaptive TU structure is applied to a first coding unit, and the method for determining whether an adaptive TU structure is applied is as described above.
[0263] Example 1
[0264] Based on whether the transform unit has an adaptive TU structure, a variable (Transform_idx) indicating a transformation method to be applied to the transform unit can be derived. A transformation method for the transform unit can be determined based on the derived variable (Transform_idx).
[0265] Specifically, when the transform unit has an adaptive TU structure, the variable (Transform_idx) can be set to a predefined value. Here, the predefined value is a value defined identically to the encoding device and the decoding device, and can be a value indicating a non-separable transform method. When the transform unit has an adaptive TU structure, the non-separable transform method can be applied without encoding the transform index (Tr_idx) described later. That is, the transform can be performed based on a transform kernel for non-separable transform for the transform unit.
[0266] If the transform unit does not have an adaptive TU structure, one or more separable transform methods other than a non-separable transform method can be selected, and the selected separable transform method can be applied to the transform unit. That is, the transform can be performed based on a transform kernel for the separable transform for the transform unit. A transform index (Tr_idx) indicating the selected separable transform method can be encoded in the bitstream. The transform index (Tr_idx) can be encoded in the bitstream based on the fact that the transform unit does not have an adaptive TU structure.
[0267] For example, if the transform unit does not have an adaptive TU structure, one or more separable transform methods other than the non-separable transform method can be selected to derive the variable (Transform_idx). The separable transform method indicated by the derived variable (Transform_idx) can be applied to the transform unit.
[0268] Alternatively, when the transform unit does not have an adaptive TU structure, any one of a plurality of transform methods including a non-separable transform method and a separable transform method may be selected, and the selected non-separable or separable transform method may be applied to the transform unit. That is, the transform may be performed based on a transform kernel for the non-separable or separable transform for the transform unit. A transform index (Tr_idx) indicating the selected non-separable or separable transform method may be encoded in the bitstream. The transform index (Tr_idx) may be encoded in the bitstream based on the transform unit not having an adaptive TU structure.
[0269] For example, if the transform unit does not have an adaptive TU structure, a variable (Transform_idx) can be derived by selecting any one of a plurality of transform methods, including a non-separable transform method and a separable transform method. Based on the derived variable (Transform_idx), either the non-separable transform method or the separable transform method can be adaptively applied.
[0270] For example, based on the value of TU_mPU_flag being 1, it can be determined that the transform unit has an adaptive TU structure. In this case, the variable (Transform_idx) can be set to a predefined value. A non-separable transform method can be applied to the transform unit without encoding the transform index (Tr_idx).
[0271] On the other hand, based on the value of TU_mPU_flag being 0, it can be determined that the transform unit does not have an adaptive TU structure. In this case, any one or more separable transform methods excluding the non-separable transform method can be selected, and the selected separable transform method can be applied to the transform unit. A transform index (Tr_idx) indicating the selected separable transform method can be encoded in the bitstream. The transform index (Tr_idx) can be encoded in the bitstream based on the value of TU_mPU_flag being 0.
[0272] For example, based on the value of TU_mPU_flag being 0, one or more separable transformation methods other than the non-separable transformation method can be selected to derive a variable (Transform_idx). The separable transformation method indicated by the derived variable (Transform_idx) can be applied to the transformation unit.
[0273] Alternatively, based on the value of TU_mPU_flag being 0, it may be determined that the transform unit does not have an adaptive TU structure. In this case, any one of a plurality of transform methods including a non-separable transform method and a separable transform method may be selected, and the selected non-separable or separable transform method may be applied to the transform unit. A transform index (Tr_idx) indicating the selected non-separable or separable transform method may be encoded in the bitstream. The transform index (Tr_idx) may be encoded in the bitstream based on the value of TU_mPU_flag being 0.
[0274] For example, based on the value of TU_mPU_flag being 0, one of a plurality of transformation methods, including a non-separable transformation method and a separable transformation method, can be selected to derive a variable (Transform_idx). The non-separable or separable transformation method indicated by the derived variable (Transform_idx) can be applied to the transformation unit.
[0275] Example 2
[0276] Based on whether the transform unit has an adaptive TU structure, a variable (NSTransform_idx) indicating whether a non-separable transform method is applied can be derived. A transform method for the transform unit can be determined based on the derived variable (NSTransform_idx). The variable (NSTransform_idx) is as described with reference to FIG. 4.
[0277] Specifically, if the transform unit has an adaptive TU structure, the variable (NSTransform_idx) may be set to 1. On the other hand, if the transform unit does not have an adaptive TU structure, it may be determined whether a non-separable transform method is applied to the transform unit. Based on the determination, the variable (NSTransform_idx) may be derived as 0 or 1. If the transform unit does not have an adaptive TU structure, an index (NSTr_idx) indicating whether a non-separable transform method is applied to the transform unit may be encoded in the bitstream.
[0278] Based on the value of the variable (NSTransform_idx) being 1, a non-separable transformation method can be applied to the transformation unit. Meanwhile, two or more non-separable transformation methods may be defined as non-separable transformation methods available to the transformation unit. In this case, the transformation can be performed based on any one of the two or more pre-defined non-separable transformation methods. A transformation index (Tr_idx) indicating any one of the two or more pre-defined non-separable transformation methods can be encoded in the bitstream. However, if one non-separable transformation method is defined as a non-separable transformation method available to the transformation unit, encoding of the transformation index (Tr_idx) may be omitted.
[0279] Based on the value of the variable (NSTransform_idx) being 0, a non-separable transformation method may not be applied to the transform unit. Instead, a separable transformation method may be applied to the transform unit. Meanwhile, two or more separable transformation methods may be defined as separable transformation methods available to the transform unit. In this case, the transformation may be performed based on any one of the two or more pre-defined separable transformation methods. A transformation index (Tr_idx) indicating any one of the two or more pre-defined separable transformation methods may be encoded in the bitstream. However, if one non-separable transformation method is defined as a separable transformation method available to the transform unit, encoding of the transformation index (Tr_idx) may be omitted.
[0280] For example, based on the value of TU_mPU_flag being 1, it can be determined that the transform unit has an adaptive TU structure. In this case, the variable (NSTransform_idx) can be set to 1. On the other hand, based on the value of TU_mPU_flag being 0, it can be determined that the transform unit does not have an adaptive TU structure. In this case, the variable (NSTransform_idx) can be derived by determining whether a non-separable transform method is applied to the transform unit. Based on the value of TU_mPU_flag being 0, an index (NSTr_idx) indicating whether a non-separable transform method is applied to the transform unit can be encoded in the bitstream.
[0281] Whether the transformation index (Tr_idx) is an index for a non-separable transformation method can be determined based on at least one of a variable (NSTransform_idx) or whether the transformation unit has an adaptive TU structure (e.g., TU_mPU_flag), as discussed with reference to FIG. 4.
[0282] Alternatively, the transformation index (Tr_idx) for the non-separable transformation method described above and the transformation index (Tr_idx) for the separable transformation method can be defined as separate syntax elements, as described with reference to FIG. 4.
[0283] The present disclosure relates to a method for signaling a transform method for efficiently encoding a transform unit having an adaptive TU structure. In particular, the transform according to the present disclosure may be composed of a primary transform and a secondary transform, and the method according to the present disclosure may be applied to the secondary transform. Here, the primary transform may be performed based on a separable transform. However, the present disclosure is not limited thereto, and the primary transform may also be performed based on a non-separable transform. The secondary transform may be performed based on a non-separable transform or a separable transform.
[0284] An encoding device may define a plurality of transformation methods for secondary transformation, and at least one of the plurality of transformation methods may be selectively used. Here, the plurality of transformation methods may include at least one of a non-separable transformation method or a separable transformation method. As the non-separable transformation method, the first to Nth non-separable transformation method(s) may be defined, and as the separable transformation method, the first to Mth separable transformation method(s) may be defined. Here, N and M may be integers greater than or equal to 1. N and M may be the same value or different values.
[0285] The transformation method according to the present disclosure can be determined based on whether the transformation unit has an adaptive TU structure, and the method for determining whether the transformation unit has an adaptive TU structure is as described above.
[0286] Example 1
[0287] Based on whether the transform unit has an adaptive TU structure, a variable (STransform_idx) indicating a transform method for a secondary transform of the transform unit can be derived. The transform method of the transform unit can be determined based on the derived variable (STransform_idx).
[0288] Specifically, when the transform unit has an adaptive TU structure, the variable (STransform_idx) can be set to a predefined value. Here, the predefined value is a value defined identically to the encoding device and the decoding device, and can be a value indicating a non-separable transform method. When the transform unit has an adaptive TU structure, the non-separable transform method can be applied without encoding the transform index (STr_idx) described later. That is, a secondary transform can be performed based on a transform kernel for non-separable transform for the transform unit.
[0289] On the other hand, if the transform unit does not have an adaptive TU structure, one or more separable transform methods other than a non-separable transform method can be selected, and the selected separable transform method can be applied to the transform unit. That is, a secondary transform can be performed on the transform unit based on a transform kernel for the separable transform. A transform index (STr_idx) indicating the selected separable transform method can be encoded in the bitstream. The transform index (STr_idx) can be encoded in the bitstream based on the fact that the transform unit does not have an adaptive TU structure.
[0290] For example, if the transform unit does not have an adaptive TU structure, one or more separable transform methods other than the non-separable transform method can be selected to derive the variable (STransform_idx). The separable transform method indicated by the derived variable (STransform_idx) can be applied to the transform unit.
[0291] Alternatively, when the transform unit does not have an adaptive TU structure, any one of a plurality of transform methods including a non-separable transform method and a separable transform method may be selected, and the selected non-separable or separable transform method may be applied to the transform unit. That is, a secondary transform may be performed based on a transform kernel for the non-separable or separable transform for the transform unit. A transform index (STr_idx) indicating the selected non-separable or separable transform method may be encoded in the bitstream. The transform index (STr_idx) may be encoded in the bitstream based on the transform unit not having an adaptive TU structure.
[0292] For example, if the transform unit does not have an adaptive TU structure, a variable (STransform_idx) can be derived by selecting any one of a plurality of transform methods, including a non-separable transform method and a separable transform method. The non-separable or separable transform method indicated by the derived variable (STransform_idx) can be applied to the transform unit.
[0293] For example, based on the value of TU_mPU_flag being 1, it can be determined that the transform unit has an adaptive TU structure. In this case, the variable (STransform_idx) can be set to a predefined value. Based on the value of TU_mPU_flag being 1, a non-separable transform method can be applied to the transform unit without encoding the transform index (STr_idx). That is, a secondary transform can be performed based on a transform kernel for non-separable transform.
[0294] On the other hand, based on the value of TU_mPU_flag being 0, it can be determined that the transform unit does not have an adaptive TU structure. In this case, any one or more separable transform methods excluding the non-separable transform method can be selected, and the selected separable transform method can be applied to the transform unit. That is, a secondary transform can be performed based on a transform kernel for the separable transform for the transform unit. A transform index (STr_idx) indicating the selected separable transform method can be encoded in the bitstream. The transform index (STr_idx) can be encoded in the bitstream based on the value of TU_mPU_flag being 0.
[0295] For example, based on the value of TU_mPU_flag being 0, one or more separable transformation methods other than the non-separable transformation method can be selected to derive a variable (STransform_idx), and the separable transformation method indicated by the variable (STransform_idx) can be applied to the transformation unit.
[0296] Alternatively, based on the value of TU_mPU_flag being 0, it may be determined that the transform unit does not have an adaptive TU structure. In this case, any one of a plurality of transform methods including a non-separable transform method and a separable transform method may be selected, and the selected non-separable or separable transform method may be applied to the transform unit. That is, the transform may be performed based on a transform kernel for the non-separable or separable transform. A transform index (STr_idx) indicating the selected non-separable or separable transform method may be encoded in the bitstream. The transform index (STr_idx) may be encoded in the bitstream based on the value of TU_mPU_flag being 0.
[0297] For example, when the value of TU_mPU_flag is 0, one of multiple transformation methods including a non-separable transformation method and a separable transformation method can be selected to derive a variable (STransform_idx), and the non-separable or separable transformation method indicated by the variable (STransform_idx) can be applied to the transformation unit.
[0298] Example 2
[0299] Based on whether the transform unit has an adaptive TU structure, a variable (STransform_idx) indicating a transform method to be applied to the transform unit among the non-separable transform methods available to the transform unit can be derived. The transform method of the transform unit can be determined based on the derived variable (STransform_idx).
[0300] Depending on whether the transform unit has an adaptive TU structure, the sets of non-separable transformation methods available to the transform unit may differ. However, if the number of non-separable transformation methods available to the transform unit is 1 depending on whether the transform unit has an adaptive TU structure, the transformation method for the transform unit may be set to the available non-separable transformation method, and the process of deriving the variable (STransform_idx) may be omitted.
[0301] Specifically, if the transform unit has an adaptive TU structure, the transform method applied to the transform unit can be determined from the first non-separable transform set. If the transform unit does not have an adaptive TU structure, the transform method applied to the transform unit can be determined from the second non-separable transform set.
[0302] For example, based on the value of TU_mPU_flag being 1, it can be determined that the transform unit has an adaptive TU structure. In this case, the transform method applied to the transform unit can be determined from the first non-separable transform set. Based on the value of TU_mPU_flag being 0, it can be determined that the transform unit does not have an adaptive TU structure. In this case, the transform method applied to the transform unit can be determined from the second non-separable transform set.
[0303] When the first non-separable transform set includes two or more non-separable transform methods, one of the two or more non-separable transform methods may be selected, and the selected non-separable transform method may be set as a transform method of the transform unit. A first transform index (STr_idx1) indicating the selected non-separable transform method may be encoded in a bitstream. The first transform index (STr_idx1) may be encoded based on whether the transform unit has an adaptive TU structure (or, whether the value of TU_mPU_flag is 1). When the first non-separable transform set includes one non-separable transform method, encoding of the first transform index (STr_idx1) may be omitted.
[0304] If the second non-separable transform set includes two or more non-separable transform methods, one of the two or more non-separable transform methods may be selected, and the selected non-separable transform method may be set as a transform method of the transform unit. A second transform index (STr_idx2) indicating the selected non-separable transform method may be encoded. The second transform index (STr_idx2) may be encoded based on whether the transform unit does not have an adaptive TU structure (or, the value of TU_mPU_flag is 0). If the second non-separable transform set includes one non-separable transform method, encoding of the second transform index (STr_idx2) may be omitted.
[0305] A transformation kernel for non-separable transformation (NST) of a transformation unit having an adaptive TU structure can be adaptively selected, and the method for selecting the transformation kernel is as described with reference to FIG. 4.
[0306] Referring to FIG. 8, a bitstream can be generated by encoding residual information regarding the transform coefficients of the transform unit (S820).
[0307] FIG. 9 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.
[0308] Referring to FIG. 9, the encoding device (200) may include a residual sample derivation unit (900), a transform coefficient derivation unit (910), and a residual information encoding unit (920). The residual sample derivation unit (900) and the transform coefficient derivation unit (910) may be provided in the residual processing unit (230) of FIG. 2. The residual information encoding unit (920) may be provided in the entropy encoding unit (240).
[0309] The residual sample derivation unit (900) can derive residual samples of a transformation unit. Here, the transformation unit may be derived based on an adaptive TU structure, and for this purpose, a transformation unit determination unit (not shown) may be provided in the residual processing unit (230).
[0310] A transform unit determination unit (not shown) can determine a transform unit, which is a unit for encoding residual samples. Specifically, it can determine whether an adaptive TU structure is applied to a coding unit to which a predetermined division mode is applied, and can determine a transform unit based on whether the adaptive TU structure is applied. This is as described with reference to FIG. 4, and any duplicate description will be omitted here.
[0311] A flag (TU_mPU_flag) related to whether an adaptive TU structure is applied can be encoded in the bitstream, and the signaling method of TU_mPU_flag is as described with reference to FIG. 4. The entropy encoding unit (240) can encode TU_mPU_flag in the bitstream based on the signaling method described above.
[0312] The transform coefficient derivation unit (910) can derive transform coefficients of the transform unit based on residual samples of the transform unit. The transform coefficients of the transform unit can be derived based on at least one of transformation or quantization of the residual samples of the transform unit.
[0313] The transform coefficient derivation unit (910) can determine a transform method for the transform, as discussed with reference to FIG. 8. If it is determined that a non-separable transform method is to be applied to the transform unit, the transform coefficient derivation unit (910) can select a transform kernel for the non-separable transform, as discussed with reference to FIG. 8.
[0314] The residual information encoding unit (920) can generate a bitstream by encoding residual information regarding the transform coefficients of the transform unit.
[0315] In the embodiments described above, the methods are described based on a flowchart as a series of steps or blocks. However, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this document.
[0316] The method according to the embodiments of the present document described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0317] When the embodiments in this document are implemented as software, the above-described method can be implemented as a module (process, function, etc.) that performs the above-described function. The module can be stored in memory and executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), another chipset, logic circuit, and / or data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored on a digital storage medium.
[0318] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0319] In addition, the processing method to which the embodiment(s) of the present specification are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present specification can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0320] Additionally, the embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0321] FIG. 10 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0322] Referring to FIG. 10, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0323] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.
[0324] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0325] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system.
[0326] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0327] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0328] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0329] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.
Claims
1. A step of deriving transform coefficients of a transform unit from a bitstream; and A step of generating residual samples of the transform unit based on inverse transformation of the above transform coefficients, A method wherein the above inverse transformation is performed based on whether the transformation unit has an adaptive TU structure.
2. In paragraph 1, A method in which whether the above transformation unit has the adaptive TU structure is determined based on a flag related to whether the adaptive TU structure is applied.
3. In paragraph 1, A method wherein the inverse transformation is performed based on a non-separable transformation method, based on the above transformation unit having the adaptive TU structure.
4. In paragraph 1, A method wherein the non-separable transformation method applied to the above transformation unit is determined based on a transformation index indicating one of the non-separable transformation methods available to the transformation unit.
5. In paragraph 1, A method wherein the inverse transformation is performed based on a separation transformation method, based on the fact that the above transformation unit does not have the adaptive TU structure.
6. In paragraph 5, A method wherein the separation transformation method applied to the above transformation unit is determined based on a transformation index indicating one of the separation transformation methods available to the transformation unit.
7. In paragraph 1, An index indicating whether a non-separable transformation method is applied to the transformation unit is signaled based on the transformation unit not having the adaptive TU structure, A method in which a conversion method for the conversion unit is determined based on the above index.
8. In paragraph 1, Based on the above transformation unit having the adaptive TU structure, the transformation method applied to the transformation unit is determined from a first non-separable transformation set, A method wherein a transformation method applied to the transformation unit is determined from a second non-separable transformation set, based on the transformation unit not having the adaptive TU structure.
9. In paragraph 8, A method wherein the number of non-separable transformation methods belonging to the first non-separable transformation set is different from the number of non-separable transformation methods belonging to the second non-separable transformation set.
10. In paragraph 3, A method in which a transformation set for non-separable transformation of the above transformation unit is determined based on at least one of a division mode of a coding unit corresponding to the transformation unit, a division direction of the coding unit, the number of division lines for dividing the transformation unit, or an interval of division lines for dividing the transformation unit.
11. A step of deriving residual samples of the conversion unit; A step of deriving transform coefficients of the transform unit based on the transform for the residual samples; and A step of encoding residual information regarding the above transformation coefficients, A method wherein the above transformation is performed based on whether the transformation unit has an adaptive TU structure.
12. A computer-readable storage medium storing a bitstream generated by the method according to Article 11.
13. A step of obtaining a bitstream for image information; wherein the bitstream is generated based on a step of deriving residual samples of a transform unit, a step of deriving transform coefficients of the transform unit based on a transform for the residual samples, and a step of encoding residual information about the transform coefficients, and A step of transmitting data including the above bitstream, A method wherein the above transformation is performed based on whether the transformation unit has an adaptive TU structure.