Image Encoding / Decoding Method and Apparatus, and Recording Medium Storing Bitstream
The method addresses inefficiencies in existing image compression technologies by adaptively determining MTS candidates based on block characteristics, reducing complexity and improving encoding efficiency for high-resolution images.
Patent Information
- Application Number
- JP2024576620
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-02
- Filing Date
- 2023-07-03
- Publication Date
- 2025-07-17
AI Technical Summary
Existing image compression technologies face challenges in efficiently adapting to the statistical characteristics of image blocks, leading to increased complexity and reduced encoding efficiency for high-resolution and high-quality images.
A method and apparatus that determine the number of available multi-transform type (MTS) candidates for a current block based on statistical characteristics such as the sum, number, magnitude, and position of conversion coefficients, allowing adaptive application of MTS to reduce transform coding complexity and improve encoding efficiency.
By adaptively controlling the range of available conversion type candidates, the method reduces transform coding complexity and enhances video encoding efficiency.
Smart Images

Figure 2025522787000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various application fields, and thus, highly efficient image compression technologies have been discussed.
[0003] As image compression technologies, there are various technologies such as an inter prediction technology that predicts pixel values included in a current picture from pictures before or after the current picture, an intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and an entropy coding technology that assigns short codes to values with high occurrence frequencies and long codes to values with low occurrence frequencies. Using such image compression technologies, image data can be effectively compressed and transmitted or stored.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present disclosure aims to provide a method and an apparatus for determining a range of available conversion type candidates for a current block.
[0005] The present disclosure aims to provide a method and an apparatus for adaptively applying MTS to a current block.
Means for Solving the Problems
[0006] The video decoding method and apparatus according to the present disclosure determine a conversion type for inverse transformation of the current block from one or more available MTS candidates of the current block, and perform inverse transformation on the conversion coefficients of the current block based on the determined conversion type to obtain residual samples of the current block. Here, the number of one or more available MTS candidates of the current block may be determined based on at least one of the sum of conversion coefficients in the current block, the number of conversion coefficients in the current block, the magnitude of one or more conversion coefficients in the current block, the position of the last valid coefficient in the current block, or the number of one or more non-zero coefficients in the current block.
[0007] In the video decoding method and apparatus according to the present disclosure, the number of one or more available MTS candidates of the current block may be determined based on a comparison between the sum of conversion coefficients belonging to the current block and a predetermined threshold.
[0008] In the video decoding method and apparatus according to the present disclosure, when the sum of conversion coefficients belonging to the current block is less than or equal to the first threshold, the number of the one or more MTS candidates is determined to be N1; when the sum of conversion coefficients belonging to the current block is greater than the first threshold and less than or equal to the second threshold, the number of the one or more MTS candidates is determined to be N2; and when the sum of conversion coefficients belonging to the current block is greater than the second threshold, the number of the one or more MTS candidates may be determined to be N3.
[0009] In the video decoding method and apparatus according to the present disclosure, the predetermined threshold is determined based on a predetermined constant factor, and the predetermined constant factor may be determined based on at least one of a slice type, a quantization parameter of the current block, a size of the current block, a form of the current block, a ratio of a width to a height of the current block, or information signaled by a bitstream.
[0010] In the video decoding method and apparatus according to the present disclosure, the number of one or more MTS candidates available for the current block may be determined based on at least one of the position of the last valid coefficient in the current block, the scaling factor, the size of the current block, or the shape of the current block.
[0011] In the video decoding method and apparatus according to the present disclosure, the number of one or more MTS candidates available for the current block may be determined based on a comparison between the number of non-zero coefficients among the transform coefficients of the current block and a third threshold.
[0012] In the video decoding method and apparatus according to the present disclosure, when the number of one or more MTS candidates available for the current block is two or more, the transform type of the current block may be determined based on the MTS index signaled by the bitstream.
[0013] In the video decoding method and apparatus according to the present disclosure, the maximum value (cMax) for the binary evolution of the MTS index may be determined based on at least one of the sum of the transform coefficients belonging to the current block, the number of transform coefficients in the current block, the magnitude of one or more transform coefficients in the current block, the position of the last valid coefficient in the current block, or the number of non-zero coefficients in the transform coefficients of the current block.
[0014] The video encoding method and apparatus according to the present disclosure may derive the transform coefficients of the current block by determining the transform type of the current block and performing a transform on the residual samples of the current block based on the transform type of the current block.
[0015] In the video encoding method and apparatus according to the present disclosure, the transform type of the current block may be selected from one or more MTS candidates belonging to any one of a plurality of predefined candidate groups.
[0016] In the video encoding method and apparatus according to the present disclosure, based on a candidate group in which the conversion type of the current block is selected, at least one of the sum of the conversion coefficients of the current block, the number of conversion coefficients in the current block, the magnitude of one or more conversion coefficients in the current block, the position of the last valid coefficient in the current block, or the number of non-zero coefficients in the current block may be required to belong to a predetermined range.
[0017] There is provided a computer-readable digital storage medium storing instructions or a program for causing a video encoding method to be performed by an encoding apparatus according to the present disclosure.
[0018] There is provided a computer-readable digital storage medium storing video / video information generated by a video encoding method according to the present disclosure.
[0019] There are provided a method and an apparatus for transmitting video / video information generated by a video encoding method according to the present disclosure.
Advantages of the Invention
[0020] According to the present disclosure, by considering the statistical characteristics of the conversion coefficients in the current block, the range of available conversion type candidates can be adaptively controlled, and the complexity of transform coding can be reduced.
[0021] Also, according to the present disclosure, by adaptively applying MTS based on the statistical characteristics of the conversion coefficients in the current block, the video encoding efficiency can be improved and the complexity of transform coding can be reduced.
Brief Description of the Drawings
[0022]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0023] The present disclosure can be modified in various ways and can have various embodiments. Specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood to include all modifications, equivalents, or alternatives included in the spirit and technical scope of the present disclosure. In the description of each figure, like reference numerals are used for like components.
[0024] Terms such as "first," "second," etc. may be used to describe various components, but these components should not be limited by these terms. These terms are only used for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes any combination of a plurality of related listed items or any one of the plurality of related listed items.
[0025] When a component is referred to as being "connected to" or "coupled to" another component, it should be understood that it may be directly connected to or directly coupled to the other component, or there may be still other components in between. On the other hand, when a component is referred to as being "directly connected to" or "directly coupled to" another component, it should be understood that there are no other components in between.
[0026] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. Singular expressions include plural expressions as well, unless otherwise specified in the context. In this application, terms such as "comprising" or "having" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0027] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein may be applied to the methods disclosed in the VVC (versatile video coding) standard. Also, the methods / embodiments disclosed in this specification may be applied to methods disclosed in video coding codecs that use multiple transforms (such as the EVC (essential video coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.)).
[0028] This specification presents various embodiments related to video / image coding, and unless otherwise specified, the above embodiments may be combined with each other.
[0029] In this specification, "video" can mean a collection of a series of "images" over time. "Picture" generally means a unit representing one image in a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (Coding Tree Units). One picture may be composed of one or more slices / tiles. One tile is a rectangular area composed of a plurality of CTUs in a specific tile column and a specific tile row of one picture. A tile column is a rectangular area of CTUs having the same height as the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs having a height specified by the picture parameter set and the same width as the width of the picture. The CTUs within one tile are continuously arranged by a CTU raster scan, while the tiles within one picture may be continuously arranged by a tile raster scan. One slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within the tiles of a picture that can be exclusively included in a single NAL unit. On the other hand, one picture may be divided into two or more sub-pictures. A sub-picture may be a rectangular area of one or more slices in a picture.
[0030] "Pixel", "pixel" or "pel" can mean the smallest unit that constitutes one picture (or image). Also, the term "sample" may be used as a term corresponding to a pixel. A sample can generally indicate a pixel or the value of a pixel, and may indicate only the pixel / pixel value of the luma component, or may indicate only the pixel / pixel value of the chroma component.
[0031] A "unit" can mean the basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. A unit may, in some cases, be used with the same meaning as terms such as "block" or "area". Generally, an MxN block may include a set (or, array) of samples (or, sample array) or transform coefficients consisting of M columns and N rows.
[0032] As used herein, "A or B" can mean "only A", "only B", or "both A and B". In other words, as used herein, "A or B" may be interpreted as "A and / or B". For example, as used herein, "A, B, or C" can mean "only A", "only B", "only C", or "any combination of A, B, and C".
[0033] The slash ( / ) or comma used herein can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0034] As used herein, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, as used herein, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0035] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0036] Also, the parentheses used in this specification can mean "for example". Specifically, when it is displayed as "prediction (intra prediction)", "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when it is displayed as "prediction (that is, intra prediction)", "intra prediction" may be proposed as an example of "prediction".
[0037] In this specification, the technical features separately described in the same drawing may be embodied separately or simultaneously.
[0038] FIG. 1 is a diagram showing a video / image coding system according to the present disclosure.
[0039] Referring to FIG. 1, the video / image coding system may include a first device (source device) and a second device (receiver device).
[0040] The source device can transmit encoded video / image information or data to the receiving device in the form of a file or a stream through a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0041] The video source can obtain video / images through processes such as capture, synthesis, or generation of video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images may be generated through a computer, etc., and in this case, the video / image capture process may be replaced during the process of generating related data.
[0042] The encoding device can encode the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0043] The transmission unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device through a digital storage medium or a network in the form of a file or a stream. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit may include elements for generating a media file according to a predefined file format and may include elements for transmission through a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0044] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0045] The renderer can render the decoded video / image. The rendered video / image may be displayed from the display unit.
[0046] FIG. 2 is a schematic block diagram of an encoding device to which the application of the embodiment of the present disclosure is applicable and in which video / image signal encoding is performed.
[0047] Referring to FIG. 2, the encoding device 200 may include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer (232), a quantizer 233, a dequantizer (234), and an inverse transformer (235). The residual processor 230 may further include a subtractor (231). The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be constituted by one or more hardware components (e.g., an encoding device chipset or a processor) according to an embodiment. Also, the memory 270 may include a DPB (decoded picture buffer) and may be constituted by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.
[0048] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0049] As an example, one coding unit may be divided into a plurality of coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary structure may be applied later. Or, the binary-tree structure may be applied earlier than the quad-tree structure. The coding procedure according to this specification may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units with a lower depth, and a coding unit having an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later.
[0050] As another example, the processing unit may further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit may be respectively divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit of deriving transform coefficients and / or a unit of deriving a residual signal from the transform coefficients.
[0051] The unit may be used in the same sense as terms such as a block or an area in some cases. Generally, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luma component, or may represent only the pixel / pixel value of the chroma component. A sample may be used as a term corresponding to one picture (or image), pixel, or pel.
[0052] The encoding device 200 can subtract a prediction signal (prediction block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoding device 200 may be called a subtraction unit 231.
[0053] The prediction unit 220 can perform a prediction on a processing target block (hereinafter referred to as the current block), and generate a predicted block including a prediction sample for the current block. The prediction unit 220 can determine whether intra prediction or inter prediction is applied in units of the current block or CU. As will be described later in the description of each prediction mode, the prediction unit 220 can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240. The information related to prediction may be encoded by the entropy encoding unit 240 and output in the form of a bit stream.
[0054] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the vicinity (neighbor) of the current block depending on the prediction mode, or may be located at a certain distance from the current block. In intra prediction, the prediction mode may include one or more non-directional modes and a plurality of directional modes. The non-directional mode may include at least one of the DC mode or the Planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail of the prediction direction. However, this is an example, and a greater or smaller number of directional modes may be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0055] The inter prediction unit 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the peripheral block may include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called by names such as a collocated reference block and a collocated CU (colCU), and the reference picture including the temporal neighboring block may also be called a collocated picture (colPic). For example, the inter prediction unit 221 can configure a motion information candidate list based on the peripheral block and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal does not need to be transmitted.In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and by signaling the motion vector difference, the motion vector of the current block can be indicated.
[0056] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This may be called the combined inter and intra prediction (CIIP) mode. Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode may be used for coding content images / videos such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction methods described in this specification. The palette mode may be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index. The prediction signal generated by the prediction unit 220 may be used to generate a restored signal or may be used to generate a residual signal.
[0057] The conversion unit 232 can generate transform coefficients by applying a conversion method to the residual signal. For example, the conversion method may include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the transform obtained from the graph when expressing the relationship information between pixels as a graph. CNT means the transform obtained based on generating a prediction signal using all previously restored pixels. Also, the conversion process may be applied to a pixel block having the same size of a square, or may also be applied to a block of variable size that is not square.
[0058] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240, and the entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients may be called residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0059] The entropy encoding unit 240 can perform various encoding methods such as exponential Golomb, CAVLC (context - adaptive variable length coding), CABAC (context - adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can encode, together or separately, not only the quantized transform coefficients but also the information necessary for video / image restoration (e.g., the values of syntax elements, etc.).
[0060] The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information may further include general constraint information. In this specification, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information is encoded by the above - described encoding procedure and may be included in the bitstream. The bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu - ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 may be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit may be included in the entropy encoding unit 240.
[0061] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients using the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block may be used as the reconstructed block. The addition unit 250 may be referred to as a restoration unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next block to be processed within the current picture and, as will be described later, may be used for inter prediction of the next picture after passing through filtering. On the other hand, LMCS (luma mapping with chroma scaling) may be applied during picture encoding and / or the restoration process.
[0062] The filtering unit 260 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240. The information related to filtering may be encoded by the entropy encoding unit 240 and output in the form of a bit stream.
[0063] The modified restored picture transmitted to the memory 270 may be used as a reference picture in the inter prediction unit 221. The encoding device can thereby avoid prediction mismatches in the encoding device 200 and the decoding device when inter prediction is applied, and can also improve the coding efficiency.
[0064] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded), and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatial neighboring blocks or temporal neighboring blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.
[0065] FIG. 3 is a schematic block diagram of a decoding apparatus to which an embodiment of the present disclosure is applicable and in which decoding of a video / image signal is performed.
[0066] Referring to FIG. 3, the decoding apparatus 300 may include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filtering unit (filter, 350), and a memory (memoery, 360). The predictor 330 may include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 may include a dequantizer (321) and an inverse transformer (inverse transformer, 321).
[0067] The entropy decoding unit 310, the residual processing unit 320, the prediction unit 330, the addition unit 340, and the filtering unit 350 described above may be configured by one hardware component (for example, a decoding apparatus chipset or a processor) according to an embodiment. Further, the memory 360 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0068] When a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information was processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device 300 may be reproduced by a reproducing device.
[0069] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal may be decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream and derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information may further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signals / received information and / or syntax elements described later in this specification may be decoded by the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding and decoding target blocks, or the information of the symbol / bin decoded in the previous stage, predicts the occurrence probability of the bin using the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value obtained by performing entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, may be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, the information related to filtering among the information decoded by the entropy decoding unit 310 may be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310.
[0070] On the other hand, the decoding device according to the present specification may be referred to as a video / image / picture decoding device, and the decoding device may be classified into an information decoding device (video / image / picture information decoding device) and a sample decoding device (video / image / picture sample decoding device). The information decoding device may include the entropy decoding unit 310, and the sample decoding device may include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0071] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output the transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain the transform coefficients.
[0072] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).
[0073] The prediction unit 320 can perform prediction on the current block to generate a predicted block including predicted samples for the current block. The prediction unit 320 can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0074] The prediction unit 320 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit 320 can apply intra prediction or inter prediction for the prediction of one block, and can also apply intra prediction and inter prediction simultaneously. This may be referred to as the CIIP (combined inter and intra prediction) mode. Further, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of the block. The IBC prediction mode or the palette mode may be used for content image / video coding such as games like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction methods described in this specification. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index may be included in and signaled in the video / image information.
[0075] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located around (neighbor) the current block depending on the prediction mode, or may be located at a certain distance from the current block. In intra prediction, the prediction mode may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit 331 can determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0076] The inter prediction unit 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the peripheral block may include a spatial neighboring block existing within the current picture and a temporal neighboring block existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on the peripheral block, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information regarding the prediction may include information indicating the inter prediction mode for the current block.
[0077] The addition unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the processing target block as in the case where the skip mode is applied, the prediction block may be used as the restored block.
[0078] The addition unit 340 may be referred to as a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, and as will be described later, may be output after filtering, or may be used for inter prediction of the next picture. On the other hand, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0079] The filtering unit 350 can apply filtering to the restoration signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.
[0080] The (modified) restored picture stored in the DPB of the memory 360 may be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the blocks for which the motion information within the current picture has been derived (or decoded), and / or the motion information of the blocks within the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 360 can store the restored samples of the restored blocks within the current picture and transmit them to the intra prediction unit 331.
[0081] In this specification, the examples described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 200 may be applied to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300 so as to be the same or corresponding, respectively.
[0082] FIG. 4 is a diagram showing an inverse transformation method performed by a decoding device according to an embodiment of the present disclosure.
[0083] The present disclosure relates to an MTS (multi transform selection)-based inverse transformation method. The MTS can mean a method of selecting any one of a plurality of transform type candidates to determine a transform type for horizontal / vertical transformation. Hereinafter, a plurality of transform type candidates for the MTS are referred to as MTS candidates.
[0084] Each of the MTS candidates may include a transform type for horizontal transformation and a transform type for vertical transformation. Each of the MTS candidates may be configured by a combination of two of the transform types already defined identically in the encoding device and the decoding device. Here, the already defined transform type may include at least one of one or more DCT-based first transform types or one or more DST-based second transform types. As an example, the first transform type candidate may include at least one of DCT-2, DCT-3, DCT-4, DCT-5, or DCT-8. The second transform type candidate may include at least one of DST-7, DST-1, or DST-4.
[0085] Alternatively, the MTS candidates may be determined separately by being classified into MTS candidates for horizontal transformation and MTS candidates for vertical transformation. In this case, each of the MTS candidates may include at least two of the transform types already defined identically in the encoding device and the decoding device. The MTS candidates for horizontal transformation may be used identically as the MTS candidates for vertical transformation.
[0086] Referring to FIG. 4, the conversion type of the current block can be determined from one or more available MTS candidates of the current block (S400).
[0087] Based on any one of the one or more MTS candidates selected from the one or more MTS candidates, the conversion types for the horizontal and / or vertical conversion of the current block may be determined respectively.
[0088] The number of one or more available MTS candidates of the current block may be determined based on at least one of the sum of the conversion coefficients in the current block, the number of conversion coefficients in the current block, the magnitude of one or more conversion coefficients in the current block, the position of the last valid coefficient, or the number of one or more non-zero coefficients in the current block. The magnitude of one or more conversion coefficients in the current block can mean the magnitude of a conversion coefficient greater than a specific threshold magnitude. The magnitude of one or more conversion coefficients in the current block can mean the magnitude of a conversion coefficient belonging to a specific position in the current block. Here, the specific position may include the upper left corner position in the current block and / or at least one sample position adjacent to the upper left corner position.
[0089] When the number of available MTS candidates for the current block is 1, this may mean that MTS is not applied to the current block. In such a case, the MTS candidate may include DCT-2 and DCT-2 as the conversion types for horizontal and vertical conversion respectively. On the other hand, when the number of available MTS candidates for the current block is 2 or more, this may mean that MTS is applied to the current block. Therefore, based on the method for determining the number of available MTS candidates for the current block described later, it may be determined whether MTS is applied to the current block. In this case, the step of determining the conversion type of the current block may include the step of determining whether MTS is applied to the current block.
[0090] The method for determining the number of one or more MTS candidates available to the current block will be described in detail below.
[0091] Method 1: Sum-based method of conversion coefficients
[0092] The number of one or more MTS candidates available to the current block may be determined based on whether the sum of the conversion coefficients belonging to the current block is greater than a predetermined threshold. Here, the sum of the conversion coefficients can mean the sum of the absolute values of the conversion coefficients.
[0093] As an example, when the sum of the conversion coefficients belonging to the current block is less than or equal to the threshold, the number of one or more MTS candidates may be N1, and when the sum of the conversion coefficients belonging to the current block is greater than the threshold, the number of one or more MTS candidates may be N2. Here, the threshold may be determined based on a predetermined constant factor (C). For example, the threshold may be defined as 6*C. Also, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 is an integer greater than N1 and may be 2, 3, 5, 6 or more.
[0094] Alternatively, when the sum of the conversion coefficients belonging to the current block is less than or equal to the threshold, the number of one or more MTS candidates may be N1, and when the sum of the conversion coefficients belonging to the current block is greater than the threshold, the number of one or more MTS candidates may be N3. Here, the threshold may be determined based on a predetermined constant factor (C). For example, the threshold may be defined as 32*C. Also, N1 may be 1 and N3 may be 6. However, it is not limited thereto, and N3 is an integer greater than N1 and may be 2, 3, 4, 5 or more.
[0095] Alternatively, when the sum of the transform coefficients belonging to the current block is less than or equal to a threshold, the number of one or more MTS candidates may be N2, and when the sum of the transform coefficients belonging to the current block is greater than the threshold, the number of one or more MTS candidates may be N3. Here, the threshold may be determined based on a predetermined constant factor (C). For example, the threshold may be defined as 32*C. Also, N2 may be 4 and N3 may be 6. However, it is not limited thereto, N2 is an integer smaller than N3, and may be 2, 3, 5 or more, and N3 is an integer greater than N2, and may be 3, 4, 5, 7 or more.
[0096] Alternatively, when the sum of the transform coefficients belonging to the current block is less than or equal to a first threshold, the number of one or more MTS candidates may be N1. When the sum of the transform coefficients belonging to the current block is greater than the first threshold and less than or equal to a second threshold, the number of one or more MTS candidates may be N2. When the sum of the transform coefficients belonging to the current block is greater than the second threshold, the number of one or more MTS candidates may be N3. Here, the first threshold and the second threshold may be determined respectively based on a predetermined constant factor (C). For example, the first threshold and the second threshold may be defined as 16*C and 32*C respectively. Also, N1, N2 and N3 may be 1, 4 and 6 respectively. However, it is not limited thereto, N2 is an integer greater than N1 and smaller than N3, and may be 2, 3, 5, or more. N3 is an integer greater than N2, and may be 3, 4, 5, 7 or more.
[0097] Information for determining at least one of the aforementioned threshold value, first threshold value, second threshold value, or constant factor (C) may be signaled by a bitstream. The constant factor (C) may be a constant value that has already been defined identically in the encoding device and the decoding device. Alternatively, the constant factor (C) may be variably determined based on at least one of the slice type, quantization parameter (QP) of the current block, size of the current block, form of the current block (i.e., whether it is non-square or not), ratio of the width and height of the current block, or information signaled by the bitstream. Here, the information to be signaled may be information encoded for determining at least one of the aforementioned threshold value, first threshold value, second threshold value, or constant factor (C).
[0098] As an example, when the quantization parameter of the current block is smaller than a first value, the constant factor may be determined to be 2, and when the quantization parameter of the current block is greater than or equal to the first value, the constant factor may be determined to be 1. Here, the first value may be 22.
[0099] Alternatively, when the quantization parameter of the current block is smaller than a first value, the constant factor may be determined to be 2, and when the quantization parameter of the current block is greater than or equal to the first value, the constant factor may be determined to be 0.5. Here, the first value may be 22.
[0100] Alternatively, when the quantization parameter of the current block is smaller than a second value, the constant factor may be determined to be 1, and when the quantization parameter of the current block is greater than or equal to the second value, the constant factor may be determined to be 0.5. Here, the second value may be 37.
[0101] Alternatively, when the quantization parameter of the current block is smaller than the first value, the constant factor may be determined as 2. When the quantization parameter of the current block is larger than or equal to the first value and smaller than the second value, the constant factor may be determined as 1. When the quantization parameter of the current block is larger than or equal to the second value, the constant factor may be determined as 0.5. Here, the first value and the second value may be 22 and 37 respectively.
[0102] As an example, when the size of the current block is smaller than the first size, the constant factor may be determined as 1. When the size of the current block is larger than or equal to the first size, the constant factor may be determined as 2. Here, the first size may be 128. However, it is not limited thereto, and the first size may be 16, 32 or 64.
[0103] Alternatively, when the size of the current block is smaller than the first size, the constant factor may be determined as 1. When the size of the current block is larger than or equal to the first size, the constant factor may be determined as 3. Here, the first size may be 128. However, it is not limited thereto, and the first size may be 16, 32 or 64.
[0104] Alternatively, when the size of the current block is smaller than the second size, the constant factor may be determined as 2. When the size of the current block is larger than or equal to the second size, the constant factor may be determined as 3. Here, the second size may be 512. However, it is not limited thereto, and the second size is larger than the first size and may be 32, 64, 128, or 256.
[0105] Alternatively, when the size of the current block is smaller than the first size, the constant factor may be determined to be 1. When the size of the current block is larger than or equal to the first size and smaller than the second size, the constant factor may be determined to be 2. When the size of the current block is larger than or equal to the second size, the constant factor may be determined to be 3. Here, the first size and the second size may be 128 and 512 respectively. However, it is not limited thereto, the first size may be 16, 32 or 64, and the second size may be larger than the first size and may be 32, 64, 128, or 256.
[0106] As an example, when the ratio of the width to the height of the current block is smaller than the first value, the constant factor may be determined to be 1. When the ratio of the width to the height of the current block is larger than or equal to the first value and smaller than the second value, the constant factor may be determined to be 2. When the ratio of the width to the height of the current block is larger than or equal to the second value, the constant factor may be determined to be 3.
[0107] The size of the current block described above may mean the size of the transform block having the same size as the current block to be decoded, and may also mean the size of the remaining region other than the high-frequency region where all the transform coefficients are 0 within the transform block, that is, the size of the low-frequency region. Here, the size is defined as the total number of samples belonging to the current block (or the low-frequency region), but is not limited thereto. For example, the size may be expressed as the width, the height, the minimum / maximum value of the width and the height, or the sum of the width and the height.
[0108] Instead of the sum of the transform coefficients in the current block, the number and / or size of the transform coefficients may be used to determine the number of MTS candidates. For this purpose, in the above-described method 1, the term "sum of transform coefficients" may be applied in place of "number of transform coefficients" or "size of one or more transform coefficients". Alternatively, at least two of the sum, number, or size of the transform coefficients may be used together to determine the number of MTS candidates.
[0109] Method 2: Position-based method of the last valid coefficient
[0110] The number of one or more MTS candidates available for the current block may be determined based on the position of the last valid coefficient in the current block. Here, the position of the last valid coefficient may be defined as the scan order (or scan position) by a predetermined scan method. The scan method may be any one of diagonal scan, horizontal scan, vertical scan, z scan, or raster scan. Alternatively, the position of the last valid coefficient may be defined as the coordinates of the last valid coefficient with reference to the top-left sample of the current block.
[0111] Specifically, the number of one or more MTS candidates available for the current block may be determined based on at least one of the position of the last valid coefficient in the current block, the size of the current block, the shape of the current block (i.e., whether it is non-square), and the ratio of the width to the height of the current block. As an example, the number of MTS candidates may be expressed as in Equation 1 below.
[0112] [Equation 1] NumMtsCand = min(MaxNumMtsCand,(LastScanPos*Scale / cbSize))
[0113] In Equation 1, NumMtsCand represents the number of one or more MTS candidates available for the current block. MaxNumMtsCand represents the maximum number of MTS candidates already defined in the encoding device and the decoding device. LastScanPos represents the position of the last valid coefficient in the current block, and Scale is the scaling factor applied to LastScanPos. The scaling factor is a value already defined identically in the encoding device and the decoding device and may be 1, 2, 3, 4, 5, 6, or more. cbSize represents the size of the current block, which is as described in "Method 1 Based on the Sum of Transformation Coefficients".
[0114] According to Equation 1, the number of one or more MTS candidates available for the current block may be set to (LastScanPos * Scale / cbSize). However, when the value of (LastScanPos * Scale / cbSize) exceeds the maximum number of predefined MTS candidates, the number of one or more MTS candidates available for the current block may be limited to the maximum number of predefined MTS candidates.
[0115] Alternatively, the number of one or more MTS candidates available for the current block may be determined based on whether the last valid coefficient exists in the upper left region within the current block (or the low-frequency region).
[0116] As an example, assume that the width and height of the current block (or the low-frequency region) are W1 and H1 respectively, and the width and height of the upper left region are W2 and H2 respectively. W2 may be smaller than or equal to W1, and H2 may be smaller than or equal to H1. In this case, when the last valid coefficient does not exist in the upper left region of the current block (i.e., when the last valid coefficient exists in another region within the current block that is not the upper left region), the number of one or more MTS candidates may be N1. On the other hand, when the last valid coefficient exists in the upper left region of the current block, the number of one or more MTS candidates may be N2. However, even when the last valid coefficient exists in the upper left region of the current block, when the last valid coefficient exists at the upper left sample position of the current block (or the low-frequency region), i.e., the DC position, the number of one or more MTS candidates may be N1. Here, N1 may be 1 and N2 may be 4. However, this is not limited thereto, and N2 is an integer greater than N1 and may be 2, 3, 5, 6 or more.
[0117] Method 3: Number-based method of non-zero coefficients
[0118] The number of one or more MTS candidates available for the current block may be determined based on whether the number of non-zero coefficients among the transform coefficients of the current block is greater than a predetermined threshold.
[0119] For this purpose, at least one of the minimum number of non-zero coefficients (MinNumNonzero) for MTS to be allowed / applied to the current block or the maximum number of non-zero coefficients (MaxNumNonzero) for MTS to be allowed / applied to the current block may be defined. MinNumNonzero and MaxNumNonzero may be the same numbers that have already been defined in the encoding device and the decoding device, respectively. Alternatively, MinNumNonzero may be a fixed number (e.g., 2) that has already been defined regardless of the size of the current block, and MaxNumNonzero may be a variable number determined based on the size of the current block. On the other hand, both MinNumNonzero and MaxNumNonzero may be considered to determine the number of one or more MTS candidates, or only one of MinNumNonzero or MaxNumNonzero may be considered.
[0120] As an example, when the number of non-zero coefficients belonging to the current block is less than MinNumNonzero, the number of one or more MTS candidates may be N1, and when the number of non-zero coefficients belonging to the current block is greater than or equal to MinNumNonzero, the number of one or more MTS candidates may be N2. Here, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 may be an integer greater than N1 and may be 2, 3, 5, 6, or more.
[0121] Alternatively, when the number of non-zero coefficients belonging to the current block is greater than MaxNumNonzero, the number of one or more MTS candidates may be N1, and when the number of non-zero coefficients belonging to the current block is less than or equal to MaxNumNonzero, the number of one or more MTS candidates may be N2. Here, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 is an integer greater than N1 and may be 2, 3, 5, 6, or more.
[0122] Alternatively, when the number of non-zero coefficients belonging to the current block is less than MinNumNonzero, the number of one or more MTS candidates may be N1. When the number of non-zero coefficients belonging to the current block is greater than or equal to MinNumNonzero and less than or equal to MaxNumNonzero, the number of one or more MTS candidates may be N2. When the number of non-zero coefficients belonging to the current block is greater than MaxNumNonzero, the number of one or more MTS candidates may be N1. Here, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 is an integer greater than N1 and may be 2, 3, 5, 6, or more.
[0123] The above-described Methods 1 to 3 are independent embodiments, and the number of one or more MTS candidates available for the current block can be determined using any one of Methods 1 to 3. Alternatively, the number of one or more MTS candidates available for the current block can also be determined based on a combination of at least two of the above-described Methods 1 to 3.
[0124] Hereinafter, a method for selecting the conversion type of the current block from one or more MTS candidates available for the current block will be described in detail.
[0125] When the number of one or more MTS candidates available for the current block is two or more, the conversion type of the current block may be determined based on the MTS index. Here, the MTS index can identify any one of the plurality of MTS candidates. Based on the MTS candidate identified by the MTS index, the conversion type for horizontal conversion and the conversion type for vertical conversion may be determined.
[0126] Alternatively, the MTS index may be defined separately for each conversion direction. In this case, the MTS index may include an MTS index for horizontal conversion and an MTS index for vertical conversion. In this case, the conversion type for horizontal conversion of the current block may be determined based on the MTS index for horizontal conversion, and the conversion type for vertical conversion of the current block may be determined based on the MTS index for vertical conversion.
[0127] The MTS index may be signaled by a bitstream. The MTS index may be encoded by context-based adaptive binary arithmetic coding (CABAC).
[0128] The maximum value (cMax) for the binary evolution of the MTS index may be variably determined based on the number of one or more MTS candidates determined based on at least one of the above-described methods 1 to 3. That is, the maximum value (cMax) for the binary evolution of the MTS index may be determined based on at least one of the sum of the conversion coefficients belonging to the current block, the number of conversion coefficients in the current block, the magnitude of one or more conversion coefficients in the current block, the position of the last valid coefficient in the current block, or the number of non-zero coefficients in the conversion coefficients of the current block. As an example, the maximum value (cMax) for the binary evolution of the MTS index may be set to a value obtained by subtracting 1 from the number of one or more MTS candidates determined based on at least one of the above-described methods 1 to 3.
[0129] The MTS index may be derived based on the encoding parameters of the current block. Here, the encoding parameters may include at least one of a prediction mode indicating an intra mode or an inter mode, the width and / or height of the current block (or, transform block), or the split information of the current block (or, transform block). The split information may include at least one of information indicating the presence or absence of splitting, information indicating the splitting direction, information regarding the number of partitions generated by block splitting, information regarding the position of the partition to which inverse transformation is applied among the partitions generated by block splitting, or information indicating the presence or absence of asymmetric splitting.
[0130] On the other hand, when the number of one or more MTS candidates available for the current block is 1, the conversion type for the horizontal conversion of the current block and the conversion type for the vertical conversion may be determined based on the MTS candidate. In this case, the MTS index may not be signaled by the bit stream, and the MTS index may be derived with a value that is already defined identically in the encoding device and the decoding device.
[0131] Referring to FIG. 4, an inverse transformation can be performed on the transform coefficients of the current block based on the conversion type of the current block, and the residual samples of the current block can be obtained (S410).
[0132] The inverse transformation may be performed based on a separable transformation. As an example, a vertical transformation can be performed on each column of the transform coefficients of the current block, and a horizontal transformation can be performed on each row of the resulting values. Alternatively, a horizontal transformation can be performed on each row of the transform coefficients of the current block, and a vertical transformation can be performed on each column of the resulting values.
[0133] The current block can be restored based on the residual samples obtained by the inverse transformation and the prediction samples of the current block.
[0134] FIG. 5 is a diagram showing a schematic configuration of an inverse transformation unit 322 that performs the inverse transformation method according to the present disclosure.
[0135] With reference to FIG. 4, the inverse conversion method performed by the decoding apparatus has been described, which may be similarly performed in the inverse conversion unit 322 of the decoding apparatus, and the following detailed description is omitted.
[0136] Referring to FIG. 5, the inverse conversion unit 322 may include a conversion type determination unit 500 and a residual sample acquisition unit 510.
[0137] The conversion type determination unit 500 can determine the conversion type of the current block from one or more available MTS candidates of the current block. That is, the conversion type determination unit 500 can select any one of the one or more MTS candidates, and based on the selected MTS candidate, can determine the conversion types for the horizontal and / or vertical conversion of the current block, respectively.
[0138] The conversion type determination unit 500 can determine the number of one or more available MTS candidates of the current block based on at least one of the sum of the conversion coefficients in the current block, the number of conversion coefficients in the current block, the magnitude of one or more conversion coefficients in the current block, the position of the last valid coefficient, or the number of one or more non-zero coefficients in the current block. For this purpose, at least one of Methods 1 to 3 described with reference to FIG. 4 may be used.
[0139] On the other hand, based on the method of determining the number of available MTS candidates of the current block, it may be determined whether MTS is applied to the current block. In this case, the conversion type determination unit 500 can also determine whether MTS is applied to the current block.
[0140] When the number of one or more MTS candidates available for the current block is two or more, the conversion type determination unit 500 can determine the conversion type of the current block based on a predetermined MTS index. The MTS index may be signaled by a bit stream or may be derived based on the encoding parameters of the current block.
[0141] On the other hand, when the number of one or more MTS candidates available for the current block is one, the conversion type determination unit 500 can determine the conversion type for the horizontal conversion and the conversion type for the vertical conversion of the current block based on the MTS candidate.
[0142] The residual sample acquisition unit 510 can perform inverse conversion on the conversion coefficients of the current block based on the conversion type of the current block and acquire the residual samples of the current block.
[0143] FIG. 6 is a diagram showing a conversion method performed by an encoding apparatus according to an embodiment of the present disclosure.
[0144] The present disclosure relates to an MTS-based conversion method. As described above, MTS can mean a method of selecting any one from a plurality of MTS candidates and determining a conversion type for horizontal / vertical conversion. Here, each of the MTS candidates may include a conversion type for horizontal conversion and a conversion type for vertical conversion. Alternatively, the MTS candidates may be defined separately as MTS candidates for horizontal conversion and MTS candidates for vertical conversion.
[0145] Referring to FIG. 6, the conversion type of the current block can be determined (S600).
[0146] The conversion type for the current block's horizontal and / or vertical conversion may be determined respectively based on any one of the MTS candidates already defined in the encoding device and the decoding device. As described with reference to FIG. 4, each of the already defined MTS candidates may be composed of a combination of two of the conversion types already defined identically in the encoding device and the decoding device. Or, the MTS candidates may be determined separately by classifying them into MTS candidates for horizontal conversion and MTS candidates for vertical conversion. Or, the MTS candidates for horizontal conversion may be used identically as the MTS candidates for vertical conversion.
[0147] Referring to FIG. 6, by performing a conversion on the residual samples of the current block based on the conversion type of the current block, the conversion coefficients of the current block can be derived (S610).
[0148] The conversion may be performed by the reverse process of the inverse conversion described above. As an example, a horizontal conversion can be performed on each row of the residual samples of the current block, and a vertical conversion can be performed on each column of the resulting values. Or, a vertical conversion can be performed on each column of the residual samples of the current block, and a horizontal conversion can be performed on each row of the resulting values.
[0149] The conversion coefficients induced by the conversion may be encoded to generate the residual information of the current block, and the generated residual information may be inserted into the bitstream.
[0150] The conversion type of the current block may be any one selected from one or more MTS candidates belonging to a predetermined candidate group. The predetermined candidate group may be any one of a plurality of candidate groups that have already been identically defined in the encoding device and the decoding device. The plurality of candidate groups may include different numbers of MTS candidates. As an example, the plurality of candidate groups may include at least two of a first candidate group, a second candidate group, or a third candidate group. Here, the first candidate group may include N1 MTS candidates, the second candidate group may include N2 MTS candidates, and the third candidate group may include N3 MTS candidates.
[0151] Based on the candidate group in which the conversion type of the current block is selected (or the number of MTS candidates belonging to the candidate group), at least one of the sum of the conversion coefficients of the current block, the number of conversion coefficients in the current block, the magnitude of one or more conversion coefficients in the current block, the position of the last valid coefficient, or the number of non-zero coefficients in the current block may be required / limited to belong to a predetermined range.
[0152] Method 1: Requirements for the sum of conversion coefficients
[0153] Depending on the candidate group in which the conversion type of the current block is selected, it may be required that the sum of the conversion coefficients of the current block belongs to a range that is less than or equal to a predetermined threshold, or it may be required that the sum of the conversion coefficients belongs to a range that is greater than the predetermined threshold. Here, the sum of the conversion coefficients can mean the sum of the absolute values of the conversion coefficients.
[0154] As an example, when the conversion type of the current block is selected from a first candidate group including N1 MTS candidates, the sum of the conversion coefficients belonging to the current block may be less than or equal to the threshold value, or may be greater than the threshold value. On the other hand, when the conversion type of the current block is selected from a second candidate group including N2 MTS candidates, the sum of the conversion coefficients belonging to the current block may be required to be greater than the threshold value. Here, the threshold value may be determined based on a predetermined constant factor (C). For example, the threshold value may be defined as 6*C. Also, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 is an integer greater than N1, and may be 2, 3, 5, 6 or more.
[0155] Alternatively, when the conversion type of the current block is selected from a first candidate group including N1 MTS candidates, the sum of the conversion coefficients belonging to the current block may be less than or equal to the threshold value, or may be greater than the threshold value. On the other hand, when the conversion type of the current block is selected from a third candidate group including N3 MTS candidates, the sum of the conversion coefficients belonging to the current block may be required to be greater than the threshold value. Here, the threshold value may be determined based on a predetermined constant factor (C). For example, the threshold value may be defined as 32*C. Also, N1 may be 1 and N3 may be 6. However, it is not limited thereto, and N3 is an integer greater than N1, and may be 2, 3, 4, 5 or more.
[0156] Alternatively, when the transformation type of the current block is selected from a second candidate group including N2 MTS candidates, the sum of the transformation coefficients belonging to the current block may be required to be less than or equal to a threshold value. On the other hand, when the transformation type of the current block is selected from a third candidate group including N3 MTS candidates, the sum of the transformation coefficients belonging to the current block may be required to be greater than the threshold value. Here, the threshold value may be determined based on a predetermined constant factor (C). For example, the threshold value may be defined as 32*C. Also, N2 may be 4, and N3 may be 6. However, it is not limited thereto, N2 is an integer smaller than N3, and may be 2, 3, 5, or more, and N3 is an integer greater than N2, and may be 3, 4, 5, 7, or more.
[0157] Alternatively, when the transformation type of the current block is selected from a first candidate group including N1 MTS candidates, the sum of the transformation coefficients belonging to the current block may be less than or equal to a first threshold value, or may be greater than the first threshold value. When the transformation type of the current block is selected from a second candidate group including N2 MTS candidates, the sum of the transformation coefficients belonging to the current block may be required to be greater than the first threshold value and less than or equal to a second threshold value. When the transformation type of the current block is selected from a third candidate group including N3 MTS candidates, the sum of the transformation coefficients belonging to the current block may be required to be greater than the second threshold value. Here, the first threshold value and the second threshold value may be determined based on a predetermined constant factor (C), respectively. For example, the first threshold value and the second threshold value may be defined as 16*C and 32*C, respectively. Also, N1, N2, and N3 may be 1, 4, and 6, respectively. However, it is not limited thereto, N2 is an integer greater than N1 and smaller than N3, and may be 2, 3, 5, or more. N3 is an integer greater than N2, and may be 3, 4, 5, 7, or more.
[0158] Information for determining at least one of the aforementioned threshold value, first threshold value, second threshold value, or constant factor (C) may be encoded and inserted into the bitstream. The aforementioned constant factor (C) may be a constant value that has already been defined identically in the encoding device and the decoding device. Alternatively, the constant factor (C) may be variably determined based on at least one of the slice type, quantization parameter (QP) of the current block, size of the current block, form of the current block (i.e., whether it is non-square), or ratio of the width to the height of the current block. In this case, information for determining the constant factor (C) may be further encoded and inserted into the bitstream. This is as described in detail with reference to FIG. 4.
[0159] Instead of the sum of the transform coefficients in the current block, the aforementioned requirements may be applied to the number and / or magnitude of the transform coefficients. For this purpose, in the method 1, the term "sum of transform coefficients" may be applied in place of "number of transform coefficients" or "magnitude of one or more transform coefficients". Alternatively, the aforementioned requirements may be applied together to at least two of the sum, number, or magnitude of the transform coefficients.
[0160] Method 2: Requirements for the position of the last valid coefficient
[0161] Depending on the candidate group in which the transform type of the current block is selected, it may be required that the last valid coefficient in the current block is located in a predetermined region within the current block or exists at a predetermined position within the current block. Here, the position of the last valid coefficient may be defined by the scan order (or scan position) according to a predetermined scan method. The scan method may be any one of diagonal scan, horizontal scan, vertical scan, z scan, or raster scan. Alternatively, the position of the last valid coefficient may be defined by the coordinates of the last valid coefficient with reference to the top-left sample of the current block.
[0162] As an example, when the conversion type of the current block is selected from a first candidate group including N1 MTS candidates, the last valid coefficient may be required to be at a position such that NumMtsCand in the above-described formula 1 has a value of N1. Similarly, when the conversion type of the current block is selected from a second candidate group including N2 MTS candidates, the last valid coefficient may be required to be at a position such that NumMtsCand has a value of N2. When the conversion type of the current block is selected from a third candidate group including N3 MTS candidates, the last valid coefficient may be required to be at a position such that NumMtsCand has a value of N3.
[0163] Alternatively, it may be required that, depending on the candidate group from which the conversion type of the current block is selected, the last valid coefficient exists in the upper left region within the current block (or the low-frequency region).
[0164] As an example, assume that the width and height of the current block (or the low-frequency region) are W1 and H1, respectively, and the width and height of the upper left region are W2 and H2, respectively. W2 may be smaller than or equal to W1, and H2 may be smaller than or equal to H1. When the conversion type of the current block is selected from a first candidate group including N1 MTS candidates, the last valid coefficient may exist in the upper left region of the current block or in another region within the current block that is not the upper left region. On the other hand, when the conversion type of the current block is selected from a first candidate group including N1 MTS candidates, the last valid coefficient may be required to exist in the upper left region of the current block. Note that even when the last valid coefficient exists in the upper left region of the current block, it may be required that the last valid coefficient does not exist at the upper left sample position of the current block (or the low-frequency region), that is, the DC position. Here, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 may be an integer greater than N1 and may be 2, 3, 5, 6, or more.
[0165] Method 3: Requirements for the number of non-zero coefficients
[0166] Depending on the candidate group in which the conversion type of the current block is selected, the number of non-zero coefficients in the current block is required to belong to a range that is less than or equal to a predetermined threshold, or may be required to belong to a range greater than the predetermined threshold. Here, the threshold may be defined as at least one of the minimum number of non-zero coefficients (MinNumNonzero) for MTS to be allowed / applied to the current block or the maximum number of non-zero coefficients (MaxNumNonzero) for MTS to be allowed / applied to the current block, as described with reference to FIG. 4.
[0167] As an example, when the conversion type of the current block is selected from a first candidate group including N1 MTS candidates, the number of non-zero coefficients belonging to the current block may be less than MinNumNonzero, or may be greater than or equal to MinNumNonzero. On the other hand, when the conversion type of the current block is selected from a second candidate group including N2 MTS candidates, the number of non-zero coefficients belonging to the current block may be required to be greater than or equal to MinNumNonzero. Here, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 is an integer greater than N1 and may be 2, 3, 5, 6 or more.
[0168] Alternatively, when the conversion type of the current block is selected from a first candidate group including N1 MTS candidates, the number of non-zero coefficients belonging to the current block may be greater than MaxNumNonzero, or may be less than or equal to MaxNumNonzero. On the other hand, when the conversion type of the current block is selected from a second candidate group including N2 MTS candidates, the number of non-zero coefficients belonging to the current block may be required to be less than or equal to MaxNumNonzero. Here, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 is an integer greater than N1 and may be 2, 3, 5, 6 or more.
[0169] Alternatively, when the transformation type of the current block is selected from the first candidate group including N1 MTS candidates, the number of non-zero coefficients belonging to the current block may be smaller than MinNumNonzero, or may be larger than or equal to MinNumNonzero and smaller than or equal to MaxNumNonzero. On the other hand, when the transformation type of the current block is selected from the second candidate group including N2 MTS candidates, the number of non-zero coefficients belonging to the current block is required to be larger than or equal to MinNumNonzero and smaller than or equal to MaxNumNonzero. Here, N1 may be 1 and N2 may be 4. However, it is not limited thereto, and N2 is an integer larger than N1 and may be 2, 3, 5, 6 or more.
[0170] The above-described Methods 1 to 3 may be applied as independent embodiments. Alternatively, based on at least two combinations of Methods 1 to 3, it may be required that at least two of the sum of the transformation coefficients, the position of the last valid coefficient, or the number of non-zero coefficients belong to a predetermined range.
[0171] The MTS index indicating the selected MTS candidate may be encoded and inserted into the bitstream. When two MTS candidates are selected, the MTS index may be encoded and inserted into the bitstream according to the transformation direction.
[0172] The MTS index may be encoded by context-based adaptive binary arithmetic coding (CABAC). The maximum value (cMax) for the binary evolution of the MTS index may be determined based on at least one of the sum of the transformation coefficients belonging to the current block, the number of transformation coefficients in the current block, the magnitude of one or more transformation coefficients in the current block, the position of the last valid coefficient in the current block, or the number of non-zero coefficients in the transformation coefficients of the current block. Alternatively, the maximum value (cMax) for the binary evolution of the MTS index may be determined based on the number of MTS candidates belonging to the candidate group when the transformation type of the current block is selected.
[0173] Alternatively, one or two MTS candidates among a plurality of MTS candidates may be selected based on the encoding parameters of the current block without encoding the MTS index. Here, the encoding parameters may include at least one of a prediction mode indicating an intra mode or an inter mode, the width and / or height of the current block (or, the transform block), or the split information of the current block (or, the transform block). The split information may include at least one of information indicating the presence or absence of splitting, information indicating the splitting direction, information regarding the number of partitions generated by the block splitting, information regarding the position of the partition to which the inverse transform is applied among the partitions generated by the block splitting, or information indicating the presence or absence of asymmetric splitting.
[0174] FIG. 7 is a diagram showing a schematic configuration of a conversion unit 232 that performs the conversion method according to the present disclosure.
[0175] The conversion method performed in the encoding device has been described with reference to FIG. 6, which may be identically performed in the conversion unit 232 of the encoding device, and the following detailed description will be omitted.
[0176] Referring to FIG. 7, the conversion unit 232 may include a conversion type determination unit 700 and a conversion coefficient derivation unit 710.
[0177] The conversion type determination unit 700 can determine the conversion type of the current block. That is, the conversion type determination unit 700 can determine the conversion type for the horizontal and / or vertical conversion of the current block based on any one of the MTS candidates already defined in the encoding device and the decoding device, respectively.
[0178] The conversion type determination unit 700 can select an MTS candidate from any one of a plurality of candidate groups and determine the conversion type for the horizontal conversion of the current block and / or the conversion type for the vertical conversion based on the selected MTS candidate.
[0179] Alternatively, the conversion type determination unit 700 can select two MTS candidates from any one of the plurality of candidate groups. In this case, the conversion type for the horizontal conversion of the current block can be determined based on either one of the two MTS candidates, and the conversion type for the vertical conversion of the current block can be determined based on the other one.
[0180] The conversion coefficient derivation unit 710 can perform conversion on the residual samples of the current block based on the conversion type of the current block, and derive the conversion coefficients of the current block.
[0181] The derived conversion coefficients may be required to satisfy at least one of the requirements of the above-described methods 1 to 3 according to the candidate group (or the number of MTS candidates belonging to the candidate group) for which the conversion type of the current block is selected.
[0182] The MTS index indicating the selected MTS candidate may be encoded by the entropy encoding unit 240 and inserted into the bit stream. When two MTS candidates are selected, the MTS index may be encoded separately for each conversion direction. The entropy encoding unit 240 can encode the MTS index by context-based adaptive binary arithmetic coding (CABAC).
[0183] Alternatively, the conversion type determination unit 700 can select one or two MTS candidates based on the encoding parameters of the current block, where the encoding parameters are as described above.
[0184] In the above-described embodiments, the method is described based on a sequence diagram by a series of steps or blocks, but the embodiments are not limited to the order of the steps. A certain step may occur in a different order or simultaneously with a step different from the above-described one. Also, those skilled in the art will understand that the steps shown in the sequence diagram are not exclusive, and other steps may be included, or one or more steps of the sequence diagram may be deleted without affecting the scope of the embodiments of this document.
[0185] The method according to the embodiments of the above-described document may be embodied in the form of software, and the encoding device and / or decoding device according to the document may be included in, for example, a device that performs image processing such as a TV, a computer, a smartphone, a set-top box, or a display device.
[0186] When the embodiments in this document are embodied as software, the above-described method may be embodied as a module (process, function, etc.) that executes the above-described functions. The module may be stored in a memory and executed by a processor. The memory may be provided inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory may include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document may be embodied and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each figure may be embodied and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for the embodiment (for example, information on instructions) or an algorithm may be stored in a digital storage medium.
[0187] In addition, the decoding device and encoding device to which the embodiments of this specification are applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a picture phone video device, a transportation means terminal (for example, a vehicle terminal (including an autonomous driving vehicle), an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and may be used to process a video signal or a data signal. For example, the over-the-top video (OTT) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0188] In addition, the processing method to which the embodiments of this specification are applied may be produced in the form of a program executed by a computer and may be stored in a computer-readable recording medium. Similarly, the multimedia data having the data structure according to the embodiments of this specification may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, the bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted through a wired or wireless communication network.
[0189] In addition, the embodiments of this specification may be embodied as a computer program product by program code, and the program code may be executed by a computer according to the embodiments of this specification. The program code may be stored on a computer-readable carrier.
[0190] FIG. 8 shows an example of a content streaming system to which the embodiments of the present disclosure are applicable.
[0191] Referring to FIG. 8, the content streaming system to which the embodiments of this specification are applied may generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0192] The encoding server is responsible for compressing the content input from a multimedia input device such as a smartphone, a camera, a camcorder, etc. into digital data to generate a bitstream and transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0193] The bitstream may be generated by an encoding method or a bitstream generation method to which the embodiments of this specification are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0194] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium to inform the user of what services are available. If the user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include a separate control server. In this case, the control server is responsible for controlling commands / responses between each device in the content streaming system.
[0195] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0196] Examples of the user device include mobile phones, smartphones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, and the like.
[0197] Each server in the content streaming system may be operated as a distributed server, and in this case, the data received by each server may be processed distributively.
[0198] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined and embodied as an apparatus, and the technical features of the apparatus claims in this specification may be combined and embodied as a method. Also, the technical features of the method claims in this specification and the technical features of the apparatus claims may be combined and embodied as an apparatus, and the technical features of the method claims in this specification and the technical features of the apparatus claims may be combined and embodied as a method.
Claims
Claim 1 Determining a conversion type for inverse conversion of the current block from one or more MTS candidates available for the current block; Performing inverse conversion on the conversion coefficients of the current block based on the determined conversion type to obtain residual samples of the current block; wherein the number of one or more MTS candidates available for the current block is determined based on at least one of the sum of conversion coefficients in the current block, the number of conversion coefficients in the current block, the magnitude of one or more conversion coefficients in the current block, the position of the last valid coefficient in the current block, or the number of one or more non-zero coefficients in the current block. A video decoding method. Claim 2 The video decoding method according to claim 1, wherein the number of one or more MTS candidates available for the current block is determined based on a comparison between the sum of conversion coefficients belonging to the current block and a predetermined threshold. Claim 3 The predetermined threshold includes a first threshold and a second threshold. When the sum of the conversion coefficients belonging to the current block is less than or equal to the first threshold, the number of the one or more MTS candidates is determined to be N 1 pieces, When the sum of the transform coefficients belonging to the current block is greater than the first threshold value and less than or equal to the second threshold value, the number of the one or more MTS candidates is determined to be N 2 and When the sum of the conversion coefficients belonging to the current block is greater than the second threshold, the number of the one or more MTS candidates is determined to be N 3 The video decoding method according to claim 2, wherein the number is N Claim 4 The predetermined threshold is determined based on a predetermined constant factor. The predetermined constant factor is determined based on at least one of a slice type, a quantization parameter of the current block, a size of the current block, a form of the current block, a ratio of a width to a height of the current block, or information signaled by a bitstream. The video decoding method according to claim 2. Claim 5 The video decoding method according to claim 1, wherein the number of one or more MTS candidates available for the current block is determined based on at least one of the position of the last valid coefficient in the current block, a scaling factor, a size of the current block, or a form of the current block. Claim 6 The video decoding method according to claim 1, wherein the number of one or more MTS candidates available for the current block is determined based on a comparison between the number of non-zero coefficients among the conversion coefficients of the current block and a third threshold. Claim 7 The video decoding method according to claim 1, wherein when the number of one or more MTS candidates available for the current block is two or more, the conversion type of the current block is determined based on an MTS index signaled by a bitstream. Claim 8 The maximum value (cMax) for the evolution of the MTS index is determined based on at least one of the sum of the transform coefficients belonging to the current block, the number of transform coefficients in the current block, the magnitude of one or more transform coefficients in the current block, the position of the last valid coefficient in the current block, or the number of non-zero coefficients in the transform coefficients of the current block. The video decoding method according to claim 7.
9. Determining a transform type of a current block; Performing a transform on the residual samples of the current block based on the transform type of the current block to derive transform coefficients of the current block, including: The transform type of the current block is selected from one or more MTS candidates belonging to any one of a plurality of predefined candidate groups. Based on the candidate group in which the transform type of the current block is selected, at least one of the sum of the transform coefficients of the current block, the number of transform coefficients in the current block, the magnitude of one or more transform coefficients in the current block, the position of the last valid coefficient in the current block, or the number of non-zero coefficients in the current block is required to belong to a predetermined range. A video encoding method.
10. A computer-readable storage medium storing a bitstream generated by a video encoding method, The video encoding method includes: Determining a transform type of a current block; Performing a transform on the residual samples of the current block based on the transform type of the current block to derive transform coefficients of the current block, including: The transform type of the current block is selected from among one or more MTS candidates belonging to any one of a plurality of predefined candidate groups. Based on the candidate group in which the transform type of the current block is selected, at least one of the sum of the transform coefficients of the current block, the number of transform coefficients in the current block, the magnitude of one or more transform coefficients in the current block, the position of the last valid coefficient in the current block, or the number of non-zero coefficients in the current block is required to belong to a predetermined range. A storage medium.
11. obtaining a bitstream for an image, the bitstream determining a conversion type of a current block, inducing conversion coefficients of the current block by performing a conversion on residual samples of the current block based on the conversion type of the current block, and encoding the induced conversion coefficients to generate; transmitting data including the bitstream; and the conversion type of the current block is selected from one or more MTS candidates belonging to any one of a plurality of predefined candidate groups; Based on the candidate group in which the conversion type of the current block is selected, at least one of the magnitude of one or more conversion coefficients of the current block, the sum of the conversion coefficients of the current block, the number of conversion coefficients in the current block, the magnitude of one or more conversion coefficients in the current block, the position of the last valid coefficient in the current block, or the number of non-zero coefficients in the current block is required to belong to a predetermined range, the method for transmitting data for an image.
12. A video decoding apparatus that performs the video decoding method according to claim 1.
13. A video encoding apparatus that performs the video encoding method according to claim 9.
14. A data transmission apparatus for an image that performs the transmission method according to claim 11.