Transform-based Image Coding Method and Apparatus
By analyzing the LFNST index and deriving the transformation coefficient, the image encoding efficiency is improved, the problem of increasing information in the transmission and storage of high-resolution and high-quality images/videos is solved, and more efficient image compression is achieved.
Patent Information
- Application Number
- CN202080092757.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-13
- Filing Date
- 2020-11-13
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-11-13
AI Technical Summary
The prior art has the problem of increasing amount of information in the transmission and storage of high resolution and high quality images/videos, especially in broadcasting of virtual reality and artificial reality content, and it is necessary to improve image encoding efficiency.
The LFNST index is parsed by skipping flag values based on the respective transformation of the color components of the current block, the modified transformation coefficient is derived, and the residual sample of the target block is derived through the inverse transformation, thereby increasing the encoding efficiency of the LFNST index.
Improves overall image/video compression efficiency, enhances encoding efficiency of LFNST indexes, and reduces transmission and storage costs.
Smart Images

Figure CN114930848B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image coding technology, and more particularly, to a method and apparatus for encoding an image based on transformation in an image coding system. Background Art
[0002] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K or higher ultra-high-definition (UHD) images / videos has been growing in various fields. As image / video data becomes higher resolution and higher quality, the amount of information or bit volume transmitted increases compared to traditional image data. Therefore, when using a medium such as a traditional wired / wireless broadband line to transmit image data or using an existing storage medium to store image / video data, its transmission cost and storage cost increase.
[0003] In addition, today, interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and broadcasting of images / videos having image characteristics different from real images such as game images is increasing.
[0004] Therefore, there is a need for an efficient image / video compression technology that effectively compresses and transmits or stores and reproduces information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention
[0005] Technical issues
[0006] A technical aspect of the present disclosure is to provide a method and apparatus for increasing image encoding efficiency.
[0007] Another technical aspect of the present disclosure is to provide a method and apparatus for increasing the efficiency of LFNST index encoding.
[0008] Yet another technical aspect of the present disclosure is to provide a method and apparatus for increasing encoding efficiency of an LFNST index based on a transform skip flag.
[0009] Technical Solution
[0010] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method may include: parsing LFNST indexes based on respective transform skip flag values of color components of a current block; deriving modified transform coefficients by applying LFNST to the transform coefficients; and deriving residual samples of a target block based on an inverse primary transform of the modified transform coefficients.
[0011] The method also includes: deriving a DC significant coefficient variable indicating whether a significant coefficient exists in a DC component of the current block based on a transform skip flag value, and based on at least one of the respective transform skip flag values being 0, the DC significant coefficient variable can be set to zero, and the LFNST index can be parsed based on the DC significant coefficient variable being 0.
[0012] The DC significant coefficient variable may be initially set to 1 at a coding unit level of the current block, and when a transform skip flag value is 0, the DC significant coefficient variable may be changed to 0 at a residual coding level.
[0013] The transform skip flag of the current block may be signaled for each color component.
[0014] Deriving the modified transform coefficients may further include setting a plurality of variables for LFNST based on whether the LFNST index is not 0 and whether respective transform skip flag values of the color components are 0.
[0015] When the tree type of the current block is single tree, the DC significant coefficient variable may be derived based on the transform skip flag value of the luma component, the transform skip flag value of the chroma Cb component, and the transform skip flag value of the chroma Cr component.
[0016] When the tree type of the current block is dual-tree luma, the DC significant coefficient variable may be derived based on the value of the transform skip flag of the luma component.
[0017] When the tree type of the current block is dual-tree chroma, the DC significant coefficient variable may be derived based on the value of the transform skip flag of the chroma Cb component and the value of the transform skip flag of the chroma Cr component.
[0018] According to another embodiment of the present disclosure, an image decoding method performed by an encoding device is provided. The method includes applying LFNST to derive modified transform coefficients from transform coefficients, wherein deriving the modified transform coefficients includes applying multiple LFNST matrices to the transform coefficients to derive DC significant coefficient variables indicating whether the significant coefficients are present in a DC component of a current block; and deriving the modified transform coefficients based on the DC significant coefficient variables indicating that the significant coefficients are present in locations other than the DC component. The DC significant coefficient variables may be derived based on transform skip flag values of respective color components of the current block.
[0019] According to still another embodiment of the present disclosure, a digital storage medium storing image data including a bit stream generated according to an image encoding method performed by an encoding device and encoded image information may be provided.
[0020] According to yet another embodiment of the present disclosure, a digital storage medium storing image data including encoded image information and a bit stream so that a decoding device performs an image decoding method may be provided.
[0021] Technical Effects
[0022] According to the present disclosure, the overall image / video compression efficiency can be increased.
[0023] According to the present disclosure, the efficiency of LFNST index encoding can be increased.
[0024] According to the present disclosure, the encoding efficiency of the LFNST index can be increased based on the transform skip flag.
[0025] The effects that can be obtained through the specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood or derived from the present disclosure by a person of ordinary skill in the relevant field. Therefore, the specific effects of the present disclosure are not limited to those explicitly described in the present disclosure, and may include various effects that can be understood or derived based on the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a diagram schematically illustrating a configuration of a video / image encoding device to which the present disclosure is applicable.
[0027] Figure 2 is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure is applicable.
[0028] Figure 3 FIG. 1 is a diagram schematically illustrating a multi-conversion technology according to an embodiment of the present disclosure.
[0029] Figure 4 This is a diagram schematically illustrating intra-frame directional modes for 65 prediction directions.
[0030] Figure 5 is a diagram for describing an RST according to an embodiment of the present disclosure.
[0031] Figure 6 is a diagram illustrating an order of arranging output data of a forward primary transform into a one-dimensional vector according to an example.
[0032] Figure 7 is a diagram illustrating an order of arranging output data of a forward quadratic transform into a one-dimensional vector according to an example.
[0033] Figure 8 exemplifies the block shape to which LFNST is applied.
[0034] Figure 9is a diagram illustrating an arrangement of output data of a forward LFNST according to an example.
[0035] Figure 10 is a diagram illustrating clearing of zeros in a block to which 4×4 LFNST is applied according to an example.
[0036] Figure 11 is a diagram illustrating clearing of zeros in a block to which 8×8 LFNST is applied according to an example.
[0037] Figure 12 is a flowchart for describing a method of decoding an image according to an example.
[0038] Figure 13 is a flowchart for describing a method of encoding an image according to an example.
[0039] Figure 14 is a diagram schematically illustrating an example of a video / image encoding system to which an embodiment of the present disclosure is applicable.
[0040] Figure 15 is a diagram exemplarily illustrating a structural diagram of a content streaming system to which the present disclosure is applied. DETAILED DESCRIPTION
[0041] Although the present disclosure may be susceptible to various modifications and includes various embodiments, its specific embodiments have been shown by way of example in the accompanying drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the technical ideas of the present disclosure. Unless the context clearly indicates otherwise, the singular form may include the plural form. Terms such as "including" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components, or combinations thereof used in the following description and should therefore not be understood as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0042] In addition, for the convenience of describing different characteristic functions, each component in the drawings described herein is illustrated independently, however, it is not intended that each component is implemented by separate hardware or software. For example, any two or more of these components can be combined to form a single component, and any single component can be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent rights of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0043] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. In addition, in the accompanying drawings, the same reference numerals are used for the same components, and repeated description of the same components will be omitted.
[0044] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).
[0045] In this document, various embodiments related to video / image encoding may be provided, and unless otherwise specified, these embodiments may be combined with each other and performed.
[0046] In this document, video can refer to a collection of images over a period of time. Generally, a picture refers to a unit that represents an image in a specific time region, and a slice / tile is a unit that constitutes a part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.
[0047] A pixel or a picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. Alternatively, a sample may refer to a pixel value in a spatial domain, or when the pixel value is transformed into a frequency domain, it may refer to a transform coefficient in the frequency domain.
[0048] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region and information related to the region. A unit may include a luminance block and two chrominance (e.g., CB, CR) blocks. Depending on the situation, terms such as unit and block, region, etc. may be used interchangeably. In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0049] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". In addition, "A / B / C" may mean "at least one of A, B, and / or C".
[0050] Additionally, in this document, the term "or" should be interpreted as meaning "and / or." For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as meaning "additionally or alternatively."
[0051] In the present disclosure, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in the present disclosure, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as “at least one of A and B”.
[0052] Furthermore, in the present disclosure, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” Furthermore, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”
[0053] In addition, the brackets used in this disclosure may indicate "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, it may mean that "intra-frame prediction" is proposed as an example of "prediction." In other words, "prediction" in this disclosure is not limited to "intra-frame prediction," and "intra-frame prediction" is proposed as an example of "prediction." In addition, when "prediction (i.e., intra-frame prediction)" is indicated, it may also mean that "intra-frame prediction" is proposed as an example of "prediction."
[0054] Technical features described separately in one drawing in the present disclosure may be implemented separately or may be implemented simultaneously.
[0055] Figure 1 Schematically illustrates a configuration of a video / image encoding device to which the present disclosure is applicable. Hereinafter, the so-called video encoding device may include an image encoding device.
[0056] Reference Figure 1, the encoding device 100 may include an image divider 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 may include an inter-frame predictor 121 and an intra-frame predictor 122. The residual processor 130 may include a transformer 132, a quantizer 133, a dequantizer 134, and an inverse transformer 135. The residual processor 130 may further include a subtractor 131. The adder 150 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image divider 110, the predictor 120, the residual processor 130, the entropy encoder 140, the adder 150, and the filter 160 described above may be configured by one or more hardware components (e.g., an encoder chipset or processor). In addition, the memory 170 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may further include the memory 170 as an internal / external component.
[0057] The image divider 110 may divide the input image (or picture or frame) input to the encoding device 100 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a maximum coding unit (LCU), the coding units may be recursively divided according to a quadtree, binary tree, ternary tree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, a coding unit may be divided into multiple coding units of a deeper depth. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that has not been further divided. In this case, based on the coding efficiency according to the image characteristics, the maximum coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively divided into coding units of a deeper depth as needed, thereby allowing the optimally sized coding unit to be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be separated or divided from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0058] Depending on the situation, terms such as unit and block, region, etc. can be used interchangeably. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a picture element (pel) of a picture (or image).
[0059] The encoding device 100 subtracts the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 121 or the intra-frame predictor 122 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 132. In this case, as shown in the figure, the unit within the encoding device 100 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as a subtractor 131. The predictor can perform prediction on a processing target block (hereinafter referred to as the "current block") and can generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later, in the description of each prediction mode, the predictor can generate various information related to the prediction, such as prediction mode information, and can send the generated information to the entropy encoder 140. The information about the prediction can be encoded in the entropy encoder 140 and output in the form of a bitstream.
[0060] The intra-frame predictor 122 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference sample can be located in a nearby area of the current block or in an area separated from the current block by a certain distance. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes can be used depending on the configuration. The intra-frame predictor 122 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0061] The inter-frame predictor 121 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information regarding the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 121 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate was used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 121 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, the residual signal cannot be sent. In the case of motion information prediction (motion vector prediction, MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0062] The predictor 120 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to predict a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor can perform prediction on the block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding in games such as screen content coding (SCC). Although IBC basically performs prediction in the current picture, the prediction can be performed similarly to inter-frame prediction, except that a reference block is derived in the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information related to the palette table or palette index.
[0063] The prediction signal generated by the predictor (including the inter-frame predictor 121 and / or the intra-frame predictor 122) can be used to generate a reconstruction signal or to generate a residual signal. The transformer 132 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT means a transform obtained from a curve graph when the relationship information between pixels is represented by a curve graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or can be applied to non-square blocks of varying sizes.
[0064] The quantizer 133 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy encoder 140. The entropy encoder 140 may encode the quantized signals (information about the quantized transform coefficients) and output the encoded signals in a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 133 may rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 140 may perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 140 may encode information required for video / image reconstruction in addition to the quantized transform coefficients (e.g., syntax element values, etc.) together or separately. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream on a network abstraction layer (NAL) unit basis. The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. In addition, the video / image information may also include general constraint information. In the present disclosure, information and / or syntax elements sent from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information may be encoded by the above-mentioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 140 or a memory (not shown) that stores the signal may be configured as an internal / external element of the encoding device 100, or the transmitter may be included in the entropy encoder 140.
[0065] The quantized transform coefficients output from the quantizer 133 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform to the quantized transform coefficients using the dequantizer 134 and the inverse transformer 135, a residual signal (residual block or residual sample) can be reconstructed. The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 121 or the intra-frame predictor 122, so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as a reconstructed block. The adder 150 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the target picture, and as described later, the generated reconstructed signal can be used for inter-frame prediction of the next picture performed by filtering.
[0066] Furthermore, in the picture encoding and / or reconstruction process, luma mapping with chroma scaling (LMCS) may be applied.
[0067] The filter 160 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the filter 160 can store the modified reconstructed picture in the memory 170, more specifically, in the DPB of the memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive ring filter, bilateral filter, etc. As discussed later in the description of each filtering method, the filter 160 can generate various information related to filtering and send the generated information to the entropy encoder 140. The information about filtering can be encoded in the entropy encoder 140 and output in the form of a bitstream.
[0068] The modified reconstructed picture sent to the memory 170 may be used as a reference picture in the inter-frame predictor 121. In doing so, the encoding apparatus may avoid prediction mismatch in the encoding apparatus 100 and the decoding apparatus when applying inter-frame prediction, and may also improve encoding efficiency.
[0069] The memory 170DPB can store the modified reconstructed picture so that it can be used as a reference picture in the inter-frame predictor 121. The memory 170 can store the motion information of the block in the current picture from which the motion information has been derived (or encoded) and / or the motion information of the block in the already (or previously) reconstructed picture. The stored motion information can be sent to the inter-frame predictor 121 to be used as the motion information of the neighboring block or the motion information of the temporally neighboring block. The memory 170 can store the reconstructed samples of the reconstructed block in the current picture and send the reconstructed samples to the intra-frame predictor 122.
[0070] Figure 2 is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure is applicable.
[0071] Reference Figure 2 , the video decoding device 200 may include an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The predictor 230 may include an inter-frame predictor 232 and an intra-frame predictor 231. The residual processor 220 may include a dequantizer 221 and an inverse transformer 222. According to an embodiment, the entropy decoder 210, the residual processor 220, the predictor 230, the adder 240, and the filter 250 described above may be configured by one or more hardware components (e.g., a decoder chipset or a processor). In addition, the memory 260 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 260 as an internal / external component.
[0072] When a bit stream including video / image information is input, the decoding device 200 can Figure 1 The image is reconstructed correspondingly to the processing of the video / image information in the encoding device. For example, the decoding device 200 can derive the unit / block based on the information related to the block segmentation obtained from the bit stream. The decoding device 200 can perform decoding by using the processing unit applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided into a quadtree structure, a binary tree structure and / or a ternary tree structure using a coding tree unit or a maximum coding unit. One or more transformation units can be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the decoding device 200 can be reproduced by a reproducer.
[0073] The decoding device 200 may receive the data from the Figure 1The signal output by the encoding device can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. In the present disclosure, the signaled / received information and / or syntax elements described later can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 210 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, use the decoded target syntax element information and the decoded information of the neighboring and decoded target blocks or the information of the symbol / bin decoded in the previous step to determine the context model, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bin to generate the symbol corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded by the context model for the next symbol / bin after determining the context model. The information about prediction among the information decoded in the entropy decoder 210 can be provided to the predictor (inter-frame predictor 232 and intra-frame predictor 231), and the residual value (i.e., quantized transform coefficient) and related parameter information for which entropy decoding has been performed in the entropy decoder 210 can be input to the residual processor 220. The residual processor 220 can derive the residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded in the entropy decoder 210 can be provided to the filter 250. In addition, a receiver (not shown) that receives a signal output from the encoding device may also configure the decoding device 200 as an internal / external element, and the receiver may be a component of the entropy decoder 210. In addition, the decoding device according to the present disclosure may be referred to as a video / image / picture encoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 210, and the sample decoder may include at least one of a dequantizer 221, an inverse transformer 222, an adder 240, a filter 250, a memory 260, an inter-frame predictor 232, and an intra-frame predictor 231.
[0074] The dequantizer 221 can output the transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 221 can rearrange the quantized transform coefficients into a two-dimensional block. In this case, the rearrangement process can be performed based on the order of coefficient scanning performed in the encoding device. The dequantizer 221 can dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0075] The inverse transformer 222 obtains a residual signal (residual block, residual sample array) by performing inverse transform on the transform coefficients.
[0076] The predictor may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 210, and more specifically, the predictor may determine an intra / inter prediction mode.
[0077] The predictor 230 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to the prediction of a block, and can also apply intra prediction and inter prediction at the same time. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform intra-block copying (IBC) for the prediction of the block. Intra-block copying can be used for content image / video coding in games such as screen content coding (SCC). Although IBC basically performs prediction in the current block, the prediction can be performed similarly to inter prediction, except that a reference block is derived in the current block. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information related to the palette table and the palette index.
[0078] The intra-frame predictor 231 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the reference samples can be located in a region near the current block or in a region separated from the current block by a certain distance. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 231 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0079] The inter-frame predictor 232 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information regarding the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 232 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the prediction information can include information indicating the inter-frame prediction mode for the current block.
[0080] The adder 240 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor 230. When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.
[0081] The adder 240 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next processing target block in the current block, and as described later, the generated reconstructed signal may be output through filtering or used for inter prediction of the next picture.
[0082] Furthermore, in the picture decoding process, luma mapping with chroma scaling (LMCS) may be applied.
[0083] The filter 250 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 250 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can send the modified reconstructed picture to the memory 260, more specifically, to the DPB of the memory 260. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0084] The (modified) reconstructed picture stored in the DPB of the memory 260 can be used as a reference picture in the inter-frame predictor 232. The memory 260 can store motion information of blocks in the current picture from which motion information has been derived (or decoded) and / or motion information of blocks in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 232 and used as motion information of neighboring blocks or motion information of temporally neighboring blocks. The memory 260 can store reconstructed samples of the reconstructed blocks in the current picture and send the reconstructed samples to the intra-frame predictor 231.
[0085] In this specification, the embodiments described in each of the filter 250, the inter-frame predictor 232, and the intra-frame predictor 231 of the decoding device 200 may be applied equally or correspondingly to the filter 160, the inter-frame predictor 121, and the intra-frame predictor 122 of the encoding device 100, respectively.
[0086] As described above, prediction is performed in order to improve compression efficiency when performing video encoding. In doing so, a prediction block including prediction samples for a current block as an encoding target block can be generated. Here, the prediction block includes prediction samples in a spatial domain (or a pixel domain). The prediction block can be derived identically in an encoding device and a decoding device, and the encoding device can improve image coding efficiency by signaling to the decoding device information (residual information) about the residual between the original block and the prediction block, rather than the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed picture including the reconstructed block.
[0087] Residual information can be generated through a transform process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients, so that it can signal the relevant residual information to the decoding device (through a bitstream). Here, the residual information may include value information, position information, transform technology, transform kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / dequantization process based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive a residual block by dequantizing / inverse transforming the quantized transform coefficients to serve as a reference for inter-frame prediction of the next picture, and can generate a reconstructed picture based on the derived residual block.
[0088] Figure 3 The multi-conversion technology according to the embodiment of the present disclosure is schematically illustrated.
[0089] Reference Figure 3 , the converter can correspond to the aforementioned Figure 1 The converter in the encoding device, and the inverse converter may correspond to the aforementioned Figure 1 The inverse transformer in the encoding device, or Figure 2 An inverse transformer in a decoding device.
[0090] The transformer may derive (primary) transform coefficients by performing a primary transform based on the residual samples (residual sample array) in the residual block (S310). This primary transform may be referred to as a core transform. In this document, the primary transform may be based on a multi-transform selection (MTS), and when multiple transforms are used as the primary transform, it may be referred to as a multi-core transform.
[0091] Multi-core transform may refer to a method for performing transforms using discrete cosine transform (DCT) type 2 and discrete sine transform (DST) type 7, DCT type 8, and / or DST type 1 in addition. In other words, multi-core transform may refer to a transform method that transforms a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this document, the primary transform coefficients may be referred to as temporary transform coefficients from the perspective of the transformer.
[0092] In other words, when a conventional transform method is applied, transform coefficients can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, when a multi-core transform is applied, transform coefficients (or primary transform coefficients) can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this document, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transform types, transform kernels, or transform cores. These DCT / DST transform types may be defined based on basis functions.
[0093] When performing multi-core transformation, a vertical transform kernel and a horizontal transform kernel for the target block can be selected from the transform kernels, a vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can indicate the transform of the horizontal component of the target block, and the vertical transform can indicate the transform of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block) including the residual block.
[0094] In addition, according to an example, if a transform is performed once by applying MTS, the mapping relationship of the transform kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transform or the horizontal transform. For example, when the horizontal transform kernel is represented as trTypeHor and the vertical transform kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT8.
[0095] In this case, the MTS index information may be encoded and signaled to the decoding device to indicate any one of a plurality of transform kernel sets. For example, an MTS index of 0 may indicate that both trTypeHor and trTypeVer values are 0, an MTS index of 1 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 2 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 3 may indicate that both trTypeHor and trTypeVer values are 1 and 2, and an MTS index of 4 may indicate that both trTypeHor and trTypeVer values are 2.
[0096] In one example, the transformation kernel set according to the MTS index information is shown in the following table.
[0097] [Table 1]
[0098] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0099] The transformer may perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients (S320). A primary transform is a transform from the spatial domain to the frequency domain, while a secondary transform refers to a transform into a more compact representation using the correlation existing between the (primary) transform coefficients. The secondary transform may include an inseparable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a pattern-dependent non-separable secondary transform (MDNSST). NSST may represent a transform that performs a secondary transform on the (primary) transform coefficients derived by the primary transform based on a non-separable transform matrix to generate modified transform coefficients (or secondary transform coefficients) for the residual signal. Here, based on the non-separable transform matrix, the transform may be applied once to the (primary) transform coefficients without separating the vertical transform and the horizontal transform (or applying the horizontal / vertical transform independently). In other words, NSST is not applied separately to (primary) transform coefficients in the vertical and horizontal directions, and can represent, for example, a transform method in which a two-dimensional signal (transform coefficient) is rearranged into a one-dimensional signal through a specific predetermined direction (e.g., a row-first direction or a column-first direction) and then the modified transform coefficients (or secondary transform coefficients) are generated based on an inseparable transform matrix. For example, the row-first order is to arrange the M×N blocks in the order of the first row, the second row, ... and the Nth row, while the column-first order is to arrange the M×N blocks in the order of the first column, the second column, ... and the Mth column. NSST can be applied to the upper left area of a block (hereinafter referred to as a transform coefficient block) configured with (primary) transform coefficients. For example, when the width W and the height H of the transform coefficient block are both 8 or larger, 8×8 NSST can be applied to the upper left 8×8 area of the transform coefficient block. In addition, while both the width (W) and the height (H) of the transform coefficient block are 4 or more, when the width (W) or the height (H) of the transform coefficient block is less than 8, the 4×4 NSST may be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block. However, the embodiment is not limited thereto, and for example, even if only the condition that the width W or the height H of the transform coefficient block is 4 or more is satisfied, the 4×4 NSST may be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block.
[0100] Specifically, for example, if a 4×4 input block is used, the non-separable secondary transform may be performed as follows.
[0101] A 4×4 input block X can be represented as follows.
[0102] [Formula 1]
[0103]
[0104] If X is represented as a vector, then the vector It can be expressed as follows.
[0105] [Formula 2]
[0106]
[0107] In Equation 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to row-major order.
[0108] In this case, the non-separable quadratic transform can be calculated as follows.
[0109] [Formula 3]
[0110]
[0111] In this formula, denotes a transform coefficient vector, and T denotes a 16x16 (non-separable) transform matrix.
[0112] By using the above formula 3, the 16×1 transform coefficient vector can be derived And the vector can be scanned in order (horizontally, vertically, diagonally, etc.) Reorganized into 4×4 blocks. However, the above calculation is an example, and Hypercube-Givens Transform (HyGT) or the like may also be used for the calculation of the inseparable secondary transform in order to reduce the computational complexity of the inseparable secondary transform.
[0113] Furthermore, in the inseparable secondary transform, the transform kernel (or transform core, transform type) may be selected to be mode-dependent. In this case, the mode may include an intra prediction mode and / or an inter prediction mode.
[0114] As described above, an inseparable secondary transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 area included in the transform coefficient block when both W and H are equal to or greater than 8, and the 8×8 area can be the upper left 8×8 area in the transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 area included in the transform coefficient block when both W and H are equal to or greater than 4, and the 4×4 area can be the upper left 4×4 area in the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0115] Here, in order to select mode-dependent transform kernels, two inseparable secondary transform kernels may be configured for each transform set used for inseparable secondary transforms for both the 8×8 transform and the 4×4 transform, and four transform sets may exist. That is, four transform sets may be configured for the 8×8 transform, and four transform sets may be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform may include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform may include two 4×4 transform kernels.
[0116] However, as the size of the transform (ie, the size of the region to which the transform is applied) may be other than 8×8 or 4×4, for example, the number of sets may be n, and the number of transform kernels in each set may be k.
[0117] The transform set may be referred to as an NSST set or a LFNST set. A specific set among the transform sets may be selected, for example, based on the intra prediction mode of the current block (CU or subblock). A low-frequency non-separable transform (LFNST) may be an example of a reduced non-separable transform, which will be described later and represents a non-separable transform for low-frequency components.
[0118] For reference, for example, the intra prediction mode may include two non-directional (or non-angle) intra prediction modes and 65 directional (or angle) intra prediction modes. The non-directional intra prediction mode may include a plane intra prediction mode No. 0 and a DC intra prediction mode No. 1, and the directional intra prediction mode may include 65 intra prediction modes No. 2 to No. 66. However, this is an example, and this document may be applied even if the number of intra prediction modes is different. In addition, in some cases, intra prediction mode No. 67 may also be used, and intra prediction mode No. 67 may represent a linear model (LM) mode.
[0119] Figure 4 The intra directional mode for 65 prediction directions is schematically shown.
[0120] Reference Figure 4 , based on the intra prediction mode 34 having the upper left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode having a horizontal directionality and an intra prediction mode having a vertical directionality. Figure 4In FIG, H and V denote horizontal and vertical directivities, respectively, and numbers -32 to 32 indicate displacements of 1 / 32 units on the sample grid position. These numbers may represent offsets for mode index values. Intra-prediction modes 2 to 33 have horizontal directivities, and intra-prediction modes 34 to 66 have vertical directivities. Strictly speaking, intra-prediction mode 34 may be considered neither horizontal nor vertical, but may be classified as belonging to horizontal directivity when determining the transform set for the secondary transform. This is because the input data is transposed for a vertically oriented mode that is symmetrical based on intra-prediction mode 34, and the input data alignment method for the horizontal mode is used for intra-prediction mode 34. Transposing the input data means switching the rows and columns of two-dimensional M×N block data to N×M data. Intra-prediction mode 18 and intra-prediction mode 50 may represent horizontal intra-prediction mode and vertical intra-prediction mode, respectively, and intra-prediction mode 2 may be referred to as an upper right diagonal intra-prediction mode because intra-prediction mode 2 has a left reference pixel and performs prediction in the upper right direction. Similarly, intra-prediction mode 34 may be referred to as a bottom-right diagonal intra-prediction mode, and intra-prediction mode 66 may be referred to as a bottom-left diagonal intra-prediction mode.
[0121] According to an example, four transform sets according to intra prediction modes may be mapped, for example, as shown in the following table.
[0122] [Table 2]
[0123] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0 2<=predModeIntra<=12 1 13<=predModeIntra<=23 2 24<=predModeIntra<=44 3 45<=predModeIntra<=55 2 56<=predModeIntra<=80 1
[0124] As shown in Table 2, any one of four transform sets, ie, lfnstTrSetIdx, may be mapped to any one of four indexes (ie, 0 to 3) according to the intra prediction mode.
[0125] When it is determined that a specific set is used for an inseparable transform, one of the k transform cores in the specific set can be selected by an inseparable secondary transform index. The encoding device can derive an inseparable secondary transform index indicating a specific transform core based on a rate-distortion (RD) check, and can signal the inseparable secondary transform index to the decoding device. The decoding device can select one of the k transform cores in the specific set based on the inseparable secondary transform index. For example, an lfnst index value 0 can refer to a first inseparable secondary transform core, an lfnst index value 1 can refer to a second inseparable secondary transform core, and an lfnst index value 2 can refer to a third inseparable secondary transform core. Alternatively, an lfnst index value 0 can indicate that the first inseparable secondary transform is not applied to the target block, and lfnst index values 1 to 3 can indicate three transform cores.
[0126] The transformer can perform a non-separable secondary transform based on the selected transform kernel and can obtain modified (secondary) transform coefficients. As described above, the modified transform coefficients can be derived as transform coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device and transmitted to the dequantizer / inverse transformer in the encoding device.
[0127] In addition, as described above, if the secondary transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device and transmitted to the dequantizer / inverse transformer in the encoding device.
[0128] The inverse transformer may perform a series of processes in the reverse order of the order already performed in the above-mentioned transformer. The inverse transformer may receive the (dequantized) transform coefficients and derive the (primary) transform coefficients by performing a secondary (inverse) transform (S350), and may obtain the residual block (residual sample) by performing a primary (inverse) transform on the (primary) transform coefficients (S360). In this regard, from the perspective of the inverse transformer, the primary transform coefficients may be referred to as modified transform coefficients. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.
[0129] The decoding device may further include a secondary inverse transform application determiner (or an element for determining whether to apply a secondary inverse transform) and a secondary inverse transform determiner (or an element for determining a secondary inverse transform). The secondary inverse transform application determiner may determine whether to apply a secondary inverse transform. For example, the secondary inverse transform may be NSST, RST, or LFNST, and the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a secondary transform flag obtained by parsing the bitstream. In another example, the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a transform coefficient of a residual block.
[0130] The secondary inverse transform determiner may determine the secondary inverse transform. In this case, the secondary inverse transform determiner may determine the secondary inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra-frame prediction mode. In an embodiment, the secondary transform determination method may be determined depending on the primary transform determination method. Various combinations of primary and secondary transforms may be determined according to the intra-frame prediction mode. In addition, in an example, the secondary inverse transform determiner may determine the area to which the secondary inverse transform is applied based on the size of the current block.
[0131] In addition, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, a (separable) inverse transform can be performed once, and a residual block (residual sample) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0132] Furthermore, in the present disclosure, reduced quadratic transform (RST) in which the size of a transformation matrix (kernel) is reduced may be applied in the concept of NSST in order to reduce the amount of calculation and storage required for an inseparable quadratic transform.
[0133] In addition, the transformation kernel, transformation matrix, and coefficients constituting the transformation kernel matrix described in the present disclosure, that is, kernel coefficients or matrix coefficients, can be represented in 8 bits. This can be a condition for implementation in decoding devices and encoding devices, and compared with existing 9 bits or 10 bits, the amount of storage required to store the transformation kernel can be reduced, and performance degradation can be reasonably adapted. In addition, representing the kernel matrix in 8 bits can allow the use of small multipliers and can be more suitable for single instruction multiple data (SIMD) instructions for optimal software implementation.
[0134] In this specification, the term "RST" may refer to a transform performed on residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. When performing a downscale transform, the amount of computation required for the transform can be reduced due to the reduction in the size of the transform matrix. In other words, RST can be used to address the computational complexity issues that arise when transforming large blocks or non-separable transforms.
[0135] RST may be referred to by various terms such as reduced transform, reduced secondary transform, downscaling transform, simplified transform, and simple transform, and the names that RST may be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in a low-frequency region including non-zero coefficients in a transform block, it may be referred to as a low-frequency non-separable transform (LFNST). The transform index may be referred to as an LFNST index.
[0136] In addition, when performing a secondary inverse transform based on an RST, the inverse transformer 135 of the encoding device 100 and the inverse transformer 222 of the decoding device 200 may include: an inverse-reduced secondary transformer that derives modified transform coefficients based on an inverse RST of the transform coefficients; and an inverse primary transformer that derives residual samples of the target block based on an inverse primary transform of the modified transform coefficients. An inverse primary transform refers to an inverse transform of a primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.
[0137] Figure 5is a diagram illustrating an RST according to an embodiment of the present disclosure.
[0138] In this disclosure, a “target block” may refer to a current block to be encoded, a residual block, or a transform block.
[0139] In the RST according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, so that a reduced transformation matrix can be determined, where R is less than N. N can refer to the square of the length of the side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the reduction factor can refer to an R / N value. The reduction factor can be referred to as a reduction factor, a shrinkage factor, a simplification factor, a simple factor, or various other terms. In addition, R can be referred to as a reduction coefficient, but depending on the situation, the reduction factor can refer to R. In addition, depending on the situation, the reduction factor can refer to an N / R value.
[0140] In the example, the reduction factor or reduction coefficient may be signaled through the bitstream, but the example is not limited thereto. For example, a predetermined value for the reduction factor or reduction coefficient may be stored in each of the encoding device 100 and the decoding device 200, and in this case, the reduction factor or reduction coefficient may not be signaled separately.
[0141] The size of the reduced transform matrix according to an example may be R×N, which is smaller than N×N (the size of a conventional transform matrix), and may be defined as in Equation 4 below.
[0142] [Formula 4]
[0143]
[0144] Figure 5 The matrix T in the reduced transform block shown in (a) may refer to the matrix T of Equation 4. R×N .like Figure 5 As shown in (a), when the reduced transformation matrix T R×N When multiplied by the residual samples of the target block, the transform coefficients of the current block can be derived.
[0145] In an example, if the size of the block to which the transform is applied is 8×8 and R=16 (ie, R / N=16 / 64=1 / 4), then according to Figure 5 The RST of (a) can be expressed as a matrix operation shown in the following Equation 5. In this case, the storage and multiplication calculations can be reduced to about 1 / 4 by a reduction factor.
[0146] In the present disclosure, a matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix provided on the left side of the column vector.
[0147] [Formula 5]
[0148]
[0149] In formula 5, r1 to r 64 ∫ may represent the residual sample of the target block, and specifically may be a transform coefficient generated by applying one transform. As a result of the calculation of Equation 5, the transform coefficient c of the target block may be derived. i , and derive c i The process can be shown as Equation 6.
[0150] [Formula 6]
[0151]
[0152] As a result of the calculation of Equation 6, the transform coefficients c1 to c R That is, when R=16, the transform coefficients c1 to c 16 . If a normal transform is applied instead of RST and a transform matrix of 64×64 (N×N) size is multiplied by a residual sample of 64×1 (N×1) size, only 16 (R) transform coefficients are derived for the target block because RST is applied, although 64 (N) transform coefficients are derived for the target block. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data transmitted by the encoding device 100 to the decoding device 200 is reduced, and thus the transmission efficiency between the encoding device 100 and the decoding device 200 can be improved.
[0153] When considering the size of the transformation matrix, the size of the conventional transformation matrix is 64×64 (N×N), but the size of the reduced transformation matrix is reduced to 16×64 (R×N). Therefore, compared with the case of performing the conventional transformation, the memory usage when performing RST can be reduced by the R / N ratio. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional transformation matrix, the number of multiplication calculations (R×N) can be reduced by the R / N ratio when using the reduced transformation matrix.
[0154] In an example, the transformer 132 of the encoding device 100 may derive transform coefficients of the target block by performing a primary transform and a secondary transform based on an RST on the residual samples of the target block. These transform coefficients may be transmitted to the inverse transformer of the decoding device 200, and the inverse transformer 222 of the decoding device 200 may derive modified transform coefficients based on an inverse reduced secondary transform (RST) for the transform coefficients, and may derive residual samples of the target block based on an inverse primary transform for the modified transform coefficients.
[0155] According to the example, the inverse RST matrix T N×RThe size of is N×R which is smaller than the size of the conventional inverse transform matrix N×N and is similar to the reduced transform matrix T shown in Equation 4. R×N Has a transposition relationship.
[0156] Figure 5 The matrix T in the reduced inverse transform block shown in (b) t Can refer to the inverse RST matrix T N×R T (The superscript T refers to transposition). Figure 5 As shown in (b), when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. R×N T It can be expressed as (T R×N ) T N×R .
[0157] More specifically, when the inverse RST is used as the secondary inverse transform, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the residual samples of the target block can be derived.
[0158] In an example, if the size of the block to which the inverse transform is applied is 8×8 and R=16 (ie, R / N=16 / 64=1 / 4), then according to Figure 5 The RST of (b) can be expressed as the matrix operation shown in the following Equation 7.
[0159] [Formula 7]
[0160]
[0161] In formula 7, c1 to c 16 As a result of the calculation of Equation 7, r representing the modified transform coefficient of the target block or the residual sample of the target block can be derived. i , and derive r i The process can be shown as formula 8.
[0162] [Formula 8]
[0163]
[0164] As a result of the calculation of Equation 8, r1 to r2 representing the modified transform coefficients of the target block or the residual samples of the target block can be derived. N . From the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduced transform matrix is reduced to 64×16 (R×N), so the storage usage in the case of performing inverse RST can be reduced by the R / N ratio compared to the case of performing the conventional inverse transform. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional inverse transform matrix, the use of the inverse reduced transform matrix can reduce the number of multiplication calculations (N×R) by the R / N ratio.
[0165] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied according to the transform set in Table 2. Since a transform set includes two or three transforms (kernels) according to the intra prediction mode, it can be configured to select one of up to four transforms, including the case where the secondary transform is not applied. In the transform where the secondary transform is not applied, the application of the identity matrix can be considered. Assuming that indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case where the identity matrix is applied, that is, the case where the secondary transform is not applied), the transform index or LFNST index as a syntax element can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the upper left 8×8 block, the 8×8 NSST in the RST configuration can be specified by the transform index, or the 8×8 LFNST can be specified when LFNST is applied. 8×8lfnst and 8×8RST refer to transforms that can be applied to an 8×8 region included in a transform coefficient block when both W and H of a target block to be transformed are equal to or greater than 8, and the 8×8 region may be the upper left 8×8 region in the transform coefficient block. Similarly, 4×4lfnst and 4×4RST refer to transforms that can be applied to a 4×4 region included in a transform coefficient block when both W and H of a target block are equal to or greater than 4, and the 4×4 region may be the upper left 4×4 region in the transform coefficient block.
[0166] According to an embodiment of the present disclosure, for the transformation in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 area. Here, “maximum” means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is, when RST is performed by applying an m×48 transform kernel matrix (m≤16) to an 8×8 area, 48 pieces of data are input and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming an 8×8 area can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on the 48 pieces of data constituting the area other than the lower right 4×4 area among the 8×8 areas. Here, when the matrix operation is performed by applying the maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper left 4×4 area according to the scanning order, and the upper right 4×4 area and the lower left 4×4 area can be filled with zeros.
[0167] For the inverse transform in the decoding process, a transposed matrix of the aforementioned transform kernel matrix may be used. That is, when inverse RST or LFNST is performed in the inverse transform process performed by the decoding device, input coefficient data to which inverse RST is applied is arranged in a one-dimensional vector according to a predetermined arrangement order, and modified coefficient vectors obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector may be arranged in a two-dimensional block according to a predetermined arrangement order.
[0168] In summary, during the transform process, when RST or LFNST is applied to an 8×8 region, the 48 transform coefficients in the upper left, upper right, and lower left regions of the 8×8 region, excluding the lower right region, are subjected to a matrix operation with the 16×48 transform kernel matrix. For the matrix operation, the 48 transform coefficients are input as a one-dimensional array. When the matrix operation is performed, 16 modified transform coefficients are derived and arranged in the upper left region of the 8×8 region.
[0169] On the contrary, in the inverse transform process, when the inverse RST or LFNST is applied to an 8×8 area, 16 transform coefficients corresponding to the upper left area of the 8×8 area among the transform coefficients in the 8×8 area can be input in a one-dimensional array according to the scanning order, and can undergo a matrix operation with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix)*(16×1 transform coefficient vector)=(48×1 modified transform coefficient vector). Here, the n×1 vector can be interpreted as having the same meaning as the n×1 matrix, and can therefore be expressed as an n×1 column vector. In addition, * represents matrix multiplication. When the matrix operation is performed, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left area, the upper right area, and the lower left area of the 8×8 area except the lower right area.
[0170] When the inverse secondary transform is based on the RST, the inverse transformer 135 of the encoding device 100 and the inverse transformer 222 of the decoding device 200 may include an inverse downscaling secondary transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients and an inverse primary transformer for deriving residual samples of the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to an inverse transform of a primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.
[0171] The non-separable transform (LFNST) described above will be described in detail as follows: LFNST may include a forward transform performed by an encoding device and an inverse transform performed by a decoding device.
[0172] The encoding device receives as input a result (or a portion of the result) derived after applying a primary (core) transform, and applies a forward secondary transform (secondary transform).
[0173] [Formula 9]
[0174] y=G T x
[0175] In Equation 9, x and y are the input and output of the quadratic transform, respectively, G is a matrix representing the quadratic transform, and the transform basis vectors consist of column vectors. In the case of inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the transpose of the matrix G becomes G T dimension.
[0176] For inverse LFNST, the dimensions of the matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of 8 transformation basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.
[0177] On the other hand, for forward LFNST, the matrix G T The dimensions are [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transformation basis vectors from the upper parts of the [16×48] matrix and the [16×16] matrix, respectively.
[0178] Therefore, in the case of forward LFNST, a [48×1] vector or a [16×1] vector can be used as input x, and a [16×1] vector or a [8×1] vector can be used as output y. In video encoding and decoding, the output of the forward primary transform is two-dimensional (2D) data, so in order to construct a [48×1] vector or a [16×1] vector as input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data as the output of the forward transform.
[0179] Figure 6 is a diagram illustrating an order of arranging output data of a forward primary transform into a one-dimensional vector according to an example. Figure 6 The left figures of (a) and (b) show the order for constructing a [48×1] vector, and Figure 6 The right figures of (a) and (b) show the order for constructing a [16×1] vector. In the case of LFNST, the 2D data can be constructed by Figure 6 The same order as in (a) and (b) is sequentially arranged to obtain a one-dimensional vector x.
[0180] The arrangement direction of the output data of the forward primary transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in the horizontal direction relative to the diagonal direction, the output data of the forward primary transform can be arranged in the following manner: Figure 6 The output data of the forward primary transform are arranged in the order of (a) and when the intra prediction mode of the current block is in a vertical direction relative to the diagonal direction, the output data of the forward primary transform can be arranged in the order of (a) Figure 6 The output data of the forward primary transformation are arranged in the order of (b).
[0181] According to the example, different Figure 6 The arrangement order of (a) and (b) is the arrangement order of (a) and (b), and in order to derive and apply Figure 6If the arrangement order of (a) and (b) is the same as the result (y vector), the column vectors of the matrix G can be rearranged according to the arrangement order. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0182] Since the output y derived by Formula 9 is a one-dimensional vector, when two-dimensional data is required as input data in a process using the result of the forward quadratic transform as input (for example, in the process of performing quantization or residual coding), the output y vector of Formula 9 needs to be properly arranged as 2D data again.
[0183] Figure 7 is a diagram illustrating an order of arranging output data of a forward quadratic transform into two-dimensional blocks according to an example.
[0184] In the case of LFNST, the output values can be arranged in 2D blocks according to a predetermined scanning order. Figure 7 (a) shows that when the output y is a [16×1] vector, the output values are arranged at 16 positions of the 2D block according to the diagonal scanning order. Figure 7 (b) shows that when the output y is an [8×1] vector, the output values are arranged at 8 positions of the 2D block according to the diagonal scanning order, and the remaining 8 positions are filled with zeros. Figure 7 The X in (b) indicates that it is filled with zeros.
[0185] According to another example, since the order of processing the output vector y when performing quantization or residual encoding can be preset, the output vector y may not be arranged in a sequence such as Figure 7 However, in the case of residual coding, data encoding can be performed in 2D block (e.g., 4×4) units (e.g., CG (coefficient group)), and in this case, according to Figure 7 The data is arranged in a specific order in the diagonal scan order of .
[0186] In addition, the decoding apparatus may configure a one-dimensional input vector y by arranging two-dimensional data output through a dequantization process according to a preset scanning order for inverse transform. The input vector y may be output as an output vector x through the following equation.
[0187] [Equation 10]
[0188] x=Gy
[0189] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or a [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.
[0190] The output vector x is based on Figure 6 The order shown in is arranged in a two-dimensional block and is arranged as two-dimensional data, and the two-dimensional data becomes input data (or a part of input data) of an inverse primary transform.
[0191] Therefore, the inverse quadratic transform is overall the reverse of the forward quadratic transform process, and in the case of the inverse transform, unlike in the forward direction, the inverse quadratic transform is applied first and then the inverse primary transform.
[0192] In the inverse LFNST, one of 8 [48×16] matrices and 8 [16×16] matrices can be selected as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.
[0193] In addition, eight matrices can be derived from the four transform sets shown in Table 2 above, and each transform set can be composed of two matrices. Which transform set to use among the four transform sets is determined according to the intra prediction mode, and more specifically, the transform set is determined based on the value of the intra prediction mode extended by considering wide-angle intra prediction (WAIP). Which matrix is selected from the two matrices constituting the selected transform set is derived by index signaling. More specifically, 0, 1, and 2 can be used as the index values sent, 0 can indicate that LFNST is not applied, and 1 and 2 can indicate either of the two transform matrices constituting the transform set selected based on the intra prediction mode value.
[0194] Furthermore, as described above, which transform matrix of the [48×16] matrix and the [16×16] matrix is applied to the LFNST is determined by the size and shape of the transform target block.
[0195] Figure 8 is a diagram illustrating a block shape to which LFNST is applied. Figure 8 (a) shows a 4×4 block, Figure 8 (b) shows a 4×8 block and an 8×4 block, Figure 8 (c) shows a 4×N block or an N×4 block, where N is 16 or greater, Figure 8 (d) shows an 8×8 block, Figure 8 (e) shows an M×N block, where M≥8, N≥8, and N>8 or M>8.
[0196] exist Figure 8 In , blocks with thick borders indicate the areas where LFNST is applied. Figure 8 For the blocks (a) and (b), LFNST is applied to the top left 4×4 region, and for Figure 8 In the block (c), LFNST is applied separately to the two upper left 4×4 regions that are arranged consecutively. Figure 8In (a), (b), and (c), since LFNST is applied in units of 4×4 regions, this LFNST will be referred to as “4×4 LFNST” hereinafter. Depending on the matrix dimension of G, a [16×16] or [16×8] matrix can be applied.
[0197] More specifically, the [16×8] matrix is applied to Figure 8 (a) 4×4 block (4×4TU or 4×4CU), and the [16×16] matrix is applied to Figure 8 This is to adjust the worst-case computational complexity to 8 multiplications per sample.
[0198] about Figure 8 In (d) and (e), LFNST is applied to the upper left 8×8 region, and this LFNST is hereinafter referred to as "8×8LFNST". As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of forward LFNST, since a [48×1] vector (the X vector in Equation 9) is input as input data, not all sample values of the upper left 8×8 region are used as input values of the forward LFNST. That is, as can be seen from Figure 6 The left order of (a) or Figure 6 As can be seen from the left order of (b), a [48×1] vector can be constructed based on samples belonging to the remaining three 4×4 blocks while leaving the lower right 4×4 block as it is.
[0199] The [48×8] matrix can be applied to Figure 8 The 8×8 block (8×8TU or 8×8CU) in (d) and the [48×16] matrix can be applied to Figure 8 The 8×8 block in (e) is used to adjust the worst-case computational complexity to 8 multiplications per sample.
[0200] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data are generated (Y vector in Equation 9, [8×1] or [16×1] vector). In the forward LFNST, since the matrix G T The characteristic is that the amount of output data is equal to or less than the amount of input data.
[0201] Figure 9 is a diagram illustrating arrangement of output data of a forward LFNST according to an example, and shows blocks in which the output data of the forward LFNST is arranged according to block shapes.
[0202] exist Figure 9The shaded area on the upper left of the block shown corresponds to the area where the output data of the forward LFNST is located, the positions marked with 0 indicate samples filled with the value 0, and the remaining area represents the area not changed by the forward LFNST. In the area not changed by the LFNST, the output data of the forward primary transform remains unchanged.
[0203] As described above, since the size of the applied transformation matrix varies according to the shape of the block, the amount of output data also varies. Figure 9 , the output data of the forward LFNST may not completely fill the upper left 4×4 block. Figure 9 In the cases of (a) and (d), the [16×8] matrix and the A[48×8] matrix are applied to the block indicated by the bold line or the partial area inside the block, respectively, and the [8×1] vector is generated as the output of the forward LFNST. That is, according to Figure 7 The scanning order shown in (b) can only fill 8 output data, such as Figure 9 As shown in (a) and (d), the remaining 8 positions can be filled with 0. Figure 8 (d) The case of LFNST application blocks, such as Figure 9 As shown in (d), the two 4×4 blocks on the upper right and lower left adjacent to the upper left 4×4 block are also filled with values of 0.
[0204] As described above, basically, by signaling the LFNST index, it is specified whether to apply LFNST and the transformation matrix to be applied. Figure 9 As shown, when LFNST is applied, since the number of output data of the forward LFNST may be equal to or less than the number of input data, an area filled with zero values occurs as follows.
[0205] 1) If Figure 9 (a) shows samples from the eighth position and subsequent positions in the scanning order in the upper left 4×4 block, that is, from the ninth to sixteenth samples.
[0206] 2) If Figure 9 As shown in (d) and (e), when the [48×16] matrix or the [48×8] matrix is applied, two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scanning order.
[0207] Therefore, if non-zero data exists by checking areas 1) and 2), it is determined that LFNST is not applied, so that signaling of the corresponding LFNST index can be omitted.
[0208] According to an example, for example, in the case of LFNST adopted in the VVC standard, since LFNST index signaling is performed after residual coding, the encoding device can determine whether non-zero data (significant coefficients) exist at all locations within the TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform LFNST index signaling based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. When non-zero data does not exist in the areas specified in 1) and 2) above, LFNST index signaling is performed.
[0209] In addition, for the adopted LFNST, the following simplified method can be applied.
[0210] (i) According to an example, the number of output data of the forward LFNST may be limited to a maximum of 16.
[0211] exist Figure 8 In the case of (c), 4×4 LFNST can be applied to two 4×4 regions adjacent to the upper left, respectively, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to a maximum of 16, in the case of a 4×N / N×4 (N≥16) block (TU or CU), 4×4 LFNST is applied only to one 4×4 region in the upper left, and LFNST can be applied only to Figure 8 By doing this, the implementation of image coding can be simplified.
[0212] (ii) According to an example, zeroing can be additionally applied to areas to which LFNST is not applied. In this document, zeroing can mean filling all positions belonging to a specific area with a value of 0. That is, zeroing can be applied to areas that are unchanged due to LFNST, and the result of the forward primary transform can be maintained. As described above, since LFNST is divided into 4×4 LFNST and 8×8 LFNST, zeroing can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.
[0213] (ii)-(A) When 4×4 LFNST is applied, an area to which 4×4 LFNST is not applied may be cleared. Figure 10 is a diagram illustrating clearing of zeros in a block to which 4×4 LFNST is applied according to an example.
[0214] like Figure 10 As shown, for the block to which 4×4 LFNST is applied, that is, for Figure 9 For all blocks in (a), (b), and (c), the entire area where LFNST is not applied can be filled with zeros.
[0215] on the other hand, Figure 10(d) shows that when the maximum value of the number of output data according to the example forward LFNST is limited to 16, clearing is performed on the remaining blocks to which the 4×4 LFNST is not applied.
[0216] (ii)-(B) When 8×8 LFNST is applied, an area to which 8×8 LFNST is not applied may be cleared. Figure 11 is a diagram illustrating clearing of zeros in a block to which 8×8 LFNST is applied according to an example.
[0217] like Figure 11 As shown, for the block where 8×8 LFNST is applied, i.e., for Figure 9 For all blocks in (d) and (e), the entire area where LFNST is not applied can be filled with zeros.
[0218] (iii) Due to the zeroing presented in (ii) above, the area filled with zeros may not be the same as when LFNST is applied. Figure 9 In the case of LFNST, the wider region performs the zeroing proposed in (ii) to check whether there is non-zero data.
[0219] For example, when (ii)-(B) is applied, when checking Figure 9 After the zero-filled areas in (d) and (e) have non-zero data, additionally check Figure 11 Whether there is non-zero data in the area filled with 0, the signaling of the LFNST index can be performed only when there is no non-zero data.
[0220] Of course, even if the clearing proposed in (ii) is applied, the presence of non-zero data can be checked in the same way as the existing LFNST index signaling. Figure 9 After checking whether there is non-zero data in the block filled with zeros in , LFNST index signaling can be applied. In this case, the encoding device only performs zero clearing and the decoding device does not assume zero clearing, that is, it only checks whether non-zero data exists only in Figure 9 In the regions explicitly marked as 0 in , LFNST index parsing can be performed.
[0221] Various embodiments of applying the combination of the simplified methods ((i), (ii)-(A), (ii)-(B), (iii)) of LFNST can be derived. Of course, the combination of the simplified methods described above is not limited to the following embodiments, and any combination can be applied to LFNST.
[0222] Implementation Method
[0223] -Limit the number of output data of the forward LFNST to a maximum of 16 → (i)
[0224] - When 4×4 LFNST is applied, all regions to which 4×4 LFNST is not applied are cleared → (II)-(A)
[0225] - When 8×8 LFNST is applied, all regions where 8×8 LFNST is not applied are cleared → (II)-(B)
[0226] - After checking whether non-zero data also exists in the existing areas filled with zero values and the areas filled with zeros due to additional clearing ((ii)-(A), (ii)-(B)), signal the LFNST index only if no non-zero data exists → (iii).
[0227] In the case of an embodiment, when LFNST is applied, the area where non-zero output data can exist is limited to the interior of the upper left 4×4 area. Figure 10 (a) and Figure 11 In the case of (a), the eighth position in the scanning order is the last position where non-zero data can exist. Figure 10 (b) and (c) and Figure 11 In the case of (b), the sixteenth position in the scanning order (ie, the position of the lower right edge of the upper left 4×4 block) is the last position in which data other than 0 may exist.
[0228] Therefore, after applying LFNST, after checking whether non-zero data exists at a position not allowed by the residual encoding process (at a position beyond the last position), it may be determined whether to signal the LFNST index.
[0229] In the case of the zeroing method proposed in (ii), the amount of data ultimately generated when both the primary transform and LFNST are applied can be reduced, thereby reducing the amount of computation required to perform the entire transform process. Specifically, when LFNST is applied, since the output data from the forward primary transform is present in areas where LFNST is not applied, there is no need to generate data for areas that were zeroed during the forward primary transform. Consequently, the amount of computation required to generate the corresponding data can be reduced. Additional benefits of the zeroing method proposed in (ii) are summarized below.
[0230] First, as mentioned above, the amount of computation required to perform the entire transformation process is reduced.
[0231] In particular, when (ii)-(B) is applied, the worst-case computational effort is reduced, making the transform process lightweight. In other words, generally speaking, a large amount of computation is required to perform a large-scale transform. By applying (ii)-(B), the amount of data derived as a result of performing forward LFNST can be reduced to 16 or less. In addition, as the size of the entire block (TU or CU) increases, the effect of reducing the number of transform operations further increases.
[0232] Second, the amount of computation required for the entire transformation process can be reduced, thereby reducing the power consumption required to perform the transformation.
[0233] Third, the delay involved in the transformation process is reduced.
[0234] Secondary transforms such as LFNST add computational complexity to the existing primary transform, thus increasing the overall latency involved in performing the transform. In particular, in the case of intra prediction, since reconstructed data from neighboring blocks is used in the prediction process, the increased latency due to the secondary transform during encoding results in an increased latency until reconstruction. This can lead to an increase in the overall latency of intra prediction encoding.
[0235] However, if the clearing proposed in (ii) is applied, the delay time for performing one transform can be greatly reduced when LFNST is applied, maintaining or reducing the delay time of the entire transform, so that the encoding device can be implemented more simply.
[0236] In conventional intra prediction, the block currently to be encoded is regarded as one coding unit, and encoding is performed without segmentation. However, intra subpartitioning (ISP) encoding means performing intra prediction encoding by dividing the block currently to be encoded in the horizontal direction or the vertical direction. In this case, a reconstructed block can be generated by performing encoding / decoding in units of divided blocks, and the reconstructed block can be used as a reference block for the next divided block. According to an embodiment, in ISP encoding, one coding block can be divided into two or four sub-blocks and encoded, and in ISP, in one sub-block, intra prediction is performed with reference to the reconstructed pixel value of the sub-block located on the adjacent left side or the adjacent upper side. Hereinafter, "encoding" may be used as a concept including both encoding performed by an encoding device and decoding performed by a decoding device.
[0237] ISP divides the block predicted as part of the luma frame into two or four sub-partitions vertically or horizontally based on the block size. For example, the minimum block size to which ISP can be applied is 4×8 or 8×4. When the block size is larger than 4×8 or 8×4, the block is divided into four sub-partitions.
[0238] When ISP is applied, subblocks are sequentially encoded from left to right or from top to bottom according to the partition type (e.g., horizontally or vertically), and after performing reconstruction processing via inverse transform and intra prediction for one subblock, encoding of the next subblock can be performed. For the leftmost or topmost subblock, the reconstructed pixels of the already encoded coding block are referenced, as in the conventional intra prediction method. In addition, when each side of the subsequent internal subblock is not adjacent to the previous subblock, in order to derive the reference pixels adjacent to the corresponding side, the reconstructed pixels of the already encoded adjacent coding block are referenced, as in the conventional intra prediction method.
[0239] In ISP coding mode, all sub-blocks can be encoded with the same intra prediction mode, and a flag indicating whether ISP coding is used and a flag indicating in which direction to split (horizontally or vertically) can be signaled. At this time, the number of sub-blocks can be adjusted to 2 or 4 according to the shape of the block. When the size (width × height) of a sub-block is less than 16, it can be restricted so that division into corresponding sub-blocks is not allowed or ISP coding itself is not applied.
[0240] In the case of the ISP prediction mode, one coding unit is divided into two or four partition blocks (ie, subblocks) and predicted, and the same intra prediction mode is applied to the divided two or four partition blocks.
[0241] As described above, in the division direction, both the horizontal direction (when an M×N coding unit having a horizontal length and a vertical length of M and N, respectively, is divided in the horizontal direction, if the M×N coding unit is divided into two, the M×N coding unit is divided into M×(N / 2) blocks, and if the M×N coding unit is divided into four blocks, the M×N coding unit is divided into M×(N / 4) blocks) and the vertical direction (when the M×N coding unit is divided in the vertical direction, if the M×N coding unit is divided into two, the M×N coding unit is divided into (M / 2)×N blocks, and if the M×N coding unit is divided into four, the M×N coding unit is divided into (M / 4)×N blocks) are possible. When the M×N coding unit is divided in the horizontal direction, the partition blocks are encoded in a top-to-bottom order, and when the M×N coding unit is divided in the vertical direction, the partition blocks are encoded in a left-to-right order. In the case of horizontal (vertical) partitioning, the currently encoded partition block can be predicted by referring to the reconstructed pixel values of the upper (left) partition block.
[0242] A transform can be applied to the residual signal generated in units of partition blocks by the ISP prediction method. A multi-transform selection (MTS) technique based on a DST-7 / DCT-8 combination and the existing DCT-2 can be applied to a forward-based primary transform (core transform), and a forward low-frequency non-separable transform (LFNST) can be applied to the transform coefficients generated from the primary transform to generate the final modified transform coefficients.
[0243] That is, LFNST can be applied to partition blocks divided by applying the ISP prediction mode, and the same intra prediction mode is applied to the partitioned partition blocks, as described above. Therefore, when an LFNST set derived based on the intra prediction mode is selected, the derived LFNST set can be applied to all partition blocks. That is, because the same intra prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks.
[0244] According to an embodiment, LFNST can be applied only to transform blocks with both horizontal and vertical lengths of 4 or greater. Therefore, when the horizontal or vertical length of a partition block divided according to the ISP prediction method is less than 4, LFNST is not applied and the LFNST index is not signaled. In addition, when LFNST is applied to each partition block, the corresponding partition block can be regarded as a transform block. When the ISP prediction method is not applied, LFNST can be applied to the coding block.
[0245] A method of applying LFNST to each partition block will be described in detail.
[0246] According to an embodiment, after applying forward LFNST to each partition block, only a maximum of 16 (8 or 16) coefficients are left in the upper left 4×4 area in the transform coefficient scanning order, and then zeroing can be applied, where the remaining positions and areas are all filled with 0.
[0247] Alternatively, according to an embodiment, when the length of one side of the partition block is 4, LFNST is applied only to the upper left 4×4 region, and when the lengths of all sides of the partition block (i.e., width and height) are 8 or greater, LFNST can be applied to the remaining 48 coefficients within the upper left 8×8 region except for the lower right 4×4 region.
[0248] Alternatively, according to an embodiment, in order to adjust the worst-case computational complexity to 8 multiplications per sample, when each partition block is 4×4 or 8×8, only 8 transform coefficients may be output after applying the forward LFNST. That is, when the partition block is 4×4, an 8×16 matrix may be applied as the transform matrix, and when the partition block is 8×8, an 8×48 matrix may be applied as the transform matrix.
[0249] In the current VVC standard, LFNST index signaling is performed in units of coding units. Therefore, in ISP prediction mode and when LFNST is applied to all partition blocks, the same LFNST index value can be applied to the corresponding partition blocks. That is, when the LFNST index value is sent once at the coding unit level, the corresponding LFNST index can be applied to all partition blocks in the coding unit. As described above, the LFNST index value can have values of 0, 1, and 2, where 0 indicates the case where LFNST is not applied, and 1 and 2 indicate two transform matrices present in one LFNST set when LFNST is applied.
[0250] As described above, the LFNST set is determined by the intra prediction mode, and in the case of the ISP prediction mode, since all partition blocks in the coding unit are predicted in the same intra prediction mode, the partition blocks can refer to the same LFNST set.
[0251] As another example, LFNST index signaling is still performed in units of coding units, but in the case of ISP prediction mode, it is not determined whether LFNST is applied uniformly to all partition blocks, and for each partition block, whether to apply the LFNST index value signaled at the coding unit level and whether to apply LFNST can be determined by a separate condition. Here, a separate condition can be signaled in the form of a flag for each partition block through the bitstream, and when the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and when the flag value is 0, LFNST may not be applied.
[0252] Hereinafter, a method of maintaining the worst-case computational complexity when applying LFNST to the ISP mode will be described.
[0253] In the case of ISP mode, when LFNST is applied, in order to keep the number of multiplications per sample (or per coefficient, per position) to a certain value or less, the application of LFNST may be limited. Depending on the size of the partition block, the number of multiplications per sample (or per coefficient, per position) can be kept to 8 or less by applying LFNST as follows.
[0254] 1. When both the horizontal length and the vertical length of the partition block are 4 or greater, the same method as the worst-case computational complexity control method for LFNST in the current VVC standard can be applied.
[0255] That is, when the partition block is a 4×4 block, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix may be applied in the forward direction instead of the 16×16 matrix, and a 16×8 matrix obtained by sampling the left 8 columns from the 16×16 matrix may be applied in the reverse direction. Furthermore, when the partition block is an 8×8 block, in the forward direction, an 8×48 matrix obtained by sampling the upper 8 rows from a 16×48 matrix may be applied instead of the 16×48 matrix, and in the reverse direction, a 48×8 matrix obtained by sampling the left 8 columns from a 48×16 matrix may be applied instead of the 48×16 matrix.
[0256] In the case of a 4×N or N×4 (N>4) block, when performing forward transform, the 16 coefficients generated after applying the 16×16 matrix only to the upper left 4×4 block can be set in the upper left 4×4 area, and the other areas can be filled with a value of 0. In addition, when performing inverse transform, the 16 coefficients located in the upper left 4×4 block are arranged in scan order to form an input vector, and then 16 output data can be generated by multiplying the 16×16 matrix. The generated output data can be set in the upper left 4×4 area, and the remaining areas except the upper left 4×4 area can be filled with a value of 0.
[0257] In the case of an 8×N or N×8 (N>8) block, when performing forward transform, the 16×48 matrix is applied to the ROI region within only the upper left 8×8 block (except for the remaining regions other than the lower right 4×4 block in the upper left 8×8 block). The generated 16 coefficients can be set in the upper left 4×4 region, and all other regions can be filled with a value of 0. In addition, when performing inverse transform, the 16 coefficients located in the upper left 4×4 region are arranged in scan order to form an input vector, and then 48 output data can be generated by multiplying the 48×16 matrix. The generated output data can be filled in the ROI region, and all other regions can be filled with a value of 0.
[0258] As another example, in order to keep the number of multiplications for each sample (or each coefficient, each position) at a certain value or less, the number of multiplications for each sample (or each coefficient, each position) based on the ISP coding unit size rather than the size of the ISP partition block can be kept at 8 or less. When only one block among the ISP partition blocks meets the conditions for applying LFNST, the worst-case complexity calculation of LFNST can be applied based on the corresponding coding unit size rather than the size of the partition block. For example, when the luminance coding block of a certain coding unit is divided into four partition blocks of size 4×4 and encoded by ISP, and there are no non-zero transform coefficients for two of the four partition blocks, it can be configured so that instead of eight transform coefficients, 16 transform coefficients are generated for the other two partition blocks (based on the encoder).
[0259] Hereinafter, a method of signaling an LFNST index in the ISP mode will be described.
[0260] As described above, the LFNST index can have any value of 0, 1, or 2, where a value of 0 indicates that LFNST is not applied, and values 1 and 2 indicate the two LFNST kernel matrices included in the selected LFNST set, respectively. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The method for transmitting the LFNST index in the current VVC standard will be described below.
[0261] 1. LFNST index can be sent once for each coding unit (CU), and in the case of dual-tree type, separate LFNST indexes can be signaled for luma blocks and chroma blocks.
[0262] 2. When the LFNST index is not signaled, the value of the LFNST index is set (inferred) to a default value of 0. The case where the value of the LFNST index is inferred to be 0 is as follows:
[0263] A. A case where a mode that does not apply transform (eg, transform skip, BDPCM, lossless coding, etc.) is used.
[0264] B. A case where one transform is not DCT-2 (DCT7 or DCT8), that is, a case where the transform in the horizontal direction or the transform in the vertical direction is not DCT-2.
[0265] C. When the horizontal length or vertical length of the luma block of a coding unit exceeds the maximum luma transformable size, for example, when the maximum luma transform size is 64, LFNST is not applied when the luma block size of the coding block is 128×16.
[0266] In the case of the dual-tree type, for each of the coding units of the luma component and the chroma component, it is determined whether the maximum luma transformable size is exceeded. That is, for the luma block, it is checked whether the maximum luma transformable size is exceeded, and for the chroma block, it is checked whether the horizontal / vertical length and the maximum luma transformable size of the corresponding luma block for the color format are exceeded. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luma block is twice the horizontal / vertical length of the chroma block, and the transform size of the corresponding luma block is twice the transform size of the chroma block. In another example, when the color format is 4:4:4, the horizontal / vertical length and transform size of the corresponding luma block are the same as those of the chroma block.
[0267] A 64-length transform or a 32-length transform refers to a transform applied horizontally or vertically with a length of 64 or 32, respectively, and the “transform size” may refer to the corresponding length 64 or 32.
[0268] In the case of a single tree type, it may be checked whether the horizontal length or vertical length of the luminance block exceeds the maximum transformable size of the luminance transform block, and when the horizontal length or vertical length of the luminance block exceeds the maximum transformable size of the luminance transform block, signaling of the LFNST index may be omitted.
[0269] D. The LFNST index may be transmitted only when both the horizontal length and the vertical length of the coding unit are greater than or equal to 4.
[0270] In case of the dual-tree type, the LFNST index may be signaled only when both the horizontal length and the vertical length of the corresponding component (ie, the luma component or the chroma component) are each greater than or equal to 4.
[0271] In case of a single tree type, the LFNST index may be signaled when both the horizontal length and the vertical length of the luma component are greater than or equal to 4, respectively.
[0272] E. If the last non-zero coefficient position is not the DC position (which is the position located at the upper left of the block), if the last non-zero coefficient position is not the DC position for the dual-tree type luma block, the LFNST index is transmitted. In the case of a dual-tree type chroma block, if either the last non-zero coefficient position of Cb or the last non-zero coefficient position of Cr is not the DC position, the corresponding LNFST index is transmitted.
[0273] In case of a single tree type, when the last non-zero coefficient position of any one of the luma component, the Cb component, and the Cr component is not the DC position, the LFNST index is transmitted.
[0274] Here, when the coded block flag (CBF) value indicating whether a transform coefficient of a transform block exists is 0, the position of the last non-zero coefficient of the corresponding transform block is not checked to determine whether to signal the LFNST index. That is, when the corresponding CBF value is 0, the transform is not applied to the corresponding block, and therefore the position of the last non-zero coefficient may not be considered when checking the conditions for LFNST index signaling.
[0275] For example, 1) in the case of a dual-tree type luma component, if the corresponding CBF value is 0, the LFNST index is not signaled; 2) in the case of a dual-tree type chroma component, if the CBF value of Cb is 0 and the CBF value of Cr is 1, only the last non-zero coefficient position of Cr is checked and the corresponding LFNST index is sent; 3) in the case of a single-tree type, the last non-zero coefficient position of any component whose CBF value is 1 among the luma component, Cb component, and Cr component is checked.
[0276] F. When a transform coefficient is found to exist at a position other than the position where the LFNST transform coefficient is allowed to exist, the signaling of the LFNST index can be omitted. In the case of 4×4 transform blocks and 8×8 transform blocks, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient can exist in 8 positions starting from the DC position and all the remaining positions are filled with 0. In addition, in cases other than 4×4 transform blocks and 8×8 transform blocks, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient can exist in 16 positions starting from the DC position and the remaining positions are all filled with 0.
[0277] Therefore, when a non-zero transform coefficient exists in a region that should be filled with 0 after residual encoding is performed, LFNST index signaling may be omitted.
[0278] In addition, the ISP mode can be applied only to the luminance block, or to both the luminance block and the chrominance block. As described above, when ISP prediction is applied, the corresponding coding unit is divided into two or four partition blocks and predicted, and a transform can be applied to each corresponding partition block. Therefore, when determining the conditions for signaling the LFNST index in units of coding units, it is necessary to consider the fact that LFNST is applicable to each of the corresponding partition blocks. In addition, when the ISP prediction mode is applied only to a specific component (e.g., a luminance block), the LFNST index should be signaled taking into account the fact that only the corresponding component is divided into partition blocks. The LFNST index signaling methods available in ISP mode can be summarized as follows.
[0279] 1. The LFNST index may be transmitted once for each coding unit (CU), and in the case of a dual-tree type, separate LFNST indexes may be signaled for luma blocks and chroma blocks, respectively.
[0280] 2. When the LFNST index is signaled, the value of the LFNST index is set (inferred) to a default value of 0. The case where the LFNST index value is inferred to be 0 is as follows.
[0281] A. A case where a mode that does not apply transform (eg, transform skip, BDPCM, lossless coding, etc.) is used.
[0282] B. When the horizontal length or vertical length of the luminance block of the coding unit exceeds the maximum luminance transformable size, for example, when the maximum transformable luminance size is 64, LFNST cannot be applied when the size of the luminance block of the coding block is 128×16.
[0283] Whether to signal the LFNST index can be determined based on the size of the partition block rather than the size of the coding unit. That is, when the horizontal length or vertical length of the partition block corresponding to the luma block exceeds the maximum luma transformable size, LFNST index signaling can be omitted and the value of the LFNST index can be inferred to be 0.
[0284] In the case of the dual tree type, it is determined whether the maximum transform block size is exceeded for each of the coding unit or partition block of the luma component and the coding unit or partition block of the chroma component. That is, the horizontal length and vertical length of the coding unit or partition block for luma are compared with the maximum luma transformable size, and if either of the horizontal length and vertical length is greater than the maximum luma transformable size, LFNST is not applied, and in the case of the coding unit or partition block for chroma, the horizontal / vertical length of the corresponding luma block for the color format is compared with the maximum luma transformable size. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luma block is twice the horizontal / vertical length of the chroma block, and the transform size of the corresponding luma block is twice the transform size of the chroma block. In another example, when the color format is 4:4:4, the horizontal / vertical length and transform size of the corresponding luma block are the same as those of the chroma block.
[0285] In case of a single tree type, it may be checked whether the horizontal length or vertical length of a luma block (coding unit or partition block) exceeds the maximum transformable size of the luma transform block, and if so, LFNST index signaling may be omitted.
[0286] C. When LFNST included in the current VVC standard is applied, the LFNST index may be transmitted only when both the horizontal length and the vertical length of the partition block are greater than or equal to 4.
[0287] When LFNST for 2×M (1×M) or M×2 (M×1) blocks is applied in addition to LFNST included in the current VVC standard, the LFNST index may be transmitted only when the size of the partition block is equal to or larger than 2×M (1×M) or M×2 (M×1) blocks. Here, if a P×Q block is equal to or larger than an R×S block, this means P ≥ R and Q ≥ S.
[0288] In summary, the LFNST index can be sent only when the partition block is equal to or larger than the minimum size for which LFNST can be applied. In the case of a dual-tree type, the LFNST index can be signaled only when the partition block of the luma or chroma component is equal to or larger than the minimum size for which LFNST can be applied. In the case of a single-tree type, the LFNST index can be signaled only when the partition block of the luma component is equal to or larger than the minimum size for which LFNST can be applied.
[0289] In this document, if an M×N block is equal to or larger than a K×L block, this means M is equal to or larger than K and N is equal to or larger than L. If an M×N block is larger than a K×L block, this means M is equal to or larger than K, N is equal to or larger than L, and either M is larger than K or N is larger than L. If an M×N block is smaller than or equal to a K×L block, this means M is smaller than or equal to K and N is smaller than or equal to L, and if an M×N block is smaller than a K×L block, this means M is smaller than or equal to K, N is smaller than or equal to L, and either M is smaller than K or N is smaller than L.
[0290] D. In the case where the last non-zero coefficient position is not the DC position (which is the position located at the upper left of the block), if the last non-zero coefficient position in any one of all partition blocks is not the DC position for the dual-tree type luminance block, the LFNST index may be transmitted. In the case of the dual-tree type chrominance block, if any of the last non-zero coefficient positions of all partition blocks for Cb (when the ISP mode is not applied to the chrominance component, the number of partition blocks is considered to be 1) and the last non-zero coefficient position of all partition blocks for Cr (when the ISP mode is not applied to the chrominance component, the number of partition blocks is considered to be 1) is not the DC position, the corresponding LNFST index may be transmitted.
[0291] In case of a single tree type, when the last non-zero coefficient position of any one of all partition blocks of luma components, Cb components, and Cr components is not a DC position, a corresponding LFNST index may be transmitted.
[0292] Here, when the coded block flag (CBF) value indicating whether a transform coefficient exists for each partition block is 0, the last non-zero coefficient position of the corresponding partition block is not checked to determine whether to signal the LFNST index. That is, when the corresponding CBF value is 0, the transform is not applied to the corresponding block, and therefore the last non-zero coefficient position of the corresponding partition block is not considered when checking the conditions for LFNST index signaling.
[0293] For example, 1) in the case of the luminance component of the dual-tree type, when the corresponding CBF value of each partition block is 0, the corresponding partition block is excluded when determining whether to signal the LFNST index; 2) in the case of the chrominance component of the dual-tree type, when the CBF value of Cb of each partition block is 0 and the CBF value of Cr is 1, only the last non-zero coefficient position of Cr is checked when determining whether to signal the LFNST index; and 3) in the case of the single-tree type, whether to signal the LFNST index can be determined by checking only the last non-zero coefficient position of the block with a CBF value of 1 for all partition blocks of the luminance component, Cb component, and Cr component.
[0294] In the case of the ISP mode, image information may be configured such that the last non-zero coefficient position is not checked, and its implementation is as follows.
[0295] i. In ISP mode, LFNST index signaling can be allowed without checking the last non-zero coefficient position of both the luma block and the chroma block. That is, even when the last non-zero coefficient position of all partition blocks is at the DC position or the corresponding CBF value is 0, the corresponding LFNST index signaling can be allowed.
[0296] ii. In the case of ISP mode, the check of the last non-zero coefficient position can be omitted only for the luma block, and the last non-zero coefficient position of the chroma block can be checked in the above manner. For example, in the case of a dual-tree type luma block, LFNST index signaling is allowed without checking the last non-zero coefficient position, while in the case of a dual-tree type chroma block, whether to signal the corresponding LFNST index can be determined by checking whether the DC position of the last non-zero coefficient position exists in the above manner.
[0297] iii. In the case of ISP mode and single tree type, the above method i or the above method ii can be applied. That is, when method i is applied to the single tree type in ISP mode, the check of the last non-zero coefficient position can be omitted for both the luminance block and the chrominance block, and LFNST index signaling is allowed. Alternatively, when method ii is applied, the check of the last non-zero coefficient position can be omitted for the partition block of the luminance component, and for the partition block of the chrominance component (when ISP is not applied to the chrominance component, the number can be considered to be 1), the last non-zero coefficient position can be checked in the above manner to determine whether to signal the corresponding LFNST index.
[0298] E. When a transform coefficient is found to exist at a position other than a position where the LFNST transform coefficient may exist for even one partition block among all partition blocks, LFNST index signaling may be omitted.
[0299] For example, in the case of a 4×4 partition block and an 8×8 partition block, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient may be present in 8 positions starting from the DC position, and all remaining positions are filled with 0. In addition, when the partition block is equal to or larger than 4×4 and is neither a 4×4 partition block nor an 8×8 partition block, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient may be present in 16 positions starting from the DC position, and all remaining positions are filled with 0.
[0300] Therefore, when there are non-zero transform coefficients in a region to which 0 padding is applied after residual encoding is performed, LFNST index signaling may be omitted.
[0301] On the other hand, in ISP mode, according to the current VVC standard, the horizontal and vertical directions are independently considered length conditions, and DST-7 is applied instead of DCT-2 without signaling the MTS index. A determination is made as to whether the horizontal length or vertical length is greater than or equal to 4 or greater than or equal to 16, and a primary transform kernel is determined based on the determination result. Therefore, for cases where LFNST can be applied in ISP mode, the following transform combinations are possible.
[0302] 1. When the LFNST index is 0 (including the case where the LFNST index is inferred to be 0), the one-time transform determination condition in the ISP mode can be followed, which is included in the current VVC standard. That is, if the length condition (which is equal to or greater than 4 and equal to or less than 16) is satisfied separately and independently for the horizontal direction and the vertical direction; if so, DST-7 can be applied instead of DCT-2, and if not, DCT-2 can be applied.
[0303] 2. When the LFNST index is greater than 0, the following two configurations can be used as a single transformation.
[0304] A. DCT-2 can be applied to both horizontal and vertical directions.
[0305] B. The conditions for determining a primary transform in ISP mode, which are included in the current VVC standard, may be followed. In other words, the horizontal and vertical directions are checked to see whether the length condition (which is equal to or greater than 4 and equal to or less than 16) is satisfied separately and independently; if so, DST-7 may be applied instead of DCT-2, and if not, DCT-2 may be applied.
[0306] In ISP mode, image information may be configured so that an LFNST index is transmitted for each partition block rather than for each coding unit. In this case, in the above-described LFNST index signaling method, whether to signal the LFNST index may be determined considering that there is only one partition block in the unit in which the LFNST index is transmitted.
[0307] Furthermore, signaling of the LFNST index and the MTS index will be described below.
[0308] The following table shows the coding unit syntax table, transform unit syntax table, and residual coding syntax table related to the signaling of LFNST index and MTS index according to an example. According to Table 3, the MTS index is moved from the transform unit level to the coding unit level syntax and is signaled after the LFNST index signaling. In addition, the constraint that does not allow LFNST when the ISP is applied to the coding unit has been removed. When the ISP is applied to the coding unit, the constraint that does not allow LFNST is removed, so that LFNST can be applied to all intra-frame prediction blocks. In addition, both the MTS index and the LFNST index are conditionally signaled at the last part of the coding unit level.
[0309] [Table 3]
[0310]
[0311] [Table 4]
[0312]
[0313] [Table 5]
[0314]
[0315] The meanings of the main variables in the table are as follows.
[0316] 1.cbWidth, cbHeight: width and height of the current encoding block
[0317] 2. log2TbWidth, log2TbHeight: The base-2 logarithmic values of the width and height of the current transform block, and reflect zeroing to reduce to the upper left area where non-zero coefficients may exist.
[0318] 3. sps_lfnst_enabled_flag: It is a flag indicating whether LFNST is enabled, if the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in the sequence parameter set (SPS).
[0319] 4. CuPredMode[chType][x0][y0]: A prediction mode corresponding to the variable chType and the (x0, y0) position, chType can have values of 0 and 1, where 0 represents a luma component and 1 represents a chroma component. The (x0, y0) position indicates a position on a picture, and MODE_INTRA (intra-frame prediction) and MODE_INTER (inter-frame prediction) can have the CuPredMode[chType][x0][y0] value.
[0320] 5. IntraSubPartitionsSplit[x0][y0]: The content of the (x0, y0) position is the same as in item 4. It indicates which ISP split is applied at the (x0, y0) position, and ISP_NO_SPLIT indicates that the coding unit corresponding to the (x0, y0) position is not divided into partition blocks.
[0321] 6. intra_mip_flag[x0][y0]: The content of the position (x0, y0) is the same as in item 4 above. intra_mip_flag is a flag indicating whether matrix-based intra prediction (MIP) prediction mode is applied. A flag value of 0 indicates that MIP is not enabled, while a flag value of 1 indicates that MIP is applied.
[0322] 7. cIdx: A value of 0 indicates luma, and values of 1 and 2 indicate Cb and Cr of the chroma components, respectively.
[0323] 8.treeType: It indicates single tree and dual tree etc. (SINGLE_TREE: single tree, DUAL_TREE_LUMA: dual tree for luma component, DUAL_TREE_CHROMA: dual tree for chroma component)
[0324] 9. lastSubBlock: It indicates the position of the subblock (coefficient group (CG)) in the scanning order where the last non-zero coefficient is located. 0 indicates a subblock containing a DC component, and if greater than 0, it is not a subblock containing a DC component.
[0325] 10.lastScanPos: It indicates where the last significant coefficient is located in the scan order within a subblock. If a subblock consists of 16 positions, it can have a value from 0 to 15.
[0326] 11. lfnst_idx[x0][y0]: LFNST index syntax element to be parsed. If it is not parsed, it is inferred to be 0. That is, the default value is set to 0, indicating that LFNST is not applied.
[0327] 12. LastSignificantCoeffX, LastSignificantCoeffY: This indicates the x-coordinate and y-coordinate of the last significant coefficient in the transform block. The x-coordinate starts at 0 and increases from left to right, and the y-coordinate starts at 0 and increases from top to bottom. If the value of both variables is 0, it means that the last significant coefficient is located at DC.
[0328] 13. cu_sbt_flag: This is a flag indicating whether sub-block transform (SBT) included in the current VVC standard is enabled. If the flag value is 0, it indicates that SBT is not enabled, and if the flag value is 1, it indicates that SBT is enabled.
[0329] 14.sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: It is a flag indicating whether explicit MTS is applied to inter CU and intra CU respectively. If the corresponding flag value is 0, it indicates that MTS is not applicable to inter CU or intra CU, and if it is 1, it indicates that MTS is applicable.
[0330] 15.tu_mts_idx[x0][y0]: This is the MTS index syntax element to be parsed. If it is not parsed, it is inferred to be 0. That is, the default value is set to 0, which indicates that DCT-2 is applied to both the horizontal and vertical directions.
[0331] As shown in Table 3, several conditions are checked when encoding mts_idx[x0][y0], and tu_mts_idx[x0][y0] is signaled only when the value of lfnst_idx[x0][y0] is 0.
[0332] In addition, tu_cbf_luma[x0][y0] is a flag indicating whether there is a significant coefficient for the luma component.
[0333] According to Table 3, when both the width and height of the coding unit of the luma component are 32 or less, mts_idx[x0][y0](Max(cbWidth,cbHeight)<=32) is signaled, that is, whether MTS is applied is determined by the width and height of the coding unit of the luma component.
[0334] In addition, according to Table 3, lfnst_idx[x0][y0] may be configured for signaling even in ISP mode (IntraSubPartitionsSplitType!=ISP_NO_SPLIT), and the same LFNST index value may be applied to all ISP partition blocks.
[0335] On the other hand, mts_idx[x0][y0] may be signaled only when not in ISP mode (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT).
[0336] In the process of determining log2ZoTbWidth and log2ZoTbHeight as shown in Table 5 (where log2ZoTbWidth and log2ZoTbHeight represent the base 2 logarithmic values of the width and height of the upper left area remaining after clearing is performed), the part of checking the mts_idx[x0][y0] value can be omitted.
[0337] Furthermore, according to an example, when log2ZoTbWidth and log2ZoTbHeight are determined in residual encoding, a condition for checking sps_mts_enable_flag may be added.
[0338] If there are valid coefficients at zero positions when LFNST is applied, the variable LfnstZeroOutSigCoeffFlag of Table 3 is 0, otherwise it is 1. The variable LfnstZeroOutSigCoeffFlag can be set according to several conditions shown in Table 5.
[0339] According to an example, the variable LfnstDcOnly in Table 3 becomes 1 when all the last significant coefficients are located at the DC position (upper left position) of the transform block having a corresponding coded block flag (CBF) with a value of 1 (1 if there is at least one significant coefficient in the block, otherwise 0), and otherwise it becomes 0. More specifically, in the case of dual-tree luma, the position of the last significant coefficient is checked for one luma transform block, and in the case of dual-tree chroma, the position of the last significant coefficient is checked for both the Cb transform block and the Cr transform block. In the case of a single tree, the position of the last significant coefficient can be checked for the transform blocks of luma, Cb, and Cr.
[0340] In Table 3, MtsZeroOutSigCoeffFlag is initially set to 1, and this value may be changed in the residual coding of Table 5. If there are significant coefficients in the region filled with zeros due to clearing (LastSignificantCoeffX>15||LastSignificantCoeffY>15), the variable MtsZeroOutSigCoeffFlag changes from 1 to 0. In this case, as shown in Table 3, the MTS index is not signaled.
[0341] In addition, as shown in Table 3, when tu_cbf_luma[x0][y0] is 0, mts_idx[x0][y0] encoding can be omitted. That is, if the CBF value of the luma component is 0, since no transform is applied, there is no need to signal the MTS index, so that the MTS index encoding can be omitted.
[0342] According to an example, the technical feature can be implemented using another conditional syntax. For example, after performing MTS, a variable indicating whether a significant coefficient exists in an area other than the DC area of the current block can be derived. If the variable indicates that a significant coefficient exists in an area other than the DC area, an MTS index can be signaled. That is, the presence of a significant coefficient in an area other than the DC area of the current block indicates that the value of tu_cbf_luma[x0][y0] is 1, and in this case, an MTS index can be signaled.
[0343] This variable may be expressed as MtsDcOnly, and after the variable MtsDcOnly is initially set to 1 at the coding unit level, the value may be changed to 0 when the residual coding level indicates that there are significant coefficients in an area other than the DC area of the current block. When the variable MtsDcOnly is 0, the image information may be configured such that the MTS index is signaled.
[0344] If tu_cbf_luma[x0][y0] is 0, the variable MtsDcOnly maintains the initial value 1 because the residual coding syntax is not called at the transform unit level in Table 4. In this case, since the variable MtsDcOnly is not changed to 0, the image information may be configured so that the MTS index is not signaled. In other words, the MTS index is not parsed and signaled.
[0345] In addition, the decoding device may determine the color index cIdx of the transformation coefficient to derive the variable MtsZeroOutSigCoeffFlag of Table 5. The color index cIdx being 0 indicates a luma component.
[0346] According to an example, since MTS may be applied only to the luma component of the current block, the decoding apparatus may determine whether the color index is luma when deriving the variable MtsZeroOutSigCoeffFlag that determines whether to parse the MTS index.
[0347] The variable MtsZeroOutSigCoeffFlag is a variable indicating whether clearing is performed when MTS is applied. It indicates whether there is a transform coefficient in an area other than the upper left area where the last significant coefficient may exist due to clearing after MTS is performed (that is, in an area other than the upper left 16×16 area). As shown in Table 3, the variable MtsZeroOutSigCoeffFlag is initially set to 1 (MtsZeroOutSigCoeffFlag=1) at the coding unit level, and if there is a transform coefficient in an area other than the 16×16 area, the value is changed from 1 to 0 at the residual encoding, as shown in Table 5. It can be changed (MtsZeroOutSigCoeffFlag=0). If the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.
[0348] As shown in Table 5, at the residual coding level, a non-cleared area where non-zero transform coefficients may exist may be set depending on whether clearing of the accompanying MTS is performed. Even in this case, when the color index (cIdx) is 0, the non-cleared area may be set to the upper left 16×16 area of the current block.
[0349] In this way, when deriving the variable for determining whether to parse the MTS index, whether the color component is luma or chroma is determined. However, since LFNST can be applied to both the luma component and the chroma component of the current block, the color component is not determined when deriving the variable for determining whether to parse the LFNST index.
[0350] For example, Table 3 shows the variable LfnstZeroOutSigCoeffFlag, which can indicate whether zeroing is performed when LFNST is applied. The variable LfnstZeroOutSigCoeffFlag indicates whether a significant coefficient exists in the second region of the current block, excluding the first region located in the upper left. This value is initially set to 1, and if a significant coefficient exists in the second region, this value can be changed to 0. The LFNST index can be parsed only when the value of the initially set variable LfnstZeroOutSigCoeffFlag remains 1. When determining and deriving whether the variable LfnstZeroOutSigCoeffFlag value is 1, the color index of the current block is not determined because LFNST can be applied to both the luma component and the chroma component of the current block.
[0351] In addition, a syntax table for signaling a coding unit of an LFNST index according to an example is as follows.
[0352] [Table 6]
[0353]
[0354] In Table 6, lfnst_idx represents the LFNST index and can have values 0, 1, and 2, as described above. As shown in Table 6, lfnst_idx is signaled only when the condition (!intra_mip_flag[x0][y0]||Min(lfnstWidth, lfnstHeight)>=16) is satisfied. Here, intra_mip_flag[x0][y0] is a flag indicating whether the matrix-based intra prediction (MIP) mode is applied to the luma block to which the (x0, y0) coordinates belong. If the MIP mode is applied to the luma block, the value is 1, and if not, the value is 0.
[0355] lfnstWidth and lfnstHeight indicate the width and height of the LFNST applied to the coding block currently being encoded (including both the luma coding block and the chroma coding block). When ISP is applied to a coding block, it can indicate the width and height of each partition block divided into two or four.
[0356] In addition, in the above conditions, when Min (lfnstWidth, lfnstHeight)> = 16 is equal to or greater than a 16 × 16 block when MIP is applied (for example, the width and height of the luma coding block to which MIP is applied are both equal to or greater than 16), it indicates that LFNST can be applied. The following briefly describes the meaning of the main variables included in Table 6 that are not repeated in the description of Table 4.
[0357] 1. IntraSubPartitionsSplitType: It indicates how the ISP partition is formed for the current coding unit, and ISP_NO_SPLIT indicates that the corresponding coding unit is not a coding unit split into sub-blocks. ISP_VER_SPLIT indicates vertical splitting, while ISP_HOR_SPLIT indicates horizontal splitting. For example, when a W×H (width W, height H) block is split horizontally into n partition blocks, it is split into W×(H / n) blocks, and when a W×H (width W, height H) block is split vertically into n partition blocks, it is split into (W / n)×H blocks.
[0358] 2. SubWidthC, SubHeightC: SubWidthC and SubHeightC are values set according to the color format (or chroma format, such as 4:2:0, 4:2:2, 4:4:4), and more specifically, they indicate the ratio of the width and height of the luma component and chroma component, respectively. (See the table below)
[0359] [Table 7]
[0360] Chroma format SubWidthC SubHeightC monochrome 1 1 4:2:0 2 2 4:2:2 2 1 4:4:4 1 1 4:4:4 1 1
[0361] 3. NumIntraSubPartitions: This indicates how many partition blocks are divided when ISP is applied. In other words, it indicates that the partition is divided into NumIntraSubPartitions partition blocks.
[0362] 4. LfnstDcOnly: For all transform blocks belonging to the current coding unit, the value of the LnfstDCOnly variable becomes 1 when each last non-zero coefficient position is the DC position (i.e., the upper left position within the corresponding transform block) or when there is no valid coefficient (i.e., when the corresponding CBF value is 0).
[0363] In the case of a luma single tree or luma dual tree, the LfnstDcOnly variable value is determined by checking the condition only for the transform block corresponding to the luma component in the corresponding coding unit, and in the case of a chroma single tree or chroma dual tree, the LfnstDcOnly variable value can be determined by checking the condition only for the transform block corresponding to the chroma component (Cb, Cr) in the corresponding coding unit. In the case of a single tree, the LfnstDcOnly variable value can be determined by checking the above condition for all transform blocks corresponding to the luma component and the chroma component (Cb, Cr) in the corresponding coding unit.
[0364] 5. LfnstZeroOutSigCoeffFlag: When LFNST is applied, if significant coefficients exist only in regions where significant coefficients can exist, it may be set to 1; otherwise, it may be set to 0.
[0365] In the case of a 4×4 transform block or an 8×8 transform block, up to 8 significant coefficients may be located starting from the (0,0) position (upper left) in the corresponding transform block according to the scanning order, and the remaining positions in the corresponding transform block may be cleared to 0. In the case of a transform block that is not 4×4 and 8×8 and whose width and height are equal to or greater than 4, respectively (i.e., a transform block to which LFNST may be applied), 16 significant coefficients may be located starting from the (0,0) position (upper left) in the corresponding transform block according to the scanning order (i.e., the significant coefficients may be located only in the upper left 4×4 block), and the remaining positions in the corresponding transform block may be cleared to 0.
[0366] In addition, as shown in Table 6, when encoded in a partition tree or dual tree, it is not checked whether MIP is applied to chroma components when signaling the LFNST index. In this way, LFNST can be appropriately applied to chroma components.
[0367] As shown in Table 6, the LFNST index is signaled when the condition (treeType == DUAL_TREE_CHROMA || !intra_mip_flag[x0][y0] || Min(lfnstWidth, lfnstHeight) > 16) is satisfied. This means that the LFNST index is signaled when the tree type is a dual-tree chroma type (treeType == DUAL_TREE_CHROMA), the MIP mode is not applied (!intra_mip_flag[x0][y0]), or the smaller of the width and height of the block to which LFNST is applied is 16 or greater (Min(lfnstWidth, lfnstHeight) > 16). That is, when the coding block is dual-tree chroma, the LFNST index is signaled without determining whether the MIP mode is applied or the width and height of the block to which LFNST is applied.
[0368] Additionally, the above condition can be interpreted as signaling the LFNST index without determining the width and height of the block to which LFNST is applied if the coding block is not dual-tree chroma and MIP is not applied.
[0369] Additionally, when the coding block is not dual-tree chroma and MIP is applied, it can be interpreted that the LFNST index can be signaled when the smaller of the width and height of the block to which LFNST is applied is 16 or greater.
[0370] On the other hand, the LFNST index is signaled only when transform skipping is not applied to the luma component as shown in Table 6 (that is, when the condition transform_skip_flag[x0][y0][0] == 0 is satisfied).
[0371] Here, x0 and y0 represent coordinates (x0, y0) when the upper left position in the picture is (0, 0) for the luma component and the horizontal X coordinate increases from left to right and the vertical Y coordinate increases from top to bottom.
[0372] (x0, y0) is a coordinate based on the luma component, but can also be used for chroma and phase components. In this case, the actual position indicated by the (x0, y0) coordinate can be scaled based on the picture for the chroma component. For example, when the chroma format is 4:2:0, the actual position of the chroma component indicated by (x0, y0) on the picture can be (x0 / 2, y0 / 2). For example, when the chroma format is 4:2:0, the actual position of the chroma component indicated by (x0, y0) on the picture can be (x0 / 2, y0 / 2).
[0373] In transform_skip_flag[x0][y0][0], the last index 0 refers to the luma component. More specifically, in transform_skip_flag[x0][y0][cIdx], cIdx refers to the component it is for, and if the cIdx value is 0, the cIdx value of 0 indicates luma, while cIdx greater than 0 (1 or 2) indicates chroma.
[0374] In addition, the variable LfnstDcOnly is initialized to a value of 1 as shown in Table 6, and can be set to 0 according to a condition in the parsing function for residual coding, as shown in the following table.
[0375] [Table 8]
[0376]
[0377]
[0378] As shown in Table 8, the LfnstDcOnly value can be set to 0 only when the transform_skip_flag[x0][y0][cIdx] value is 0 (that is, only when transform skipping is not applied to the component indicated by cIdx). If it is not ISP mode, as shown in Table 6, the LFNST index is signaled only when the LfnstDcOnly value is 0, and when the LFNST index is not signaled, the LFNST index value can be inferred to be 0.
[0379] For reference, the residual coding function presented in Table 8 is called when the transform_tree called in Table 6 is executed, and for a single tree, the residual coding functions for luma (cIdx=0) and chroma (cIdx=1 or 2, corresponding to the Cb component and the Cr component) are all called, and for a dual tree, in the case of the luma dual tree (DUAL_TREE_LUMA), only the residual coding function of luma (cIdx=0) is called, and in the case of the chroma dual tree (DUAL_TREE_CHROMA), only the residual coding function of chroma (cIdx=1 or 2, corresponding to the Cb and Cr components) is called.
[0380] The conditions for signaling the LFNST index for the case not in ISP mode are summarized as follows (here, it can be assumed that other conditions for signaling the LFNST index are met, for example, assuming that the condition Max(cbWidth, cbHeight)<=MaxTbSizeY is met).
[0381] 1. When transform_skip_flag[x0][y0][0] is 1
[0382] - LFNST index is inferred to be 0 and not signaled
[0383] 2. When transform_skip_flag[x0][y0][0] is 0
[0384] 2-A. When transform_skip_flag[x0][y0][1] is 0 and transform_skip_flag[x0][y0][2] is 0
[0385] - In Table 8, for all cIdx (for cIdx 0, 1, 2), the LfnstDcOnly value may be set to 0
[0386] If the LfnstDcOnly value is 0, the LFNST index is signaled; otherwise, the LFNST index is not signaled and the value is inferred to be 0.
[0387] 2-B. When transform_skip_flag[x0][y0][1] is 0 and transform_skip_flag[x0][y0][2] is 1
[0388] - In Table 8, the LfnstDcOnly value can be set to 0 only when cIdx is 0 and 1
[0389] If the LfnstDcOnly value is 0, the LFNST index is signaled; otherwise, the LFNST index is not signaled and the value is inferred to be 0.
[0390] 2-C. When transform_skip_flag[x0][y0][1] is 1 and transform_skip_flag[x0][y0][2] is 0
[0391] - In Table 8, the LfnstDcOnly value can be set to 0 only when cIdx is 0 and 2
[0392] If the LfnstDcOnly value is 0, the LFNST index is signaled; otherwise, the LFNST index is not signaled and the value is inferred to be 0.
[0393] 2-D. When transform_skip_flag[x0][y0][1] is 1 and transform_skip_flag[x0][y0][2] is 1
[0394] - In Table 8, the LfnstDcOnly value can be set to 0 only when cIdx is 0
[0395] If the LfnstDcOnly value is 0, the LFNST index is signaled; otherwise, the LFNST index is not signaled and the value is inferred to be 0.
[0396] In the case of a single tree, the values of transform_skip_flag[x0][y0][0], transform_skip_flag[x0][y0][1], transform_skip_flag[x0][y0][2] are checked for the above cases, in the case of a luma dual tree, only transform_skip_flag[x0][y0][0] is checked, and in the case of a chroma dual tree, the values of transform_skip_flag[x0][y0][1] and transform_skip_flag[x0][y0][2] are checked.
[0397] In case of ISP mode (IntraSubPartitionsSplitType!=ISP_NO_SPLIT condition in Table 6, ie, horizontal split or vertical split), as shown in Table 6, the LfnstDcOnly variable is not checked and the LFNST index is signaled.
[0398] Therefore, in the case of ISP mode under luma dual-tree and single-tree, regardless of the value of the LfnstDcOnly variable, when the transform_skip_flag[x0][y0][0] value is 0 (transform skipping is not applied for the luma component), the LFNST index is signaled (when the LFNST index is not signaled, the LFNST index value can be inferred to be 0).
[0399] In the case of chroma dual tree, based on the fact that ISP prediction is applied only to luma in the current VVC standard, it is considered that ISP is not applied to chroma, and the LFNST index can be signaled by checking the LfnstDcOnly variable in the above method, and as shown in Table 8, the LfnstDcOnly variable can be set to 0 only when the transform_skip_flag[x0][y0][cIdx] value is 0.
[0400] Of course, the application of ISP mode to luma affects even the chroma dual tree, so even in the case of chroma dual tree, regardless of the LfnstDcOnly variable, the LFNST index is signaled when the transform_skip_flag[x0][y0][0] value is 0.
[0401] The conditions for signaling the LFNST index when the ISP mode is applied and the transform_skip_flag[x0][y0][0] value is 0 are summarized as follows. If the transform_skip_flag[x0][y0][0] value is 1, the LFNST index is not signaled and the LFNST index is inferred to be 0. Of course, other conditions required for signaling the LFNST index in Table 6 may be assumed to be satisfied. For example, it may be assumed that the condition such as Max(cbWidth, cbHeight)<=MaxTbSizeY is satisfied.
[0402] 1. In the case of a single tree
[0403] - Signals LFNST index regardless of LfnstDcOnly variable value
[0404] 2. In the case of double trees
[0405] 2-A. In the case of a double brightness tree
[0406] - Signals LFNST index regardless of LfnstDcOnly variable value
[0407] 2-B. In the case of chroma dual trees
[0408] - Depending on the values of transform_skip_flag[x0][y0][1] and transform_skip_flag[x0][y0][2], when the LfnstDcOnly variable value is set to 0, that is, the cIdx value in transform_skip_flag[x0][y0][cIdx] is 1, the LfnstDcOnly variable value can be set to 0 only when the transform_skip_flag[x0][y0][1] value is 0, and when the cIdx value is 2, the LfnstDcOnly variable value can be set to 0 only when the transform_skip_flag[x0][y0][2] value is 0.
[0409] If the LfnstDcOnly value is 0, the LFNST index is signaled; otherwise, the LFNST index is not signaled and the value is inferred to be 0.
[0410] In the above case, the case of chroma dual tree is the same as the case where ISP is not applied.
[0411] Furthermore, according to an example, although transform skipping is allowed for chroma components in the current VVC standard, a transform skip flag corresponding to each chroma component is added as shown in the following table.
[0412] [Table 9]
[0413]
[0414] In Table 9, except for the case of luma dual tree, it can be confirmed that transform_skip_flag[xC][yC][1] corresponding to whether transform skip is applied to Cb and transform_skip_flag[xC][yC][2] corresponding to whether transform skip is applied to Cr can be signaled. If the transform_skip_flag[xC][yC][1] value is 1, transform skip is applied to Cb, and if it is 0, transform skip is not applied to Cb, and if the transform_skip_flag[xC][yC][2] value is 1, transform skip is applied to Cr, and if it is 0, transform skip is not applied to Cr.
[0415] Therefore, even if the LFNST index value is greater than 0 (i.e., when LFNST is applied), each transform_skip_flag[x0][y0][cIdx] value for the luma component (Y component) and chroma components (Cb component and Cr component) can be different. According to Table 6, since the LFNST index value can be greater than 0 only when the transform_skip_flag[x0][y0][0] value is 0, when the LFNST index value is greater than 0, transform_skip_flag[x0][y0][0] is always zero.
[0416] Therefore, the cases where LFNST can be applied according to the transform_skip_flag[x0][y0][cIdx] value are summarized as follows. Here, the LFNST index is greater than 0 and the transform_skip_flag[x0][y0][0] value is 0. It can be assumed that other conditions for applying LFNST are met, for example, the width and height of the corresponding block can be greater than or equal to 4.
[0417] 1. Single tree
[0418] -LFNST is applied to the luma component
[0419] - If the transform_skip_flag[x0][y0][1] value is 0, LFNST is applied to the Cb component, and if it is 1, LFNST is not applied to the Cb component.
[0420] - If the transform_skip_flag[x0][y0][2] value is 0, LFNST is applied to the Cr component, and if it is 1, LFNST is not applied to the Cr component.
[0421] 2. Brightness double tree
[0422] -LFNST is applied to the luma component
[0423] 3. Chroma Dual Tree
[0424] - If the transform_skip_flag[x0][y0][1] value is 0, LFNST is applied to the Cb component, and if it is 1, LFNST is not applied to the Cb component.
[0425] - If the transform_skip_flag[x0][y0][2] value is 0, LFNST is applied to the Cr component, and if it is 1, LFNST is not applied to the Cr component.
[0426] As described above, in order to selectively apply LFNST according to the transform_skip_flag[x0][y0][cIdx] value, the following condition should be added to the specification text of LFNST.
[0427] [Table 10]
[0428]
[0429]
[0430]
[0431] As shown in Table 10, when the LFNST index (lfnst_idx) value is not 0 (ie, when LFNST is applied), by checking the transform_skip_flag[xTbY][yTbY][cIdx] value of the component specified by cIdx (when lfnst_idx is not equal to 0 and transform_skip_flag[xTbY][yTbY][cIdx] is equal to 0 and both nTbW and nTbH are greater than or equal to 4, the following applies), it can be configured so that the subsequent encoding process is performed only when the transform_skip_flag[xTbY][yTbY][cIdx] value is 0 (ie, LFNST is applied).
[0432] Furthermore, the LFNST index may be signaled according to whether transform is skipped for each color component.
[0433] As an example, when compared with Table 6, in Table 11, only when the transform_skip_flag[x0][y0][0] value is 0, the condition for transmitting the LFNST index may be removed.
[0434] [Table 11]
[0435]
[0436] However, since the method of setting the LfnstDcOnly variable value described in Table 11 is the same as that in Table 8, the setting of the LfnstDcOnly variable value is changed according to the transform_skip_flag[x0][y0][cIdx] value, and finally whether the LFNST index is signaled will also change.
[0437] When ISP mode is not applied, how to signal the LFNST index through the transform_skip_flag[x0][y0][cIdx] value is summarized as follows. It can be assumed that other conditions for signaling the LFNST index have been met, for example, conditions such as Max(cbWidth, cbHeight)<=MaxTbSizeY are met.
[0438] 1. In the case of a single tree
[0439] -When the transform_skip_flag[x0][y0][0] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 8.
[0440] -When the transform_skip_flag[x0][y0][1] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 8.
[0441] -When the transform_skip_flag[x0][y0][2] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 8.
[0442] - The LFNST index may be signaled when the LfnstDcOnly value is 0. If the LFNST index is not signaled, it may be inferred to be 0.
[0443] 2. In the case of dual-tree of luminance component
[0444] -When the transform_skip_flag[x0][y0][0] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 8.
[0445] - The LFNST index may be signaled when the LfnstDcOnly value is 0. If the LFNST index is not signaled, it may be inferred to be 0.
[0446] 3. In the case of chroma component dual tree
[0447] -When the transform_skip_flag[x0][y0][1] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 8.
[0448] -When the transform_skip_flag[x0][y0][2] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 8.
[0449] - The LFNST index may be signaled when the LfnstDcOnly value is 0. If the LFNST index is not signaled, it may be inferred to be 0.
[0450] As shown in Table 11, the LfnstDcOnly value is initialized to 1, and in the case of dual trees, the LFNST index corresponding to the luma dual tree and the LFNST index corresponding to the chroma dual tree can be signaled separately. This means that different LFNST kernels can be applied to luma and chroma.
[0451] In addition, the dual trees shown in Tables 6 to 11 may include DUAL_TREE_LUMA (corresponding to the luma component) and DUAL_TREE_CHROMA (corresponding to the chroma component) appearing in the current VVC specification document, which may include a syntax parsing tree for luma and a syntax parsing tree for chroma that are different due to the size condition of the coding unit, etc. For example, a single tree case may be included.
[0452] When the ISP mode is applied, transform_skip_flag[x0][y0][0] is not signaled and is inferred to be 0 as shown in Table 9. That is, as shown in Table 9, transform_skip_flag[x0][y0]][0] is signaled only when the condition IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT is satisfied, which is the case when the ISP mode is not applied.
[0453] In addition, as shown in Table 9, whether the ISP mode is applied or not, transform_skip_flag[x0][y0][1] and transform_skip_flag[x0][y0][2] may be signaled.
[0454] Therefore, the LFNST signaling in the case of applying the ISP mode can be summarized as follows: It can be assumed that other conditions for sending the LFNST index have been met, for example, conditions such as Max(cbWidth, cbHeight)<=MaxTbSizeY can be met.
[0455] 1. In the case of a single tree
[0456] - The LFNST index may be signaled. If no LFNST index is signaled, it may be inferred to be 0.
[0457] 2. In the case of dual-tree of luminance component
[0458] - The LFNST index may be signaled. If no LFNST index is signaled, it may be inferred to be 0.
[0459] 3. In the case of chroma component dual tree
[0460] -When the transform_skip_flag[x0][y0][1] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 8.
[0461] -When the transform_skip_flag[x0][y0][2] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 8.
[0462] - The LFNST index may be signaled when the LfnstDcOnly value is 0. If the LFNST index is not signaled, it may be inferred to be 0.
[0463] When the ISP mode is applied, the LfnstDcOnly condition is not checked, as shown in Table 11. Therefore, in the case of the first single tree and the second luma component dual tree, the LFNST index can be signaled without checking the LfnstDcOnly condition. In the case of the chroma dual tree, the LFNST index can be signaled according to the same conditions as in the case where the ISP mode is not applied in Case 3 above. That is, the LFNST index can be signaled according to the LfnstDcOnly condition.
[0464] According to an example, since the transform_skip_flag[x0][y0][cIdx] value can be assigned to the luma component and the two chroma components respectively, when the LFNST index value is greater than 0, that is, even when LFNST is applied, LFNST can be applied to the component indicated by cIdx only when the transform_skip_flag[x0][y0][cIdx] value is 0. The correspondingly changed normative text content is the same as Table 10.
[0465] According to an example, if the condition of checking whether the transform_skip_flag[x0][y0][0] value is 0 only for the dual tree case compared to Table 6 is removed, the LFNST index signaling may be configured as shown in Table 12. The LfnstDcOnly variable shown in Table 12 may be set to 0 according to the condition shown in Table 8.
[0466] [Table 12]
[0467]
[0468] When configured as shown in Table 12, in the case of a single tree, the LFNST index signaling method shown in Tables 6 to 8 can be applied, and in the case of a dual tree, the method shown in Table 11 can be applied. In addition, since the transform_skip_flag[x0][y0][cIdx] values are assigned to the luma component and the two chroma components, respectively, as shown in Tables 6 to 8, even if the LFNST index value is greater than 0 (i.e., when LFNST is applied), it can be configured so that LFNST is applied to the component indicated by cIdx only when the transform_skip_flag[x0][y0][cIdx] value is 0. The correspondingly changed normative text content is the same as that in Table 10.
[0469] The following figures are created to illustrate specific examples of this specification. Since the names of specific devices or the names of specific signals / messages / fields described in the figures are presented as examples, the technical features of this specification are not limited to the specific names used in the following figures.
[0470] Figure 12 is a flowchart illustrating the operation of a video decoding apparatus according to an embodiment of the present disclosure.
[0471] Figure 12 Each step disclosed in the above is based on Figures 1 to 11 Therefore, some of the contents described in the above will be omitted or simplified. Figures 1 to 11 A detailed description of what is described in .
[0472] The decoding apparatus 200 according to an embodiment may receive information about an intra prediction mode, residual information, and an LFNST index from a bitstream ( S1210 ).
[0473] More specifically, the decoding device 200 can decode information about the quantized transform coefficients of the current block from the bitstream and derive the quantized transform coefficients of the target block based on the information about the quantized transform coefficients of the current block. The information about the quantized transform coefficients of the target block may be included in a sequence parameter set (SPS) or a slice header, and may include at least one of information about a simplification factor, information about a minimum transform size to which a simplified transform is applied, information about a maximum transform size to which a simplified transform is applied, and information about a transform index indicating any one of a transform kernel matrix and a simplified inverse transform size included in a transform set.
[0474] In addition, the decoding device may also receive information about the intra-frame prediction mode of the current block and information about whether ISP is applied to the current block. The decoding device can derive whether the current block is divided into a predetermined number of sub-partition transform blocks by receiving and parsing flag information indicating whether ISP encoding or ISP mode is applied. Here, the current block may be a coding block. In addition, the decoding device can derive the size and number of the sub-partition blocks after division by using flag information indicating the direction in which the current block is to be divided.
[0475] The decoding apparatus 200 may induce a transform coefficient by performing inverse quantization on residual information (ie, a quantized transform coefficient) about the current block ( S1220 ).
[0476] The derived transform coefficients can be arranged in reverse diagonal scanning order in units of 4×4 blocks, and the transform coefficients within the 4×4 block can also be arranged in reverse diagonal scanning order. That is, the transform coefficients that have been inversely quantized can be arranged according to the reverse scanning order used in a video codec (such as in VVC or HEVC).
[0477] The decoding apparatus may derive modified transform coefficients by applying LFNST to the transform coefficients.
[0478] Unlike the first transform, which separates and transforms transform target coefficients in the vertical or horizontal direction, LFNST is a non-separable transform that applies the transform without separating the coefficients in a specific direction. This non-separable transform can be a low-frequency non-separable transform that applies the forward transform only to the low-frequency region rather than the entire block region.
[0479] The LFNST index information may be received as syntax information, and the syntax information may be received as a binarized bin string including 0 and 1.
[0480] The syntax element of the LFNST index according to this embodiment can indicate whether inverse LFNST or inverse non-separate transform is applied and any one of the transform kernel matrices included in the transform set, and when two transform kernel matrices are included in the transform set, the syntax element of the transform index can have three values.
[0481] That is, according to an embodiment, the syntax element value of the LFNST index may include: 0, which indicates that inverse LFNST is not applied to the target block; 1, which indicates the first transform kernel matrix among the transform kernel matrices; and 2, which indicates the second transform kernel matrix among the transform kernel matrices.
[0482] Intra prediction mode information and LFNST index information may be signaled at a coding unit level.
[0483] The decoding apparatus may parse the LFNST index of the current block based on the respective transform skip flag values of the color components of the current block ( S1230 ).
[0484] The decoding device may derive a variable (DC significant coefficient variable) indicating whether a significant coefficient exists in the DC component of the current block based on the transform skip flag value, and may derive the variable based on the respective transform skip flags of the color components of the current block. According to an example, the DC significant coefficient variable may be set to 0 based on the fact that at least one respective transform skip flag value is 0, and the LFNST index may be parsed based on the DC significant coefficient variable being 0.
[0485] A variable indicating whether there is a significant coefficient in the DC component of the current block may be represented as a variable LfnstDcOnly, and becomes 0 when there is a non-zero coefficient at a non-DC component for at least one transform block in one coding unit, and becomes 1 when there is no non-zero coefficient at a position other than the DC component for all transform blocks in one coding unit. In the present disclosure, the DC component refers to the upper left position or (0, 0) which is a position reference of a 2D component.
[0486] There may be several transform blocks within a coding unit. For example, in the case of chrominance components, there may be transform blocks for Cb and Cr, and in the case of a single tree type, there may be transform blocks for luma, Cb, and Cr. According to an example, when a non-zero coefficient other than the DC component position is found in even one transform block among the transform blocks constituting the current coding block, the value of the variable LnfstDcOnly may be set to 0.
[0487] Furthermore, since residual coding is not performed on a transform block if there are no non-zero coefficients in the corresponding transform block, the value of the variable LfnstDcOnly is not changed by the corresponding transform block. Therefore, if there are no non-zero coefficients in the non-DC component of the transform block, the value of the variable LfnstDcOnly remains unchanged and maintains its previous value. For example, when a coding unit is encoded using a single-tree type and the value of the variable LfnstDcOnly changes to 0 due to the luma transform block, the value of the variable LfnstDcOnly remains 0 if non-zero coefficients exist only in the DC component of the Cb transform block or if there are no non-zero coefficients in the Cb transform block. The value of the variable LfnstDcOnly is initially initialized to 1. If no component in the current coding unit updates the value of the variable LfnstDcOnly to 0, it remains 1. If one of the transform blocks constituting the coding unit sets the value of the variable LfnstDcOnly to 0, it ultimately remains 0.
[0488] Furthermore, the variable LfnstDcOnly can be derived based on the respective transform skip flag values of the color components of the current block. The transform skip flag of the current block can be signaled for each color component, and if the tree type of the current block is single-tree, the transform skip flag value of the luma component, the transform skip flag value of the chroma Cb component, and the transform skip flag value of the chroma Cr component can be derived based on the transform skip flag value of the luma component. Alternatively, if the tree type of the current block is dual-tree luma, the variable LfnstDcOnly can be derived based on the transform skip flag value of the luma component, and if the tree type of the current block is dual-tree chroma, the variable LfnstDcOnly can be derived based on the transform skip flag value of the chroma Cb component and the transform skip flag value of the chroma Cr component.
[0489] According to an example, based on the transform skip flag value of the color component being 0, the variable LfnstDcOnly may indicate the presence of a significant coefficient at a position other than the DC component. That is, if the tree type of the current block is single-tree, the transform skip flag value of the luma component is 0, then the variable LfnstDcOnly may be derived to be 0 based on the fact that at least one of the transform skip flag value of the luma component, the transform skip flag value of the chroma Cb component, and the transform skip flag value of the chroma Cr component is 0. Alternatively, if the tree type of the current block is dual-tree luma, the variable LfnstDcOnly may be derived based on the transform skip flag value of the luma component, and if the tree type of the current block is dual-tree chroma, the variable LfnstDcOnly may be derived based on the transform skip flag value of the chroma Cb component and the transform skip flag value of the chroma Cr component.
[0490] As described above, the variable LfnstDcOnly may be initially set to 1 at the coding unit level of the current block, and if the transform skip flag value is 0, the variable LfnstDcOnly may be changed to 0 at the residual coding level.
[0491] The decoding apparatus may parse the LFNST index based on the variable LfnstDcOnly indicating that a significant coefficient exists at a position other than a DC component, ie, the variable LfnstDcOnly is 0.
[0492] Furthermore, in the case of a luma block to which an intra subpartitioning (ISP) mode may be applied, the LFNST index may be parsed without deriving the variable LfnstDcOnly.
[0493] Specifically, when the ISP mode is applied and the transform skip flag of the luma component (i.e., transform_skip_flag[x0][y0][0] value) is 0, and the tree type of the current block is a single tree or a luma dual tree, the LFNST index can be signaled regardless of the variable LfnstDcOnly value.
[0494] On the other hand, in the case of a chroma component to which the ISP mode is not applied, the variable LfnstDcOnly value may be set to 0 according to transform_skip_flag[x0][y0][1] as the transform skip flag of the chroma Cb component and transform_skip_flag[x0][y0][2] as the transform skip flag of the chroma Cr component. That is, in transform_skip_flag[x0][y0][cIdx], when the cIdx value is 1, the variable LfnstDcOnly value may be set to 0 only when the transform_skip_flag[x0][y0][1] value is 0, and when the cIdx value is 2, the transform_skip_flag[x0][y0][2] value may be set to 0 only when the transform_skip_flag[x0][y0][2] value is 0. If the variable LfnstDcOnly value is 0, the decoding device may parse the LFNST index, otherwise the LFNST index may not be signaled and may be inferred to be a value of 0.
[0495] Thereafter, the decoding apparatus may derive a modified transform coefficient from the transform coefficient based on the LFNST index and the LFNST matrix used for LFNST ( S1240 ).
[0496] The decoding device may set a plurality of variables for LFNST based on whether the LFNST index is not 0 (ie, whether the LFNST index is greater than 0) and respective transform skip flag values of the color components are 0.
[0497] For example, in the step of applying LFNST after parsing the LFNST index, the decoding device may again determine whether the transform skip flag value of each color component is 0, and may set various variables for applying LFNST. For example, the intra prediction mode for selecting the LFNST set, the number of transform coefficients output after applying LFNST, the size of the block to which LFNST is applied, etc. may be set.
[0498] In the case of blocks encoded with BDPCM, the transform skip flag may be automatically set to 1, and in this case, the transform skip flag may be 1 even if the LFNST index is not 0, so when LFNST is actually applied, the transform skip flag value of each color component may be checked again.
[0499] Alternatively, according to an example, when the flag value indicating whether a coding significant coefficient exists in a transform block is 0, there may be a case where the transform skip flag value is not checked. In this case as well, since the transform skip flag value is not guaranteed to be 0 simply because the LFNST index is not 0, when LFNST is actually applied, the transform skip flag value of each color component may be checked again.
[0500] That is, the decoding device may check the transform skip flag value of each color component in the LFNST index parsing step, and may check the transform skip flag value of each color component again when LFNST is actually applied.
[0501] The decoding apparatus may determine an LFNST set including an LFNST matrix based on the intra prediction mode derived from the intra prediction mode information, and select any one of a plurality of LFNST matrices based on the LFNST set and the LFNST index.
[0502] In this case, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks split from the current block. That is, because the same intra prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra prediction mode can be applied equally to all sub-partition transform blocks. In addition, because the LFNST index is signaled at the coding unit level, the same LFNST matrix can be applied to the sub-partition transform blocks split from the current block.
[0503] As described above, a transform set can be determined according to the intra prediction mode of the transform block to be transformed, and inverse LFNST can be performed based on any one of the transform kernel matrices (i.e., LFNST matrices) included in the transform set indicated by the LFNST index. The matrix applied to inverse LFNST can be referred to as an inverse LFNST matrix or an LFNST matrix, and such a matrix can have any name as long as it has a transposed relationship with the matrix used for forward LFNST.
[0504] In one example, the inverse LFNST matrix may be a non-square matrix in which the number of columns is less than the number of rows.
[0505] The decoding apparatus may induce residual samples of the current block based on one inverse transform of the modified transform coefficient ( S1250 ).
[0506] In this case, as the inverse primary transform, a conventional separate transform can be used, and the above-mentioned MTS can be used.
[0507] Subsequently, the decoding apparatus 200 may generate reconstructed samples based on the residual samples of the current block and the predicted samples of the current block.
[0508] The following figures are created to describe specific examples of this specification. Since the names of specific devices or the names of specific signals / messages / fields shown in the figures are presented by way of example, the technical features of this specification are not limited to the specific names used in the following figures.
[0509] Figure 13 is a flowchart illustrating the operation of a video encoding apparatus according to an embodiment of the present disclosure.
[0510] Figure 13 Each step disclosed in the above is based on Figures 4 to 11 Therefore, some of the contents described in the above will be omitted or simplified. Figure 2 as well as Figures 4 to 11 A detailed description of what is described in .
[0511] The encoding apparatus 100 according to an embodiment may induce a prediction sample of a current block based on an intra prediction mode applied to the current block.
[0512] The encoding apparatus may perform prediction for each sub-partitioned transform block when applying the ISP to the current block.
[0513] The encoding device can determine whether to apply ISP encoding or ISP mode to the current block (i.e., the encoding block), and determine in which direction the current block is to be divided according to the determination result, and derive the size and number of the divided sub-blocks.
[0514] The same intra prediction mode is applied to the sub-partition transform blocks divided from the current block, and the encoding device can derive prediction samples for each sub-partition transform block. That is, the encoding device performs intra prediction in sequence according to the division form of the sub-partition transform blocks, such as horizontally or vertically, from left to right or from top to bottom. For the leftmost or topmost sub-block, the reconstructed pixels of the already encoded coding block are referenced in the traditional intra prediction method. In addition, for each edge of the subsequent internal sub-partition transform block, when it is not adjacent to the previous sub-partition transform block, in order to derive the reference pixels adjacent to the corresponding edge, the adjacent coding block that has been encoded as in the traditional intra prediction method refers to the reconstructed pixels.
[0515] The encoding apparatus 100 may induce residual samples of the current block based on the prediction samples ( S1310 ).
[0516] The encoding apparatus 100 may induce a transformation coefficient of a current block by applying at least one of LFNST and MTS to residual samples, and may arrange the transformation coefficients according to a predetermined scanning order.
[0517] The encoding apparatus may induce a transformation coefficient of the current block based on one transformation of the residual sample ( S1320 ).
[0518] One transform may be performed by a plurality of transform cores like MTS, and in this case, the transform core may be selected based on the intra prediction mode.
[0519] The encoding apparatus 100 may determine whether to perform a secondary transform or a non-separate transform (specifically, LFNST) on a transform coefficient of a current block, and apply LFNST to the transform coefficient to induce a modified transform coefficient.
[0520] Unlike the first transform that separates and transforms the transform target coefficients in the vertical or horizontal direction, LFNST is a non-separable transform that applies the transform without separating the coefficients in a specific direction. The non-separable transform may be a low-frequency non-separable transform that applies the transform only to the low-frequency region rather than the entire target block to be transformed.
[0521] The encoding device may apply multiple LFNST matrices to the transform coefficients to derive a variable (DC significant coefficient variable) indicating whether a significant coefficient exists in the DC component of the current block, and may derive the variable based on the respective transform skip flag values of the color components of the current block (S1330).
[0522] The encoding apparatus may induce variables after applying LFNST to each LFNST matrix candidate, or in a state where LFNST is not applied when LFNST is not applied.
[0523] Specifically, the encoding device may apply a plurality of LFNST candidates (i.e., LFNST matrices) to exclude corresponding LFNST matrices in which the significant coefficients of all transform blocks exist only in the DC position (of course, when the CBF is 0, the corresponding variable is excluded from the process), and compare RD values only between LFNST matrices in which the variable LfnstDcOnly value is 0. For example, when LFNST is not applied, it is included in the comparison process because it is irrelevant to the variable LfnstDcOnly value (in this case, since LFNST is not applied, the variable LfnstDcOnly value can be determined based on the transform coefficient obtained as a result of one transform), and the LFNST matrix corresponding to the LfnstDcOnly value of 0 is also included in the comparison process of the RD value.
[0524] A variable indicating whether there is a significant coefficient in the DC component of the current block can be represented as a variable LfnstDcOnly, and for at least one transform block in one coding unit, becomes 0 when there is a non-zero coefficient at a non-DC component, and for all transform blocks in one coding unit, becomes 1 when there is no non-zero coefficient in a position other than the DC component.
[0525] There may be several transform blocks in one coding unit. For example, in the case of chrominance components, there may be transform blocks for Cb and Cr, and in the case of a single tree type, there may be transform blocks for luma, Cb, and Cr. According to an example, when a non-zero coefficient other than the DC component position is found in even one transform block among the transform blocks constituting the current coding block, the variable LnfstDcOnly value may be set to 0.
[0526] Furthermore, since residual coding is not performed on a transform block if there are no non-zero coefficients in the corresponding transform block, the value of the variable LfnstDcOnly is not changed by the corresponding transform block. Therefore, if there are no non-zero coefficients in the non-DC component of the transform block, the value of the variable LfnstDcOnly does not change and maintains its previous value. For example, when a coding unit is encoded using a single-tree type and the value of the variable LfnstDcOnly is changed to 0 due to the luma transform block, the value of the variable LfnstDcOnly remains 0 if non-zero coefficients exist only in the DC component of the Cb transform block or if there are no non-zero coefficients in the Cb transform block. The value of the variable LfnstDcOnly is initially initialized to 1 and remains 1 if no component in the current coding unit updates the value of the variable LfnstDcOnly to 0. It ultimately remains 0 if one of the transform blocks constituting the coding unit sets the value of the variable LfnstDcOnly to 0.
[0527] Furthermore, the variable LfnstDcOnly can be derived based on the respective transform skip flag values of the color components of the current block. The transform skip flag of the current block can be signaled for each color component, and if the tree type of the current block is single-tree, the transform skip flag value of the luma component, the transform skip flag value of the chroma Cb component, and the transform skip flag value of the chroma Cr component can be derived based on the transform skip flag value of the luma component. Alternatively, if the tree type of the current block is dual-tree luma, the variable LfnstDcOnly is derived based on the transform skip flag value of the luma component, and if the tree type of the current block is dual-tree chroma, the variable LfnstDcOnly can be derived based on the transform skip flag value of the chroma Cb component and the transform skip flag value of the chroma Cr component.
[0528] According to an example, based on the transform skip flag value of the color component being 0, the variable LfnstDcOnly may indicate the presence of a significant coefficient at a position other than the DC component. That is, if the tree type of the current block is single-tree, the transform skip flag value of the luma component is 0, then the variable LfnstDcOnly may be derived to be 0 based on the fact that at least one of the transform skip flag value of the luma component, the transform skip flag value of the chroma Cb component, and the transform skip flag value of the chroma Cr component is 0. Alternatively, if the tree type of the current block is dual-tree luma, the variable LfnstDcOnly may be derived based on the transform skip flag value of the luma component, and if the tree type of the current block is dual-tree chroma, the variable LfnstDcOnly may be derived based on the transform skip flag value of the chroma Cb component and the transform skip flag value of the chroma Cr component.
[0529] As described above, the variable LfnstDcOnly may be initially set to 1 at the coding unit level of the current block, and if the transform skip flag value is 0, the variable LfnstDcOnly may be changed to 0 at the residual coding level.
[0530] The encoding apparatus may select an optimal LFNST matrix based on the variable indicating that a significant coefficient exists at a position other than a DC component, and may derive a modified transform coefficient based on the selected LFNST matrix ( S1340 ).
[0531] The encoding apparatus may set a plurality of variables for LFNST based on whether respective transform skip flag values of color components are 0 in the step of deriving the modified transform coefficient.
[0532] For example, after determining whether to apply LFNST, the encoding device may again determine whether the transform skip flag value of each color component is 0 in the step of applying LFNST, and may set various variables for applying LFNST. For example, the intra prediction mode for selecting the LFNST set, the number of transform coefficients output after applying LFNST, the size of the block to which LFNST is applied, etc. may be set.
[0533] In the case of a block encoded by BDPCM, since the transform skip flag may be automatically set to 1, when LFNST is actually applied, the transform skip flag value of each color component may be checked again.
[0534] Alternatively, according to an example, when the flag value indicating the presence of a significant coding coefficient in a transform block is 0, there may be a case where the transform skip flag value is not checked. In this case as well, since the transform skip flag value is not guaranteed to be 0 simply because the LFNST index is not 0, when LFNST is actually applied, the transform skip flag value of each color component may be checked again.
[0535] That is, the encoding device may check the transform skip flag value of each color component in the step of determining whether to apply LFNST, and may check the transform skip flag value of each color component again when actually applying LFNST.
[0536] Furthermore, in the case of a luma block to which an intra subpartitioning (ISP) mode may be applied, LFNST may be applied without deriving the variable LfnstDcOnly.
[0537] Specifically, when the ISP mode is applied and the transform skip flag for the luma component (i.e., transform_skip_flag[x0][y0][0] value) is 0, and the tree type of the current block is a single tree or a luma dual tree, LFNST can be applied regardless of the variable LfnstDcOnly value.
[0538] On the other hand, in the case of a chroma component to which the ISP mode is not applied, the variable LfnstDcOnly value may be set to 0 according to transform_skip_flag[x0][y0][1] as the transform skip flag for the chroma Cb component and transform_skip_flag[x0][y0][2] as the transform skip flag for the chroma Cr component. That is, in transform_skip_flag[x0][y0][cIdx], when the cIdx value is 1, the variable LfnstDcOnly value may be set to 0 only when the transform_skip_flag[x0][y0][1] value is 0, and when the cIdx value is 2, the transform_skip_flag[x0][y0][2] value may be set to 0 only when the transform_skip_flag[x0][y0][2] value is 0. If the variable LfnstDcOnly value is 0, the encoding device may apply LFNST, otherwise LFNST is not applied.
[0539] The encoding apparatus 100 may determine an LFNST set based on a mapping relationship according to an intra prediction mode applied to a current block, and perform LFNST, ie, inseparable transform, based on one of two LFNST matrices included in the LFNST set.
[0540] In this case, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks split from the current block. That is, because the same intra prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra prediction mode can also be applied equally to all sub-partition transform blocks. In addition, because the LFNST index is encoded in units of coding units, the same LFNST matrix can be applied to the sub-partition transform blocks split from the current block.
[0541] As described above, the transform set may be determined according to the intra prediction mode of the transform block to be transformed.The matrix applied to LFNST has a transposed relationship with the matrix used for inverse LFNST.
[0542] In one example, the LFNST matrix may be a non-square matrix in which the number of rows is less than the number of columns.
[0543] The encoding device can construct the image information so that the LFNST index indicating the LFNST matrix applied to LFNST is parsed based on the following facts: the variable LfnstDcOnly is initially set to 1 in the coding unit level of the current block, and when the transform skip flag value is 0, the variable LfnstDcOnly is changed to 0 at the residual coding level, and the variable LfnstDcOnly is 0.
[0544] The encoding apparatus may perform quantization based on the modified transform coefficient of the current block to induce a quantized transform coefficient, and encode the LFNST index.
[0545] That is, the encoding device can generate residual information including information about the quantized transform coefficients. The residual information can include the above-mentioned transform-related information / syntax elements. The encoding device can encode the image / video information including the residual information and output the encoded image / video information in the form of a bitstream.
[0546] More specifically, the encoding apparatus 100 may generate information about the quantized transform coefficient and encode the generated information about the quantized transform coefficient.
[0547] The syntax element of the LFNST index according to this embodiment may indicate whether (inverse) LFNST is applied and any one LFNST matrix included in the LFNST set, and when the LFNST set includes two transform kernel matrices, the syntax element of the LFNST index may have three values.
[0548] According to an embodiment, when the partition tree structure of the current block is a dual tree type, an LFNST index may be encoded for each of the luma block and the chroma block.
[0549] According to an embodiment, the syntax element value of the transform index may be derived as 0, 1, and 2, 0 indicating that (inverse) LFNST is not applied to the current block, 1 indicating the first LFNST matrix in the LFNST matrix, and 2 indicating the second LFNST matrix in the LFNST matrix.
[0550] In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency.
[0551] In addition, in the present disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled through residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.
[0552] In the above embodiments, the method is explained based on a flowchart with the aid of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may be performed in an order or step different from the above order or step, or a certain step may be performed concurrently with other steps. In addition, it will be understood by those skilled in the art that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0553] The above-mentioned method according to the present disclosure may be implemented in a software form, and the encoding device and / or decoding device according to the present disclosure may be included in a device for image processing such as a television, a computer, a smart phone, a set-top box, and a display device.
[0554] When the embodiments in the present disclosure are implemented by software, the above methods can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium and / or other storage devices. In other words, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units shown in each of the accompanying drawings can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip.
[0555] In addition, the decoding device and encoding device to which the present disclosure is applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device (such as video communication), a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0556] In addition, the processing method of the present invention can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network. In addition, the embodiments of the present invention can be implemented as a computer program product through program code, and the program code can be executed on a computer according to the embodiments of the present invention. The program code can be stored on a computer-readable carrier.
[0557] Figure 14An example of a video / image encoding system to which the present disclosure is applicable is schematically illustrated.
[0558] Reference Figure 14 The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may transmit the encoded video / image information or data to the receive device in the form of a file or stream via a digital storage medium or a network.
[0559] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0560] The video source can obtain the video / image by capturing, synthesizing, or generating a video / image. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.
[0561] An encoding device can encode input video / images. It can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0562] The transmitter can transmit the encoded video / image information or data, output as a bitstream, to a receiver in a receiving device via a digital storage medium or network in the form of a file or stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received / extracted bitstream to a decoding device.
[0563] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0564] The renderer can render the decoded video / image, and the rendered video / image can be displayed on a display.
[0565] Figure 15 The structure of the content streaming system to which the present disclosure is applied is illustrated.
[0566] Furthermore, a content streaming system to which the present disclosure is applied may generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0567] The encoding server is used to compress content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream, and then transmit the bitstream to the streaming server. As another example, if the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. The streaming server can also temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0568] The streaming server transmits multimedia data to user devices via a web server based on user requests. The web server serves as a tool for notifying users of available services. When a user requests a desired service, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case controls commands and responses between the various devices in the content streaming system.
[0569] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to smoothly provide a streaming service, the streaming server can store the bitstream for a predetermined time.
[0570] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system may operate as a distributed server, and in this case, data received by each server may be processed in a distributed manner.
[0571] The claims disclosed herein may be combined in various ways. For example, the technical features of the method claims of the present disclosure may be combined to be implemented or performed in a device, and the technical features of the device claims may be combined to be implemented or performed in a method. Furthermore, the technical features of method claims and device claims may be combined to be implemented or performed in a device, and the technical features of method claims and device claims may be combined to be implemented or performed in a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtain residual information from the bitstream; deriving a transform coefficient for a current block based on the residual information; deriving a DC significant coefficient variable indicating whether a significant coefficient exists only in a DC component of the current block based on a first transform skip flag for a chroma Cb component of the current block and a second transform skip flag for a chroma Cr component of the current block; Resolving LFNST indexes based on the DC significant coefficient variables; deriving modified transform coefficients by applying a low frequency non-separable transform (LFNST) to the transform coefficients; deriving residual samples for the current block based on an inverse primary transform of the modified transform coefficients; as well as Generate a reconstructed picture based on the residual samples, The step of deriving the modified transform coefficients includes: applying the LFNST to at least one of a first transform coefficient associated with the chroma Cb component or a second transform coefficient associated with the chroma Cr component, Wherein, whether to apply the LFNST to the chroma Cb component is determined based on the LFNST index and the first transform skip flag, Wherein, whether to apply the LFNST to the chroma Cr component is determined based on the LFNST index and the second transform skip flag, and The LFNST index is parsed based on that the tree type of the current block is dual-tree chroma.
2. The image decoding method according to claim 1, in, The value of the DC effective coefficient variable is initially set to 1, wherein, based on at least one of the value of the first transform skip flag and the value of the second transform skip flag being 0, the value of the DC significant coefficient variable is derived to be 0, and The LFNST index is parsed based on the value of the DC significant coefficient variable being 0.
3. The image decoding method according to claim 1, in, The value of the DC significant coefficient variable is initially set to 1 at the coding unit level of the current block, and Wherein, based on at least one of the value of the first transform skip flag and the value of the second transform skip flag being 0, the value of the DC significant coefficient variable is set to 0 at the residual coding level.
4. The image decoding method according to claim 1, in, Based on the value of the LFNST index being greater than 0 and the value of the first transform skip flag for the chroma Cb component being 0, applying the LFNST to the chroma Cb component of the current block, and Wherein, based on the value of the LFNST index being greater than 0 and the value of the first transform skip flag for the chroma Cb component being 1, the LFNST is not applied to the chroma Cb component of the current block.
5. The image decoding method according to claim 1, in, Based on the value of the LFNST index being greater than 0 and the value of the second transform skip flag for the chroma Cr component being 0, applying the LFNST to the chroma Cr component of the current block, and Wherein, based on the value of the LFNST index being greater than 0 and the value of the second transform skip flag for the chroma Cr component being 1, the LFNST is not applied to the chroma Cr component of the current block.
6. The image decoding method according to claim 1, in, Based on the value of the first transform skip flag for the chroma Cb component being 1 and the value of the second transform skip flag for the chroma Cr component being 0, the value of the DC significant coefficient variable is derived to be 0.
7. The image decoding method according to claim 1, in, Based on the value of the first transform skip flag for the chroma Cb component being 0 and the value of the second transform skip flag for the chroma Cr component being 1, the value of the DC significant coefficient variable is derived to be 0.
8. The image decoding method according to claim 1, in, Based on the value of one of the first transform skip flag for the chroma Cb component and the second transform skip flag for the chroma Cr component being 1 and the other being 0, the value of the DC significant coefficient variable is derived to be 0.
9. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Deriving prediction samples for the current block; deriving residual samples for the current block based on the prediction samples; deriving transform coefficients for the current block based on a transform for the residual samples; deriving modified transform coefficients by applying a low frequency non-separable transform (LFNST) to the transform coefficients; generating residual information based on the modified transform coefficients; as well as encoding the image information including the residual information, wherein the image information includes an LFNST index based on a DC significant coefficient variable indicating whether a significant coefficient exists only in a DC component of the current block, The DC significant coefficient variable is derived based on a first transform skip flag for a chroma Cb component of the current block and a second transform skip flag for a chroma Cr component of the current block, The step of deriving the modified transform coefficients comprises the following steps: applying the LFNST to at least one of a first transform coefficient associated with the chroma Cb component or a second transform coefficient associated with the chroma Cr component, Wherein, whether to apply the LFNST to the chroma Cb component is determined based on the LFNST index and the first transform skip flag, Wherein, whether to apply the LFNST to the chroma Cr component is determined based on the LFNST index and the second transform skip flag, and The tree type based on the current block is dual-tree chroma, and the image information includes the LFNST index.
10. The image encoding method according to claim 9, in, The value of the DC effective coefficient variable is initially set to 1, wherein, based on at least one of the value of the first transform skip flag and the value of the second transform skip flag being 0, the value of the DC significant coefficient variable is derived to be 0, and Wherein, based on the value of the DC effective coefficient variable being 0, the image information includes the LFNST index.
11. The image encoding method according to claim 9, in, The value of the DC significant coefficient variable is initially set to 1 at the coding unit level of the current block, and Wherein, based on at least one of the value of the first transform skip flag and the value of the second transform skip flag being 0, the value of the DC significant coefficient variable is set to 0 at the residual coding level.
12. The image encoding method according to claim 9, in, Based on the value of the LFNST index being greater than 0 and the value of the first transform skip flag for the chroma Cb component being 0, applying the LFNST to the chroma Cb component of the current block, and Wherein, based on the value of the LFNST index being greater than 0 and the value of the first transform skip flag for the chroma Cb component being 1, the LFNST is not applied to the chroma Cb component of the current block.
13. The image encoding method according to claim 9, in, Based on the value of the LFNST index being greater than 0 and the value of the second transform skip flag for the chroma Cr component being 0, applying the LFNST to the chroma Cr component of the current block, and Wherein, based on the value of the LFNST index being greater than 0 and the value of the second transform skip flag for the chroma Cr component being 1, the LFNST is not applied to the chroma Cr component of the current block.
14. The image encoding method according to claim 9, in, Based on the value of one of the first transform skip flag for the chroma Cb component and the second transform skip flag for the chroma Cr component being 1 and the other being 0, the value of the DC significant coefficient variable is derived to be 0.
15. A method for transmitting a bit stream, the method comprising the steps of: Obtaining a bitstream, wherein the bitstream is generated based on the following operations: deriving prediction samples for a current block, deriving residual samples for the current block based on the prediction samples, deriving transform coefficients for the current block based on a primary transform of the residual samples, deriving modified transform coefficients by applying a low-frequency non-separable transform (LFNST) to the transform coefficients, generating residual information based on the modified transform coefficients, and encoding image information including the residual information to output the bitstream; and sending the bitstream, wherein the image information includes an LFNST index based on a DC significant coefficient variable indicating whether a significant coefficient exists only in a DC component of the current block, The DC significant coefficient variable is derived based on a first transform skip flag for a chroma Cb component of the current block and a second transform skip flag for a chroma Cr component of the current block, The step of deriving the modified transform coefficients comprises the following steps: applying the LFNST to at least one of a first transform coefficient associated with the chroma Cb component or a second transform coefficient associated with the chroma Cr component, Wherein, whether to apply the LFNST to the chroma Cb component is determined based on the LFNST index and the first transform skip flag, Wherein, whether to apply the LFNST to the chroma Cr component is determined based on the LFNST index and the second transform skip flag, and The tree type based on the current block is dual-tree chroma, and the image information includes the LFNST index.