Image decoding method, image encoding method, and data transmission method
Through the LFNST index coding method, the problem of high transmission and storage costs of high-resolution, high-quality images/videos is solved, and more efficient image/video compression and encoding is achieved, which is suitable for immersive media and broadcasting with different image characteristics.
Patent Information
- Application Number
- CN202310699346.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-23
- Filing Date
- 2020-09-21
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2040-09-21
AI Technical Summary
When transmitting and storing high-resolution, high-quality images/videos, the existing technology increases the amount of information, resulting in high costs, and the broadcasting demand for immersive media and images/videos with different image characteristics is not effectively met.
The LFNST index coding method is adopted to improve the image coding efficiency by deriving modified transform coefficients and LFNST matrices, and applying them to sub-partition blocks to enhance the efficiency of the encoding and decoding processes.
It improves the overall image/video compression efficiency and increases the efficiency of transform index coding, making it suitable for the efficient compression and transmission of high-resolution and high-quality images/videos.
Smart Images

Figure CN116527937B_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application number 202080078893.5 (International application number: PCT / KR2020 / 012695, application date: September 21, 2020, invention name: Transformation-based image coding method and device thereof). Technical Field
[0002] The present disclosure relates to an image coding technology, and more particularly, to a method and apparatus for encoding an image based on transformation in an image coding system. Background Art
[0003] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K or higher ultra-high-definition (UHD) images / videos has been growing in various fields. As image / video data becomes higher resolution and higher quality, the amount of information or bit volume transmitted increases compared to traditional image data. Therefore, when using a medium such as a traditional wired / wireless broadband line to transmit image data or using an existing storage medium to store image / video data, its transmission cost and storage cost increase.
[0004] In addition, today, interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and broadcasting of images / videos having image characteristics different from real images such as game images is increasing.
[0005] Therefore, there is a need for an efficient image / video compression technology that effectively compresses and transmits or stores and reproduces information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention
[0006] Technical issues
[0007] A technical aspect of the present disclosure is to provide a method and apparatus for increasing image encoding efficiency.
[0008] Another technical aspect of the present disclosure is to provide a method and apparatus for increasing the efficiency of transform index encoding.
[0009] Yet another technical aspect of the present disclosure is to provide an image encoding method and apparatus using LFNST.
[0010] Yet another technical aspect of the present disclosure is to provide an image encoding method and apparatus for applying LFNST to a sub-partitioned block.
[0011] Technical Solution
[0012] In one aspect of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes deriving a modified transform coefficient, wherein deriving the modified transform coefficient includes deriving a first variable indicating whether the transform coefficient exists in a region other than a DC position of a current block; parsing an LFNST index based on the derivation result; and deriving the modified transform coefficient based on the LFNST index and an LFNST matrix. Based on the current block being divided into a plurality of sub-partition blocks, the LFNST index is parsed without deriving the first variable.
[0013] When the current block is not divided into a plurality of sub-partition blocks and the first variable indicates that a transform coefficient exists in a region other than a DC position, the LFNST index may be parsed.
[0014] The deriving of the modified transform coefficient may further include determining whether the transform coefficient exists in a second region other than the upper left first region of the current block, and when the transform coefficient does not exist in the second region, resolving the LFNST index.
[0015] When no transform coefficients exist at all in the second region of each of the plurality of sub-partition blocks, the LFNST index may be parsed.
[0016] The current block may be a coding block, and when the width and height of each sub-partition block are each greater than or equal to 4, the LFNST index of the current block may be applied to a plurality of sub-partition blocks.
[0017] When each divided sub-partition block is a 4×4 block or an 8×8 block, LFNST may be applied to up to the 8th transform coefficient in the scan direction from the upper left of the sub-partition block.
[0018] In another aspect of the present disclosure, an image encoding method performed by an encoding device is provided. The method includes: deriving transform coefficients for a current block based on a primary transform of residual samples of the current block; deriving modified transform coefficients from the transform coefficients based on an LFNST matrix used for LFNST; configuring image information so that an LFNST index indicating the LFNST matrix is signaled based on the presence of the transform coefficient in a region other than a DC position of the current block; and encoding quantized residual information and the LFNST index. Based on the current block being divided into a plurality of sub-partition blocks, the image information is configured so that the LFNST index is signaled regardless of whether the transform coefficient exists in a region other than a DC position of the current block.
[0019] According to still another embodiment of the present disclosure, a digital storage medium storing image data including a bit stream generated according to an image encoding method performed by an encoding device and encoded image information may be provided.
[0020] According to yet another embodiment of the present disclosure, a digital storage medium storing image data including encoded image information and a bit stream so that a decoding device performs an image decoding method may be provided.
[0021] Technical Effects
[0022] According to the present disclosure, the overall image / video compression efficiency can be increased.
[0023] According to the present disclosure, the efficiency of transform index encoding can be increased.
[0024] Technical aspects of the present disclosure may provide an image encoding method and apparatus using LFNST.
[0025] Another technical aspect of the present disclosure may provide an image encoding method and apparatus for applying LFNST to a sub-partitioned block.
[0026] The effects that can be obtained through the specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood or derived from the present disclosure by a person of ordinary skill in the relevant field. Therefore, the specific effects of the present disclosure are not limited to those explicitly described in the present disclosure, and may include various effects that can be understood or derived based on the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 An example of a video / image encoding system to which the present disclosure is applicable is schematically illustrated.
[0028] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which the present disclosure is applicable.
[0029] Figure 3 is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure is applicable.
[0030] Figure 4 Schematically illustrates multiple transformation schemes according to the embodiments of this document.
[0031] Figure 5 The intra directional mode for 65 prediction directions is schematically shown.
[0032] Figure 6 is a diagram for explaining RST according to an embodiment of this document.
[0033] Figure 7 is a diagram illustrating an order of arranging output data of a forward primary transform into a one-dimensional vector according to an example.
[0034] Figure 8is a diagram illustrating an order of arranging output data of a forward quadratic transform into two-dimensional blocks according to an example.
[0035] Figure 9 is a diagram illustrating a wide-angle intra prediction mode according to an embodiment of this document.
[0036] Figure 10 is a diagram illustrating a block shape to which LFNST is applied.
[0037] Figure 11 is a diagram illustrating arrangement of output data of a forward LFNST according to an embodiment.
[0038] Figure 12 is a diagram illustrating that the number of output data of the forward LFNST according to an example is limited to a maximum of 16.
[0039] Figure 13 is a diagram illustrating clearing of zeros in a block to which 4×4 LFNST is applied according to an example.
[0040] Figure 14 is a diagram illustrating clearing of zeros in a block to which 8×8 LFNST is applied according to an example.
[0041] Figure 15 is a diagram illustrating clearing in a block to which 8×8 LFNST is applied according to another example.
[0042] Figure 16 This is a diagram illustrating an example of sub-blocks into which one coding block is divided.
[0043] Figure 17 is a diagram illustrating another example of sub-blocks into which one coding block is divided.
[0044] Figure 18 is a diagram illustrating symmetry between an M×2 (M×1) block and a 2×M (1×M) block according to an embodiment.
[0045] Figure 19 is a diagram illustrating an example of transposing a 2×M block according to an embodiment.
[0046] Figure 20 The scanning order of 8×2 or 2×8 areas according to an embodiment is illustrated.
[0047] Figure 21 is a flowchart illustrating a method of decoding an image according to an embodiment.
[0048] Figure 22 is a flowchart illustrating a method of encoding an image according to an embodiment.
[0049] Figure 23This is a diagram illustrating the structure of a content stream system to which this document is applied. DETAILED DESCRIPTION
[0050] Although the present disclosure may be susceptible to various modifications and includes various embodiments, its specific embodiments have been shown by way of example in the accompanying drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the technical ideas of the present disclosure. Unless the context clearly indicates otherwise, the singular form may include the plural form. Terms such as "including" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components, or combinations thereof used in the following description and should therefore not be understood as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0051] In addition, for the convenience of describing different characteristic functions, each component in the drawings described herein is illustrated independently, however, it is not intended that each component is implemented by separate hardware or software. For example, any two or more of these components can be combined to form a single component, and any single component can be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent rights of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0052] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. In addition, in the accompanying drawings, the same reference numerals are used for the same components, and repeated description of the same components will be omitted.
[0053] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).
[0054] In this document, various embodiments related to video / image encoding may be provided, and unless otherwise specified, these embodiments may be combined with each other and performed.
[0055] In this document, video can refer to a collection of images over a period of time. Generally, a picture refers to a unit that represents an image in a specific time region, and a slice / tile is a unit that constitutes a part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.
[0056] A pixel or a picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. Alternatively, a sample may refer to a pixel value in a spatial domain, or when the pixel value is transformed into a frequency domain, it may refer to a transform coefficient in the frequency domain.
[0057] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region and information related to the region. A unit may include a luminance block and two chrominance (e.g., CB, CR) blocks. Depending on the situation, terms such as unit and block, region, etc. may be used interchangeably. In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0058] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". In addition, "A / B / C" may mean "at least one of A, B, and / or C".
[0059] Additionally, in this document, the term "or" should be interpreted as meaning "and / or." For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as meaning "additionally or alternatively."
[0060] In the present disclosure, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in the present disclosure, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as “at least one of A and B”.
[0061] Furthermore, in the present disclosure, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” Furthermore, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”
[0062] In addition, the brackets used in this disclosure may indicate "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, it may mean that "intra-frame prediction" is proposed as an example of "prediction." In other words, "prediction" in this disclosure is not limited to "intra-frame prediction," and "intra-frame prediction" is proposed as an example of "prediction." In addition, when "prediction (i.e., intra-frame prediction)" is indicated, it may also mean that "intra-frame prediction" is proposed as an example of "prediction."
[0063] Technical features described separately in one drawing in the present disclosure may be implemented separately or may be implemented simultaneously.
[0064] Figure 1 An example of a video / image encoding system to which the present disclosure is applicable is schematically illustrated.
[0065] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may transmit the encoded video / image information or data to the receive device in the form of a file or stream via a digital storage medium or a network.
[0066] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0067] The video source can obtain the video / image by capturing, synthesizing, or generating a video / image. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.
[0068] An encoding device can encode input video / images. It can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0069] The transmitter can transmit the encoded video / image information or data, output as a bitstream, to a receiver in a receiving device via a digital storage medium or network in the form of a file or stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received / extracted bitstream to a decoding device.
[0070] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0071] The renderer can render the decoded video / image, and the rendered video / image can be displayed on a display.
[0072] Figure 2 Schematically illustrates a configuration of a video / image encoding device to which the present disclosure is applicable. Hereinafter, the so-called video encoding device may include an image encoding device.
[0073] Reference Figure 2 , the encoding device 200 may include an image divider 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. Depending on the embodiment, the image divider 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be composed of one or more hardware components (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0074] The image divider 210 may divide the input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a maximum coding unit (LCU), the coding units may be recursively divided according to a quadtree, binary tree, ternary tree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, a coding unit may be divided into multiple coding units of a deeper depth. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that has not been further divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively divided into coding units of a deeper depth as needed, thereby allowing the optimally sized coding unit to be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be separated or divided from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0075] Depending on the situation, terms such as unit and block, region, etc. may be used instead of each other. In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample may be used as a term corresponding to a pixel or a picture element (pel) of a picture (or image).
[0076] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from the predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. The predictor 220 can perform prediction on the processing target block (hereinafter referred to as "current block") and can generate a prediction block including prediction samples of the current block. The predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. Information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0077] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference sample can be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0078] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate was used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, the residual signal cannot be sent. In the case of motion information prediction (motion vector prediction, MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0079] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction at the same time. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode to perform prediction on the block. The IBC prediction mode or the palette mode can be used for content image / video encoding such as games such as screen content coding (SCC). Although IBC basically performs prediction in the current block, its execution method is similar to inter-frame prediction in that it derives a reference block in the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0080] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate a reconstructed signal or a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT means a transform obtained from a curve graph when the relationship information between pixels is represented by a curve graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size rather than square blocks.
[0081] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream. The information about the quantized transform coefficients can be called residual information. The quantizer 233 can rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode information required for video / image reconstruction in addition to the quantized transform coefficients (e.g., syntax element values, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream on a unit basis of the network abstraction layer (NAL). The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. In addition, the video / image information may also include general constraint information. In the present disclosure, information and / or syntax elements sent / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded by the above-mentioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 or a memory (not shown) that stores it may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0082] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual sample) can be reconstructed. The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222, so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the target picture, and as described later, can be used for inter-frame prediction of the next picture performed by filtering.
[0083] Furthermore, in the picture encoding and / or reconstruction process, luma mapping with chroma scaling (LMCS) may be applied.
[0084] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, especially in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive ring filter, bilateral filter, etc. As discussed later in the description of each filtering method, the filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0085] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. Accordingly, the encoding device can avoid prediction mismatch in the encoding device 100 and the decoding device when applying inter-frame prediction, and can also improve encoding efficiency.
[0086] The memory 270DPB can store the modified reconstructed picture so that it can be used as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the blocks in the current picture from which the motion information has been derived (or encoded) and / or the motion information of the blocks in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 221 to be used as the motion information of the neighboring blocks or the motion information of the temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.
[0087] Figure 3is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure is applicable.
[0088] Reference Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-frame predictor 331 and an inter-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be composed of one or more hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0089] When a bit stream including video / image information is input, the decoding device 300 can Figure 2 The image is reconstructed correspondingly to the processing of the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the information related to the block segmentation obtained from the bit stream. The decoding device 300 can perform decoding by using the processing unit applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided into a quadtree structure, a binary tree structure and / or a ternary tree structure using a coding tree unit or a maximum coding unit. One or more transformation units can be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproducer.
[0090] The decoding device 300 may receive the data from the Figure 2The received signal is output by the encoding device, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. In the present disclosure, the signaled / received information and / or syntax elements described later can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, use the decoded target syntax element information and the decoded information of the neighboring and decoded target blocks or the information of the symbol / bin decoded in the previous step to determine the context model, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bin to generate the symbol corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded by the context model for the next symbol / bin after determining the context model. The information about prediction among the information decoded in the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (i.e., quantized transform coefficient) and associated parameter information for which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. In addition, a receiver (not shown) that receives a signal output from the encoding device may also constitute the decoding device 300 as an internal / external element, and the receiver may be a component of the entropy decoder 310. In addition, the decoding device according to the present disclosure may be referred to as a video / image / picture encoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0091] The dequantizer 321 can output the transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the order of coefficient scanning performed in the encoding device. The dequantizer 321 can dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0092] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing inverse transformation on the transformation coefficients.
[0093] The predictor may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and specifically may determine the intra / inter prediction mode.
[0094] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction at the same time. This can be called combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor can perform intra-block copying (IBC) for the prediction of the block. Intra-block copying can be used for content image / video coding such as games such as screen content coding (SCC). Although IBC basically performs prediction in the current block, its execution method is similar to inter-frame prediction in that it derives a reference block in the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0095] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0096] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the inter-frame prediction mode for the current block.
[0097] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor 330. When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.
[0098] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next processing target block in the current block, and as described later, may be output through filtering or used for inter prediction of the next picture.
[0099] Furthermore, in the picture decoding process, luma mapping with chroma scaling (LMCS) may be applied.
[0100] The filter 350 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can send the modified reconstructed picture to the memory 360, in particular, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive ring filter, bilateral filter, etc.
[0101] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block in the current picture from which the motion information has been derived (or decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the neighboring block or the motion information of the temporally neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.
[0102] In this specification, the examples described in the predictor 330, dequantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 may be similarly or correspondingly applied to the predictor 220, dequantizer 234, inverse transformer 235, and filter 260 of the encoding device 200, respectively.
[0103] As described above, prediction is performed in order to improve compression efficiency when performing video encoding. Accordingly, a prediction block including prediction samples for a current block as an encoding target block can be generated. Here, the prediction block includes prediction samples in a spatial domain (or a pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve image coding efficiency by signaling to the decoding device information (residual information) about the residual between the original block and the prediction block, rather than the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed picture including the reconstructed block.
[0104] Residual information can be generated through a transform process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients, so that it can signal the associated residual information to the decoding device (through a bitstream). Here, the residual information may include value information, position information, transform technology, transform kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / dequantization process based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive a residual block by dequantizing / inverse transforming the quantized transform coefficients to serve as a reference for inter-frame prediction of the next picture, and can generate a reconstructed picture based on this.
[0105] Figure 4 The multi-conversion technology according to the embodiment of the present disclosure is schematically illustrated.
[0106] Reference Figure 4 , the converter can correspond to the aforementioned Figure 2 The converter in the encoding device, and the inverse converter may correspond to the aforementioned Figure 2 The inverse transformer in the encoding device, or Figure 3 An inverse transformer in a decoding device.
[0107] The transformer may derive (primary) transform coefficients by performing a primary transform based on the residual samples (residual sample array) in the residual block (S410). This primary transform may be referred to as a core transform. In this document, the primary transform may be based on a multi-transform selection (MTS), and when multiple transforms are used as the primary transform, it may be referred to as a multi-core transform.
[0108] Multi-core transform may refer to a method for performing transforms using discrete cosine transform (DCT) type 2 and discrete sine transform (DST) type 7, DCT type 8, and / or DST type 1 in addition. In other words, multi-core transform may refer to a transform method that transforms a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this document, the primary transform coefficients may be referred to as temporary transform coefficients from the perspective of the transformer.
[0109] In other words, when a conventional transform method is applied, transform coefficients can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, when a multi-core transform is applied, transform coefficients (or primary transform coefficients) can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this document, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transform types, transform kernels, or transform cores. These DCT / DST transform types may be defined based on basis functions.
[0110] When performing multi-core transformation, a vertical transform kernel and a horizontal transform kernel for the target block can be selected from the transform kernels, a vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can indicate the transform of the horizontal component of the target block, and the vertical transform can indicate the transform of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block) including the residual block.
[0111] In addition, according to an example, if a transform is performed once by applying MTS, the mapping relationship of the transform kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transform or the horizontal transform. For example, when the horizontal transform kernel is represented as trTypeHor and the vertical transform kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT8.
[0112] In this case, the MTS index information may be encoded and signaled to the decoding device to indicate any one of a plurality of transform kernel sets. For example, an MTS index of 0 may indicate that both trTypeHor and trTypeVer values are 0, an MTS index of 1 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 2 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 3 may indicate that both trTypeHor and trTypeVer values are 1 and 2, and an MTS index of 4 may indicate that both trTypeHor and trTypeVer values are 2.
[0113] In one example, the transformation kernel set according to the MTS index information is shown in the following table.
[0114] [Table 1]
[0115] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0116] The transformer may perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients (S420). A primary transform is a transform from the spatial domain to the frequency domain, while a secondary transform refers to a transform into a more compact representation using the correlation existing between the (primary) transform coefficients. The secondary transform may include an inseparable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a pattern-dependent non-separable secondary transform (MDNSST). NSST may represent a transform that performs a secondary transform on the (primary) transform coefficients derived by the primary transform based on a non-separable transform matrix to generate modified transform coefficients (or secondary transform coefficients) for the residual signal. Here, based on the non-separable transform matrix, the transform may be applied once to the (primary) transform coefficients without separating the vertical transform and the horizontal transform (or applying the horizontal / vertical transform independently). In other words, NSST is not applied separately to (primary) transform coefficients in the vertical and horizontal directions, and can represent, for example, a transform method in which a two-dimensional signal (transform coefficient) is rearranged into a one-dimensional signal through a specific predetermined direction (e.g., a row-first direction or a column-first direction) and then the modified transform coefficients (or secondary transform coefficients) are generated based on an inseparable transform matrix. For example, the row-first order is to arrange the M×N blocks in the order of the first row, the second row, ... and the Nth row, while the column-first order is to arrange the M×N blocks in the order of the first column, the second column, ... and the Mth column. NSST can be applied to the upper left area of a block (hereinafter referred to as a transform coefficient block) configured with (primary) transform coefficients. For example, when the width W and the height H of the transform coefficient block are both 8 or larger, 8×8 NSST can be applied to the upper left 8×8 area of the transform coefficient block. Furthermore, while both the width (W) and the height (H) of the transform coefficient block are 4 or greater, when the width (W) or the height (H) of the transform coefficient block is less than 8, the 4×4 NSST may be applied to the upper left min(8, w)×min(8, H) region of the transform coefficient block. However, embodiments are not limited thereto, and for example, even if only the condition that the width W or the height H of the transform coefficient block is 4 or greater is satisfied, the 4×4 NSST may be applied to the upper left min(8, W)×min(8, H) region of the transform coefficient block.
[0117] Specifically, for example, if a 4×4 input block is used, the non-separable secondary transform may be performed as follows.
[0118] A 4×4 input block X can be represented as follows.
[0119] [Formula 1]
[0120]
[0121] If X is represented as a vector, then the vector It can be expressed as follows.
[0122] [Formula 2]
[0123]
[0124] In Equation 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of formula 1 according to row-major order.
[0125] In this case, the non-separable quadratic transform can be calculated as follows.
[0126] [Formula 3]
[0127]
[0128] In this formula, denotes a transform coefficient vector, and T denotes a 16x16 (non-separable) transform matrix.
[0129] By using the above formula 3, the 16×1 transform coefficient vector can be derived And the vector can be scanned in order (horizontally, vertically, diagonally, etc.) Reorganized into 4×4 blocks. However, the above calculation is an example, and Hypercube-Givens Transform (HyGT) or the like may also be used for the calculation of the inseparable secondary transform in order to reduce the computational complexity of the inseparable secondary transform.
[0130] Furthermore, in the inseparable secondary transform, the transform kernel (or transform core, transform type) may be selected to be mode-dependent. In this case, the mode may include an intra prediction mode and / or an inter prediction mode.
[0131] As described above, an inseparable secondary transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 area included in the transform coefficient block when both W and H are equal to or greater than 8, and the 8×8 area can be the upper left 8×8 area in the transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 area included in the transform coefficient block when both W and H are equal to or greater than 4, and the 4×4 area can be the upper left 4×4 area in the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0132] Here, in order to select mode-dependent transform kernels, two inseparable secondary transform kernels may be configured for each transform set used for inseparable secondary transforms for both the 8×8 transform and the 4×4 transform, and four transform sets may exist. That is, four transform sets may be configured for the 8×8 transform, and four transform sets may be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform may include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform may include two 4×4 transform kernels.
[0133] However, as the size of the transform (ie, the size of the region to which the transform is applied) may be other than 8×8 or 4×4, for example, the number of sets may be n, and the number of transform kernels in each set may be k.
[0134] The transform set may be referred to as an NSST set or a LFNST set. A specific set among the transform sets may be selected, for example, based on the intra prediction mode of the current block (CU or subblock). A low-frequency non-separable transform (LFNST) may be an example of a reduced non-separable transform, which will be described later and represents a non-separable transform for low-frequency components.
[0135] For reference, for example, the intra prediction mode may include two non-directional (or non-angle) intra prediction modes and 65 directional (or angle) intra prediction modes. The non-directional intra prediction mode may include a plane intra prediction mode No. 0 and a DC intra prediction mode No. 1, and the directional intra prediction mode may include 65 intra prediction modes No. 2 to No. 66. However, this is an example, and this document may be applied even if the number of intra prediction modes is different. In addition, in some cases, intra prediction mode No. 67 may also be used, and intra prediction mode No. 67 may represent a linear model (LM) mode.
[0136] Figure 5 The intra directional mode for 65 prediction directions is schematically shown.
[0137] Reference Figure 5 , based on the intra prediction mode 34 having the upper left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode having a horizontal directionality and an intra prediction mode having a vertical directionality. Figure 5In FIG, H and V denote horizontal and vertical directivities, respectively, and numbers -32 to 32 indicate displacements of 1 / 32 units on the sample grid position. These numbers may represent offsets for mode index values. Intra-prediction modes 2 to 33 have horizontal directivities, and intra-prediction modes 34 to 66 have vertical directivities. Strictly speaking, intra-prediction mode 34 may be considered neither horizontal nor vertical, but may be classified as belonging to horizontal directivity when determining the transform set for the secondary transform. This is because the input data is transposed for a vertically oriented mode that is symmetrical based on intra-prediction mode 34, and the input data alignment method for the horizontal mode is used for intra-prediction mode 34. Transposing the input data means switching the rows and columns of two-dimensional M×N block data to N×M data. Intra-prediction mode 18 and intra-prediction mode 50 may represent horizontal intra-prediction mode and vertical intra-prediction mode, respectively, and intra-prediction mode 2 may be referred to as an upper right diagonal intra-prediction mode because intra-prediction mode 2 has a left reference pixel and performs prediction in the upper right direction. Similarly, intra-prediction mode 34 may be referred to as a bottom-right diagonal intra-prediction mode, and intra-prediction mode 66 may be referred to as a bottom-left diagonal intra-prediction mode.
[0138] According to an example, four transform sets according to intra prediction modes may be mapped, for example, as shown in the following table.
[0139] [Table 2]
[0140] lfnstPredModeIntra lfnstTrSetIdx lfnstPredModeIntra<0 1 0<=lfnstPredModeIntra<=1 0 2<=lfnstPredModeIntra<=12 1 13<=lfnstPredModeIntra<=23 2 24<=1fnstPredModeIntra<=44 3 45<=lfnstPredModeIntra<=55 2 56<=lfnstPredModeIntra<=80 1 81<=lfnstPredModeIntra<=83 0
[0141] As shown in Table 2, any one of four transform sets, ie, lfnstTrSetIdx, may be mapped to any one of four indexes (ie, 0 to 3) according to the intra prediction mode.
[0142] When it is determined that a specific set is used for an inseparable transform, one of the k transform cores in the specific set can be selected by an inseparable secondary transform index. The encoding device can derive an inseparable secondary transform index indicating a specific transform core based on a rate-distortion (RD) check, and can signal the inseparable secondary transform index to the decoding device. The decoding device can select one of the k transform cores in the specific set based on the inseparable secondary transform index. For example, an lfnst index value 0 can refer to a first inseparable secondary transform core, an lfnst index value 1 can refer to a second inseparable secondary transform core, and an lfnst index value 2 can refer to a third inseparable secondary transform core. Alternatively, an lfnst index value 0 can indicate that the first inseparable secondary transform is not applied to the target block, and lfnst index values 1 to 3 can indicate three transform cores.
[0143] The transformer can perform a non-separable secondary transform based on the selected transform kernel and can obtain modified (secondary) transform coefficients. As described above, the modified transform coefficients can be derived as transform coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device and transmitted to the dequantizer / inverse transformer in the encoding device.
[0144] In addition, as described above, if the secondary transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device and transmitted to the dequantizer / inverse transformer in the encoding device.
[0145] The inverse transformer may perform a series of processes in the reverse order of the order already performed in the above-mentioned transformer. The inverse transformer may receive the (dequantized) transform coefficients and derive the (primary) transform coefficients by performing a secondary (inverse) transform (S450), and may obtain the residual block (residual sample) by performing a primary (inverse) transform on the (primary) transform coefficients (S460). In this regard, from the perspective of the inverse transformer, the primary transform coefficients may be referred to as modified transform coefficients. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.
[0146] The decoding device may further include a secondary inverse transform application determiner (or an element for determining whether to apply a secondary inverse transform) and a secondary inverse transform determiner (or an element for determining a secondary inverse transform). The secondary inverse transform application determiner may determine whether to apply a secondary inverse transform. For example, the secondary inverse transform may be NSST, RST, or LFNST, and the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a secondary transform flag obtained by parsing the bitstream. In another example, the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a transform coefficient of a residual block.
[0147] The secondary inverse transform determiner may determine the secondary inverse transform. In this case, the secondary inverse transform determiner may determine the secondary inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra-frame prediction mode. In an embodiment, the secondary transform determination method may be determined depending on the primary transform determination method. Various combinations of primary and secondary transforms may be determined according to the intra-frame prediction mode. In addition, in an example, the secondary inverse transform determiner may determine the area to which the secondary inverse transform is applied based on the size of the current block.
[0148] In addition, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, a (separable) inverse transform can be performed once, and a residual block (residual sample) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0149] Furthermore, in the present disclosure, reduced quadratic transform (RST) in which the size of a transformation matrix (kernel) is reduced may be applied in the concept of NSST in order to reduce the amount of calculation and storage required for an inseparable quadratic transform.
[0150] In addition, the transformation kernel, transformation matrix, and coefficients constituting the transformation kernel matrix described in the present disclosure, that is, kernel coefficients or matrix coefficients, can be represented in 8 bits. This can be a condition for implementation in decoding devices and encoding devices, and compared with existing 9 bits or 10 bits, the amount of storage required to store the transformation kernel can be reduced, and performance degradation can be reasonably adapted. In addition, representing the kernel matrix in 8 bits can allow the use of small multipliers and can be more suitable for single instruction multiple data (SIMD) instructions for optimal software implementation.
[0151] In this specification, the term "RST" may refer to a transform performed on residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. When performing a downscale transform, the amount of computation required for the transform can be reduced due to the reduction in the size of the transform matrix. In other words, RST can be used to address the computational complexity issues that arise when transforming large blocks or non-separable transforms.
[0152] RST may be referred to by various terms such as reduced transform, reduced secondary transform, downscaling transform, simplified transform, and simple transform, and the names that RST may be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in a low-frequency region including non-zero coefficients in a transform block, it may be referred to as a low-frequency non-separable transform (LFNST). The transform index may be referred to as an LFNST index.
[0153] In addition, when performing a secondary inverse transform based on an RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include: an inverse downscaled secondary transformer that derives modified transform coefficients based on an inverse RST of the transform coefficients; and an inverse primary transformer that derives residual samples of the target block based on an inverse primary transform of the modified transform coefficients. An inverse primary transform refers to an inverse transform of a primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.
[0154] Figure 6is a diagram illustrating an RST according to an embodiment of the present disclosure.
[0155] In this disclosure, a “target block” may refer to a current block to be encoded, a residual block, or a transform block.
[0156] In the RST according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, so that a reduced transformation matrix can be determined, where R is less than N. N can refer to the square of the length of the side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the reduction factor can refer to an R / N value. The reduction factor can be referred to as a reduction factor, a shrinkage factor, a simplification factor, a simple factor, or various other terms. In addition, R can be referred to as a reduction coefficient, but depending on the situation, the reduction factor can refer to R. In addition, depending on the situation, the reduction factor can refer to an N / R value.
[0157] In the example, the reduction factor or reduction coefficient may be signaled through the bitstream, but the example is not limited thereto. For example, a predetermined value for the reduction factor or reduction coefficient may be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or reduction coefficient may not be signaled separately.
[0158] The size of the reduced transform matrix according to an example may be R×N, which is smaller than N×N (the size of a conventional transform matrix), and may be defined as in Equation 4 below.
[0159] [Formula 4]
[0160]
[0161] Figure 6 The matrix T in the reduced transform block shown in (a) may refer to the matrix T of Equation 4. R×N .like Figure 6 As shown in (a), when the reduced transformation matrix T R×N When multiplied by the residual samples of the target block, the transform coefficients of the current block can be derived.
[0162] In an example, if the size of the block to which the transform is applied is 8×8 and R=16 (ie, R / N=16 / 64=1 / 4), then according to Figure 6 The RST of (a) can be expressed as a matrix operation shown in the following Equation 5. In this case, the storage and multiplication calculations can be reduced to about 1 / 4 by a reduction factor.
[0163] In the present disclosure, a matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix provided on the left side of the column vector.
[0164] [Formula 5]
[0165]
[0166] In formula 6, r1 to r 64 ∫ may represent the residual sample of the target block, and specifically may be a transform coefficient generated by applying one transform. As a result of the calculation of Equation 5, the transform coefficient c of the target block may be derived. i , and derive c i The process can be shown as Equation 6.
[0167] [Formula 6]
[0168]
[0169] As a result of the calculation of Equation 6, the transform coefficients c1 to c R That is, when R=16, the transform coefficients c1 to c 16 If a normal transform is applied instead of RST and a transform matrix of 64×64 (N×N) size is multiplied by a residual sample of 64×1 (N×1) size, only 16 (R) transform coefficients are derived for the target block because RST is applied, even though 64 (N) transform coefficients are derived for the target block. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data transmitted by the encoding device 200 to the decoding device 300 is reduced, and thus the transmission efficiency between the encoding device 200 and the decoding device 300 can be improved.
[0170] When considering the size of the transformation matrix, the size of the conventional transformation matrix is 64×64 (N×N), but the size of the reduced transformation matrix is reduced to 16×64 (R×N). Therefore, compared with the case of performing the conventional transformation, the memory usage when performing RST can be reduced by the R / N ratio. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional transformation matrix, the number of multiplication calculations (R×N) can be reduced by the R / N ratio when using the reduced transformation matrix.
[0171] In an example, the transformer 232 of the encoding device 200 may derive transform coefficients of the target block by performing a primary transform and a secondary transform based on an RST on the residual samples of the target block. These transform coefficients may be transmitted to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 may derive modified transform coefficients based on an inverse reduced secondary transform (RST) for the transform coefficients, and may derive residual samples of the target block based on an inverse primary transform for the modified transform coefficients.
[0172] According to the example, the inverse RST matrix T N×RThe size of is N×R which is smaller than the size of the conventional inverse transform matrix N×N and is similar to the reduced transform matrix T shown in Equation 4. R×N Has a transposition relationship.
[0173] Figure 6 The matrix T in the reduced inverse transform block shown in (b) t It can refer to the inverse RST matrix T N×R T (The superscript T refers to transposition). Figure 6 As shown in (b), when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. R×N T It can be expressed as (T R×N ) T N×R .
[0174] More specifically, when the inverse RST is used as the secondary inverse transform, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the residual samples of the target block can be derived.
[0175] In an example, if the size of the block to which the inverse transform is applied is 8×8 and R=16 (ie, R / N=16 / 64=1 / 4), then according to Figure 6 The RST of (b) can be expressed as the matrix operation shown in the following Equation 7.
[0176] [Formula 7]
[0177]
[0178] In formula 7, c1 to c 16 As a result of the calculation of Equation 7, r representing the modified transform coefficient of the target block or the residual sample of the target block can be derived. j , and derive r j The process can be shown as formula 8.
[0179] [Formula 8]
[0180]
[0181] As a result of the calculation of Equation 8, r1 to r2 representing the modified transform coefficients of the target block or the residual samples of the target block can be derived. N . From the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduced transform matrix is reduced to 64×16 (R×N), so the storage usage in the case of performing inverse RST can be reduced by the R / N ratio compared to the case of performing the conventional inverse transform. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional inverse transform matrix, the use of the inverse reduced transform matrix can reduce the number of multiplication calculations (N×R) by the R / N ratio.
[0182] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied according to the transform set in Table 2. Since a transform set includes two or three transforms (kernels) according to the intra prediction mode, it can be configured to select one of up to four transforms, including the case where the secondary transform is not applied. In the transform where the secondary transform is not applied, the application of the identity matrix can be considered. Assuming that indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case where the identity matrix is applied, that is, the case where the secondary transform is not applied), the transform index or LFNST index as a syntax element can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the upper left 8×8 block, the 8×8 NSST in the RST configuration can be specified by the transform index, or the 8×8 LFNST can be specified when LFNST is applied. 8×8lfnst and 8×8RST refer to transforms that can be applied to an 8×8 region included in a transform coefficient block when both W and H of a target block to be transformed are equal to or greater than 8, and the 8×8 region may be the upper left 8×8 region in the transform coefficient block. Similarly, 4×4lfnst and 4×4RST refer to transforms that can be applied to a 4×4 region included in a transform coefficient block when both W and H of a target block are equal to or greater than 4, and the 4×4 region may be the upper left 4×4 region in the transform coefficient block.
[0183] According to an embodiment of the present disclosure, for the transformation in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 area. Here, “maximum” means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is, when RST is performed by applying an m×48 transform kernel matrix (m≤16) to an 8×8 area, 48 pieces of data are input and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming an 8×8 area can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on the 48 pieces of data constituting the area other than the lower right 4×4 area among the 8×8 areas. Here, when the matrix operation is performed by applying the maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper left 4×4 area according to the scanning order, and the upper right 4×4 area and the lower left 4×4 area can be filled with zeros.
[0184] For the inverse transform in the decoding process, a transposed matrix of the aforementioned transform kernel matrix may be used. That is, when inverse RST or LFNST is performed in the inverse transform process performed by the decoding device, input coefficient data to which inverse RST is applied is arranged in a one-dimensional vector according to a predetermined arrangement order, and modified coefficient vectors obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector may be arranged in a two-dimensional block according to a predetermined arrangement order.
[0185] In summary, during the transform process, when RST or LFNST is applied to an 8×8 region, the 48 transform coefficients in the upper left, upper right, and lower left regions of the 8×8 region, excluding the lower right region, are subjected to a matrix operation with the 16×48 transform kernel matrix. For the matrix operation, the 48 transform coefficients are input as a one-dimensional array. When the matrix operation is performed, 16 modified transform coefficients are derived and arranged in the upper left region of the 8×8 region.
[0186] On the contrary, in the inverse transform process, when the inverse RST or LFNST is applied to an 8×8 area, 16 transform coefficients corresponding to the upper left area of the 8×8 area among the transform coefficients in the 8×8 area can be input in a one-dimensional array according to the scanning order, and can undergo a matrix operation with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix)*(16×1 transform coefficient vector)=(48×1 modified transform coefficient vector). Here, the n×1 vector can be interpreted as having the same meaning as the n×1 matrix, and can therefore be expressed as an n×1 column vector. In addition, * represents matrix multiplication. When the matrix operation is performed, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left area, the upper right area, and the lower left area of the 8×8 area except the lower right area.
[0187] When the secondary inverse transform is based on the RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include an inverse downscaling secondary transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients and an inverse primary transformer for deriving residual samples of the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.
[0188] The non-separable transform (LFNST) described above will be described in detail as follows: LFNST may include a forward transform performed by an encoding device and an inverse transform performed by a decoding device.
[0189] The encoding device receives as input a result (or a portion of the result) derived after applying a primary (core) transform, and applies a forward secondary transform (secondary transform).
[0190] [Formula 9]
[0191] y=G T x
[0192] In Equation 9, x and y are the input and output of the quadratic transform, respectively, G is a matrix representing the quadratic transform, and the transform basis vectors consist of column vectors. In the case of inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the transpose of the matrix G becomes G T dimension.
[0193] For inverse LFNST, the dimensions of the matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of 8 transformation basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.
[0194] On the other hand, for forward LFNST, the matrix G T The dimensions are [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transformation basis vectors from the upper parts of the [16×48] matrix and the [16×16] matrix, respectively.
[0195] Therefore, in the case of forward LFNST, a [48×1] vector or a [16×1] vector can be used as input x, and a [16×1] vector or a [8×1] vector can be used as output y. In video encoding and decoding, the output of the forward primary transform is two-dimensional (2D) data, so in order to construct a [48×1] vector or a [16×1] vector as input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data as the output of the forward transform.
[0196] Figure 7 is a diagram illustrating an order of arranging output data of a forward primary transform into a one-dimensional vector according to an example. Figure 7 The left figures of (a) and (b) show the order for constructing a [48×1] vector, and Figure 7 The right figures of (a) and (b) show the order for constructing a [16×1] vector. In the case of LFNST, the 2D data can be constructed by Figure 7 The same order as in (a) and (b) is sequentially arranged to obtain a one-dimensional vector x.
[0197] The arrangement direction of the output data of the forward primary transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in the horizontal direction relative to the diagonal direction, the output data of the forward primary transform can be arranged in the following manner: Figure 7 The output data of the forward primary transform are arranged in the order of (a) and when the intra prediction mode of the current block is in a vertical direction relative to the diagonal direction, the output data of the forward primary transform can be arranged in the order of (a) Figure 7 The output data of the forward primary transformation are arranged in the order of (b).
[0198] According to the example, different Figure 7 The arrangement order of (a) and (b) is the arrangement order of (a) and (b), and in order to derive and apply Figure 7If the arrangement order of (a) and (b) is the same as the result (y vector), the column vectors of the matrix G can be rearranged according to the arrangement order. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0199] Since the output y derived by Formula 9 is a one-dimensional vector, when two-dimensional data is required as input data in a process using the result of the forward quadratic transform as input (for example, in the process of performing quantization or residual coding), the output y vector of Formula 9 needs to be properly arranged as 2D data again.
[0200] Figure 8 is a diagram illustrating an order of arranging output data of a forward quadratic transform into two-dimensional blocks according to an example.
[0201] In the case of LFNST, the output values can be arranged in 2D blocks according to a predetermined scanning order. Figure 8 (a) shows that when the output y is a [16×1] vector, the output values are arranged at 16 positions of the 2D block according to the diagonal scanning order. Figure 8 (b) shows that when the output y is an [8×1] vector, the output values are arranged at 8 positions of the 2D block according to the diagonal scanning order, and the remaining 8 positions are filled with zeros. Figure 8 The X in (b) indicates that it is filled with zeros.
[0202] According to another example, since the order of processing the output vector y when performing quantization or residual encoding can be preset, the output vector y may not be arranged in a sequence such as Figure 8 However, in the case of residual coding, data encoding can be performed in 2D block (e.g., 4×4) units (e.g., CG (coefficient group)), and in this case, according to Figure 8 The data is arranged in a specific order in the diagonal scan order of .
[0203] In addition, the decoding apparatus may configure a one-dimensional input vector y by arranging two-dimensional data output through a dequantization process according to a preset scanning order for inverse transform. The input vector y may be output as an output vector x through the following equation.
[0204] [Equation 10]
[0205] X=Gy
[0206] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or a [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.
[0207] The output vector x is based on Figure 7 The order shown in is arranged in a two-dimensional block and is arranged as two-dimensional data, and the two-dimensional data becomes input data (or a part of input data) of an inverse primary transform.
[0208] Therefore, the inverse quadratic transform is overall the reverse of the forward quadratic transform process, and in the case of the inverse transform, unlike in the forward direction, the inverse quadratic transform is applied first and then the inverse primary transform.
[0209] In the inverse LFNST, one of 8 [48×16] matrices and 8 [16×16] matrices can be selected as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.
[0210] In addition, eight matrices can be derived from the four transform sets shown in Table 2 above, and each transform set can be composed of two matrices. Which transform set to use among the four transform sets is determined according to the intra prediction mode, and more specifically, the transform set is determined based on the value of the intra prediction mode extended by considering wide-angle intra prediction (WAIP). Which matrix is selected from the two matrices constituting the selected transform set is derived by index signaling. More specifically, 0, 1, and 2 can be used as the index values sent, 0 can indicate that LFNST is not applied, and 1 and 2 can indicate either of the two transform matrices constituting the transform set selected based on the intra prediction mode value.
[0211] Figure 9 is a diagram illustrating a wide-angle intra prediction mode according to an embodiment of this document.
[0212] The general intra prediction mode value may have values from 0 to 66 and from 81 to 83, and the intra prediction mode value extended due to WAIP may have values from -14 to 83 as shown. The values from 81 to 83 indicate the CCLM (Cross Component Linear Model) mode, and the values from -14 to -1 and the values from 67 to 80 indicate the intra prediction mode extended due to WAIP application.
[0213] When the width of the current prediction block is greater than its height, the upper reference pixel is generally closer to the position inside the block to be predicted. Therefore, prediction in the lower left direction can be more accurate than in the upper right direction. Conversely, when the height of the block is greater than its width, the left reference pixel is generally closer to the position inside the block to be predicted. Therefore, prediction in the upper right direction can be more accurate than in the lower left direction. Therefore, applying remapping (i.e., mode index modification) to the index of the Wide Intra prediction mode can be advantageous.
[0214] When wide-angle intra prediction is applied, information about existing intra predictions may be signaled, and after the information is parsed, the information may be remapped to the index of the wide-angle intra prediction mode. Therefore, the total number of intra prediction modes for a specific block (e.g., a non-square block of a specific size) may not be changed, that is, the total number of intra prediction modes is 67, and the intra prediction mode encoding for the specific block may not be changed.
[0215] Table 3 below shows a process of deriving a modified intra mode by remapping the intra prediction mode to the wide-angle intra prediction mode.
[0216] [Table 3]
[0217]
[0218] In Table 3, the extended intra prediction mode value is finally stored in the predModeIntra variable, and ISP_NO_SPLIT indicates that the CU block is not divided into sub-partitions by the intra sub-partitioning (ISP) technology currently adopted in the VVC standard, and the cIdx variable values 0, 1, and 2 indicate the cases of the luma component, Cb component, and Cr component, respectively. The log2 function shown in Table 3 returns a log value with a base of 2, and the Abs function returns an absolute value.
[0219] The variable predModeIntra indicating the intra prediction mode and the height and width of the transform block are used as input values for the wide-angle intra prediction mode mapping process, and the output value is the modified intra prediction mode predModeIntra. The height and width of the transform block or coding block can be the height and width of the current block used for remapping of the intra prediction mode. At this time, the variable whRatio reflecting the ratio of width to width can be set to Abs(Log2(nW / nH)).
[0220] For non-square blocks, the intra prediction mode can be divided into two cases and modified.
[0221] First, if all of conditions (1) to (3) are satisfied, (1) the width of the current block is greater than the height, (2) the intra prediction mode before modification is equal to or greater than 2, and (3) the intra prediction mode is less than a value derived as (8+2*whRatio) when the variable whRatio is greater than 1 and is less than 8 when the variable whRatio is less than or equal to 1 (predModeIntra is less than (whRatio>1)?(8+2*whRatio):8), the intra prediction mode is set to a value 65 greater than predModeIntra [predModeIntra is set equal to (predModeIntra+65)].
[0222] If different from the above, that is, if conditions (1) to (3) are satisfied, (1) the height of the current block is greater than the width, (2) the intra-frame prediction mode before modification is less than or equal to 66, and (3) the intra-frame prediction mode is greater than the value derived as (60-2*whRatio) when whRatio is greater than 1 and is greater than 60 when whRatio is less than or equal to 1 (predModeIntra is greater than (whRatio>1)?(60-2*whRatio):60), then the intra-frame prediction mode is set to a value that is 67 less than predModeIntra [predModeIntra is set equal to (predModeIntra-67)].
[0223] Table 2 above shows how to select a transform set based on the intra prediction mode value extended by WAIP in LFNST. Figure 9 As shown, modes 14 to 33 and modes 35 to 80 are symmetric about the prediction direction around mode 34. For example, mode 14 and mode 54 are symmetric about the direction corresponding to mode 34. Therefore, the same set of transforms is applied to modes located in mutually symmetric directions, and this symmetry is also reflected in Table 2.
[0224] In addition, it is assumed that the forward LFNST input data of mode 54 is symmetric with the forward LFNST input data of mode 14. For example, for mode 14 and mode 54, according to Figure 7 (a) and Figure 7 The arrangement order shown in (b) rearranges the two-dimensional data into one-dimensional data. In addition, it can be seen that Figure 7 (a) and Figure 7 The pattern of the sequence shown in (b) is symmetrical about the direction indicated by the pattern 34 (diagonal direction).
[0225] Furthermore, as described above, which transform matrix of the [48×16] matrix and the [16×16] matrix is applied to the LFNST is determined by the size and shape of the transform target block.
[0226] Figure 10 is a diagram illustrating a block shape to which LFNST is applied. Figure 10 (a) shows a 4×4 block, Figure 10 (b) shows a 4×8 block and an 8×4 block, Figure 10 (c) shows a 4×N block or an N×4 block, where N is 16 or greater, Figure 10 (d) shows an 8×8 block, Figure 10 (e) shows an M×N block, where M≥8, N≥8, and N>8 or M>8.
[0227] exist Figure 10In , blocks with thick borders indicate the areas where LFNST is applied. Figure 10 For the blocks (a) and (b), LFNST is applied to the top left 4×4 region, and for Figure 10 In the block (c), LFNST is applied separately to the two upper left 4×4 regions that are arranged consecutively. Figure 10 In (a), (b), and (c), since LFNST is applied in units of 4×4 regions, this LFNST will be referred to as “4×4 LFNST” hereinafter. Depending on the matrix dimension of G, a [16×16] or [16×8] matrix can be applied.
[0228] More specifically, the [16×8] matrix is applied to Figure 10 (a) 4×4 block (4×4TU or 4×4CU), and the [16×16] matrix is applied to Figure 10 This is to adjust the worst-case computational complexity to 8 multiplications per sample.
[0229] about Figure 10 In (d) and (e), LFNST is applied to the upper left 8×8 region, and this LFNST is hereinafter referred to as "8×8LFNST". As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of forward LFNST, since a [48×1] vector (the X vector in Equation 9) is input as input data, not all sample values of the upper left 8×8 region are used as input values of the forward LFNST. That is, as can be seen from Figure 7 The left order of (a) or Figure 7 As can be seen from the left order of (b), a [48×1] vector can be constructed based on samples belonging to the remaining three 4×4 blocks while leaving the lower right 4×4 block as it is.
[0230] The [48×8] matrix can be applied to Figure 10 The 8×8 block (8×8TU or 8×8CU) in (d) and the [48×16] matrix can be applied to Figure 10 The 8×8 block in (e) is used to adjust the worst-case computational complexity to 8 multiplications per sample.
[0231] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data are generated (Y vector in Equation 9, [8×1] or [16×1] vector). In the forward LFNST, since the matrix G T The characteristic is that the amount of output data is equal to or less than the amount of input data.
[0232] Figure 11 is a diagram illustrating arrangement of output data of a forward LFNST according to an example, and shows blocks in which the output data of the forward LFNST is arranged according to block shapes.
[0233] exist Figure 11 The shaded area on the upper left of the block shown corresponds to the area where the output data of the forward LFNST is located, the positions marked with 0 indicate samples filled with the value 0, and the remaining area represents the area not changed by the forward LFNST. In the area not changed by the LFNST, the output data of the forward primary transform remains unchanged.
[0234] As described above, since the size of the applied transformation matrix varies according to the shape of the block, the amount of output data also varies. Figure 11 , the output data of the forward LFNST may not completely fill the upper left 4×4 block. Figure 11 In the cases of (a) and (d), the [16×8] matrix and the A[48×8] matrix are applied to the block indicated by the bold line or the partial area inside the block, respectively, and the [8×1] vector is generated as the output of the forward LFNST. That is, according to Figure 8 The scanning order shown in (b) can only fill 8 output data, such as Figure 11 As shown in (a) and (d), the remaining 8 positions can be filled with 0. Figure 10 (d) The case of LFNST application blocks, such as Figure 11 As shown in (d), the two 4×4 blocks on the upper right and lower left adjacent to the upper left 4×4 block are also filled with values of 0.
[0235] As described above, basically, by signaling the LFNST index, it is specified whether to apply LFNST and the transformation matrix to be applied. Figure 11 As shown, when LFNST is applied, since the number of output data of the forward LFNST may be equal to or less than the number of input data, an area filled with zero values occurs as follows.
[0236] 1) If Figure 11 As shown in (a), samples from the eighth position and subsequent positions in the scanning order in the upper left 4×4 block, that is, samples from the ninth to sixteenth positions.
[0237] 2) If Figure 11 As shown in (d) and (e), when the [48×16] matrix or the [48×8] matrix is applied, two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scanning order.
[0238] Therefore, if non-zero data exists by checking areas 1) and 2), it is determined that LFNST is not applied, so that signaling of the corresponding LFNST index can be omitted.
[0239] According to an example, for example, in the case of LFNST adopted in the VVC standard, since LFNST index signaling is performed after residual coding, the encoding device can determine whether non-zero data (significant coefficients) exist at all locations within the TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform LFNST index signaling based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. When non-zero data does not exist in the areas specified in 1) and 2) above, LFNST index signaling is performed.
[0240] Since truncated unary codes are applied as the binarization method of the LFNST index, the LFNST index consists of up to two bins, and 0, 10, and 11 are assigned as binary codes for possible LFNST index values 0, 1, and 2, respectively. In the case of LFNST currently used for VVC, context-based CABAC coding is applied to the first bin (conventional coding), and bypass coding is applied to the second bin. The total number of contexts for the first bin is 2. When (DCT-2, DCT-2) is applied as a transform pair for the horizontal and vertical directions and the luminance component and the chrominance component are encoded in a dual-tree type, one context is allocated and the other context is applied to the remaining cases. The encoding of the LFNST index is shown in the following table.
[0241] [Table 4]
[0242]
[0243] In addition, for the adopted LFNST, the following simplified method can be applied.
[0244] (i) According to an example, the number of output data of the forward LFNST may be limited to a maximum of 16.
[0245] exist Figure 10 In the case of (c), 4×4 LFNST can be applied to two 4×4 regions adjacent to the upper left, respectively, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to a maximum of 16, in the case of a 4×N / N×4 (N≥16) block (TU or CU), 4×4 LFNST is applied only to one 4×4 region in the upper left, and LFNST can be applied only to Figure 10 By doing this, the implementation of image coding can be simplified.
[0246] Figure 12 It is shown that the number of output data of the forward LFNST according to the example is limited to a maximum of 16. Figure 12 , when LFNST is applied to the upper leftmost 4×4 region in a 4×N or N×4 block (where N is 16 or greater), the output data of the forward LFNST becomes 16.
[0247] (ii) According to an example, zeroing can be additionally applied to areas to which LFNST is not applied. In this document, zeroing can mean filling all positions belonging to a specific area with a value of 0. That is, zeroing can be applied to areas that are unchanged due to LFNST, and the result of the forward primary transform can be maintained. As described above, since LFNST is divided into 4×4 LFNST and 8×8 LFNST, zeroing can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.
[0248] (ii)-(A) When 4×4 LFNST is applied, an area to which 4×4 LFNST is not applied may be cleared. Figure 13 is a diagram illustrating clearing of zeros in a block to which 4×4 LFNST is applied according to an example.
[0249] like Figure 13 As shown, for the block to which 4×4 LFNST is applied, that is, for Figure 11 For all blocks in (a), (b), and (c), the entire area where LFNST is not applied can be filled with zeros.
[0250] on the other hand, Figure 13 (d) shows that when the maximum value of the number of output data of the forward LFNST is limited to 16 (as Figure 12 ), zeroing is performed on the remaining blocks to which 4×4 LFNST is not applied.
[0251] (ii)-(B) When 8×8 LFNST is applied, an area to which 8×8 LFNST is not applied may be cleared. Figure 14 is a diagram illustrating clearing of zeros in a block to which 8×8 LFNST is applied according to an example.
[0252] like Figure 14 As shown, for the block where 8×8 LFNST is applied, i.e., for Figure 11 For all blocks in (d) and (e), the entire area where LFNST is not applied can be filled with zeros.
[0253] (iii) Due to the zeroing presented in (ii) above, the area filled with zeros may not be the same as when LFNST is applied. Figure 11In the case of LFNST, the wider region performs the zeroing proposed in (ii) to check whether there is non-zero data.
[0254] For example, when (ii)-(B) is applied, when checking Figure 11 After the zero-filled areas in (d) and (e) have non-zero data, additionally check Figure 14 Whether there is non-zero data in the area filled with 0, the signaling of the LFNST index can be performed only when there is no non-zero data.
[0255] Of course, even if the clearing proposed in (ii) is applied, the presence of non-zero data can be checked in the same way as the existing LFNST index signaling. Figure 11 After checking whether there is non-zero data in the block filled with zeros in , LFNST index signaling can be applied. In this case, the encoding device only performs zero clearing and the decoding device does not assume zero clearing, that is, it only checks whether non-zero data exists only in Figure 11 In the regions explicitly marked as 0 in , LFNST index parsing can be performed.
[0256] Alternatively, according to another example, the following may be performed: Figure 15 Cleared as shown. Figure 15 is a diagram illustrating clearing in a block to which 8×8 LFNST is applied according to another example.
[0257] like Figure 13 and Figure 14 As shown, zeroing can be applied to all regions except the region where LFNST is applied, or zeroing can be applied only to local regions, as shown in Figure 15 Clearing to zero applies only to Figure 15 For areas outside the upper left 8×8 area, clearing may not be applied to the lower right 4×4 block within the upper left ×8 area.
[0258] Various embodiments of applying the combination of the simplified methods ((i), (ii)-(A), (ii)-(B), (iii)) of LFNST can be derived. Of course, the combination of the simplified methods described above is not limited to the following embodiments, and any combination can be applied to LFNST.
[0259] Implementation Method
[0260] -Limit the number of output data of the forward LFNST to a maximum of 16 → (i)
[0261] - When 4×4 LFNST is applied, all regions to which 4×4 LFNST is not applied are cleared → (II)-(A)
[0262] - When 8×8 LFNST is applied, all regions where 8×8 LFNST is not applied are cleared → (II)-(B)
[0263] - After checking whether non-zero data also exists in the existing areas filled with zero values and the areas filled with zeros due to additional clearing ((ii)-(A), (ii)-(B)), signal the LFNST index only if no non-zero data exists → (iii).
[0264] In the case of an embodiment, when LFNST is applied, the area where non-zero output data can exist is limited to the interior of the upper left 4×4 area. In more detail, in Figure 13 (a) and Figure 14 In the case of (a), the eighth position in the scanning order is the last position where non-zero data can exist. Figure 13 (b) and (c) and Figure 14 In the case of (b), the sixteenth position in the scanning order (ie, the position of the lower right edge of the upper left 4×4 block) is the last position in which data other than 0 may exist.
[0265] Therefore, after applying LFNST, after checking whether non-zero data exists at a position not allowed by the residual encoding process (at a position beyond the last position), it may be determined whether to signal the LFNST index.
[0266] In the case of the zeroing method proposed in (ii), the amount of data ultimately generated when both the primary transform and LFNST are applied can be reduced, thereby reducing the amount of computation required to perform the entire transform process. Specifically, when LFNST is applied, since the output data from the forward primary transform is present in areas where LFNST is not applied, there is no need to generate data for areas that were zeroed during the forward primary transform. Consequently, the amount of computation required to generate the corresponding data can be reduced. Additional benefits of the zeroing method proposed in (ii) are summarized below.
[0267] First, as mentioned above, the amount of computation required to perform the entire transformation process is reduced.
[0268] In particular, when (ii)-(B) is applied, the worst-case computational effort is reduced, making the transform process lightweight. In other words, generally speaking, a large amount of computation is required to perform a large-scale transform. By applying (ii)-(B), the amount of data derived as a result of performing forward LFNST can be reduced to 16 or less. In addition, as the size of the entire block (TU or CU) increases, the effect of reducing the number of transform operations further increases.
[0269] Second, the amount of computation required for the entire transformation process can be reduced, thereby reducing the power consumption required to perform the transformation.
[0270] Third, the delay involved in the transformation process is reduced.
[0271] Secondary transforms such as LFNST add computational complexity to the existing primary transform, thus increasing the overall latency involved in performing the transform. In particular, in the case of intra prediction, since reconstructed data from neighboring blocks is used in the prediction process, the increased latency due to the secondary transform during encoding results in an increased latency until reconstruction. This can lead to an increase in the overall latency of intra prediction encoding.
[0272] However, if the clearing proposed in (ii) is applied, the delay time for performing one transform can be greatly reduced when LFNST is applied, maintaining or reducing the delay time of the entire transform, so that the encoding device can be implemented more simply.
[0273] In conventional intra prediction, the block currently to be encoded is regarded as one coding unit, and encoding is performed without segmentation. However, intra subpartitioning (ISP) encoding means performing intra prediction encoding by dividing the block currently to be encoded in the horizontal direction or the vertical direction. In this case, a reconstructed block can be generated by performing encoding / decoding in units of divided blocks, and the reconstructed block can be used as a reference block for the next divided block. According to an embodiment, in ISP encoding, one coding block can be divided into two or four sub-blocks and encoded, and in ISP, in one sub-block, intra prediction is performed with reference to the reconstructed pixel value of the sub-block located on the adjacent left side or the adjacent upper side. Hereinafter, "encoding" may be used as a concept including both encoding performed by an encoding device and decoding performed by a decoding device.
[0274] Table 5 represents the number of sub-blocks divided according to the block size when ISP is applied, and the sub-partitions divided according to ISP may be referred to as transform blocks (TUs).
[0275] [Table 5]
[0276] Block size (CU) Number of divisions 4×4 Unavailable 4×8、8×4 2 All other cases 4
[0277] ISP divides the block predicted as part of the luma frame into two or four sub-partitions vertically or horizontally based on the block size. For example, the minimum block size to which ISP can be applied is 4×8 or 8×4. When the block size is larger than 4×8 or 8×4, the block is divided into four sub-partitions.
[0278] Figure 16 and 17 illustrates an example of sub-blocks into which a coding block is divided, and more specifically, Figure 16An example of partitioning in which the coding block (width (W) × height (H)) is a 4×8 block or an 8×4 block is illustrated, and Figure 17 An example of division is illustrated for a case where the coding block is not a 4×8 block, an 8×4 block, or a 4×4 block.
[0279] When ISP is applied, subblocks are sequentially encoded from left to right or from top to bottom according to the partition type (e.g., horizontally or vertically), and after performing reconstruction processing via inverse transform and intra prediction for one subblock, encoding of the next subblock can be performed. For the leftmost or topmost subblock, the reconstructed pixels of the already encoded coding block are referenced, as in the conventional intra prediction method. In addition, when each side of the subsequent internal subblock is not adjacent to the previous subblock, in order to derive the reference pixels adjacent to the corresponding side, the reconstructed pixels of the already encoded adjacent coding block are referenced, as in the conventional intra prediction method.
[0280] In ISP coding mode, all sub-blocks can be encoded with the same intra prediction mode, and a flag indicating whether ISP coding is used and a flag indicating in which direction to split (horizontally or vertically) can be signaled. Figure 16 and Figure 17 As shown, the number of sub-blocks can be adjusted to 2 or 4 according to the shape of the block. When the size (width × height) of a sub-block is less than 16, it can be restricted so that division into corresponding sub-blocks is not allowed or ISP encoding itself is not applied.
[0281] In the case of the ISP prediction mode, one coding unit is divided into two or four partition blocks (ie, subblocks) and predicted, and the same intra prediction mode is applied to the divided two or four partition blocks.
[0282] As described above, in the division direction, both the horizontal direction (when an M×N coding unit having a horizontal length and a vertical length of M and N, respectively, is divided in the horizontal direction, if the M×N coding unit is divided into two, the M×N coding unit is divided into M×(N / 2) blocks, and if the M×N coding unit is divided into four blocks, the M×N coding unit is divided into M×(N / 4) blocks) and the vertical direction (when the M×N coding unit is divided in the vertical direction, if the M×N coding unit is divided into two, the M×N coding unit is divided into (M / 2)×N blocks, and if the M×N coding unit is divided into four, the M×N coding unit is divided into (M / 4)×N blocks) are possible. When the M×N coding unit is divided in the horizontal direction, the partition blocks are encoded in a top-to-bottom order, and when the M×N coding unit is divided in the vertical direction, the partition blocks are encoded in a left-to-right order. In the case of horizontal (vertical) partitioning, the currently encoded partition block can be predicted by referring to the reconstructed pixel values of the upper (left) partition block.
[0283] A transform can be applied to the residual signal generated in units of partition blocks by the ISP prediction method. A multi-transform selection (MTS) technique based on a DST-7 / DCT-8 combination and the existing DCT-2 can be applied to a forward-based primary transform (core transform), and a forward low-frequency non-separable transform (LFNST) can be applied to the transform coefficients generated from the primary transform to generate the final modified transform coefficients.
[0284] That is, LFNST can be applied to partition blocks divided by applying the ISP prediction mode, and the same intra prediction mode is applied to the partitioned partition blocks, as described above. Therefore, when an LFNST set derived based on the intra prediction mode is selected, the derived LFNST set can be applied to all partition blocks. That is, because the same intra prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks.
[0285] According to an embodiment, LFNST can be applied only to transform blocks with both horizontal and vertical lengths of 4 or greater. Therefore, when the horizontal or vertical length of a partition block divided according to the ISP prediction method is less than 4, LFNST is not applied and the LFNST index is not signaled. In addition, when LFNST is applied to each partition block, the corresponding partition block can be regarded as a transform block. When the ISP prediction method is not applied, LFNST can be applied to the coding block.
[0286] A method of applying LFNST to each partition block will be described in detail.
[0287] According to an embodiment, after applying forward LFNST to each partition block, only a maximum of 16 (8 or 16) coefficients are left in the upper left 4×4 area in the transform coefficient scanning order, and then zeroing can be applied, where the remaining positions and areas are all filled with 0.
[0288] Alternatively, according to an embodiment, when the length of one side of the partition block is 4, LFNST is applied only to the upper left 4×4 region, and when the lengths of all sides of the partition block (i.e., width and height) are 8 or greater, LFNST can be applied to the remaining 48 coefficients within the upper left 8×8 region except for the lower right 4×4 region.
[0289] Alternatively, according to an embodiment, in order to adjust the worst-case computational complexity to 8 multiplications per sample, when each partition block is 4×4 or 8×8, only 8 transform coefficients may be output after applying the forward LFNST. That is, when the partition block is 4×4, an 8×16 matrix may be applied as the transform matrix, and when the partition block is 8×8, an 8×48 matrix may be applied as the transform matrix.
[0290] In the current VVC standard, LFNST index signaling is performed in units of coding units. Therefore, in ISP prediction mode and when LFNST is applied to all partition blocks, the same LFNST index value can be applied to the corresponding partition blocks. That is, when the LFNST index value is sent once at the coding unit level, the corresponding LFNST index can be applied to all partition blocks in the coding unit. As described above, the LFNST index value can have values of 0, 1, and 2, where 0 indicates the case where LFNST is not applied, and 1 and 2 indicate two transform matrices present in one LFNST set when LFNST is applied.
[0291] As described above, the LFNST set is determined by the intra prediction mode, and in the case of the ISP prediction mode, since all partition blocks in the coding unit are predicted in the same intra prediction mode, the partition blocks can refer to the same LFNST set.
[0292] As another example, LFNST index signaling is still performed in units of coding units, but in the case of ISP prediction mode, it is not determined whether LFNST is applied uniformly to all partition blocks, and for each partition block, whether to apply the LFNST index value signaled at the coding unit level and whether to apply LFNST can be determined by a separate condition. Here, a separate condition can be signaled in the form of a flag for each partition block through the bitstream, and when the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and when the flag value is 0, LFNST may not be applied.
[0293] In a coding unit to which the ISP mode is applied, an example of applying LFNST when the length of one side of a partition block is less than 4 is described as follows.
[0294] First, when the size of the partition block is N×2 (2×N), LFNST can be applied to the upper left M×2 (2×M) region (where M≤N). For example, when M=8, the upper left region becomes 8×2 (2×8), so the region with 16 residual signals can be the input of the forward LFNST, and an R×16 (R≤16) forward transform matrix can be applied.
[0295] Here, the forward LFNST matrix can be a separate additional matrix in addition to the matrix included in the current VVC standard. In addition, for worst-case complexity control, an 8×16 matrix in which only the upper 8 rows of the 16×16 matrix are sampled can be used for the transformation. The complexity control method will be described in detail later.
[0296] Secondly, when the size of the partition block is N×1 (1×N), LFNST can be applied to the upper left M×1 (1×M) region (where M≤N). For example, when M=16, the upper left region becomes 16×1 (1×16), so the region with 16 residual signals can be the input of the forward LFNST, and an R×16 (R≤16) forward transformation matrix can be applied.
[0297] Here, the corresponding forward LFNST matrix can be a separate additional matrix in addition to the matrix included in the current VVC standard. In addition, in order to control the complexity of the worst case, an 8×16 matrix in which only the upper 8 row vectors of the 16×16 matrix are sampled can be used for the transformation. The complexity control method will be described in detail later.
[0298] The first and second embodiments can be applied simultaneously, or either embodiment can be applied. In particular, in the case of the second embodiment, because a single transformation is considered in LFNST, it was experimentally observed that the compression performance improvement that can be obtained in existing LFNST is relatively small compared to the LFNST index signaling cost. However, in the case of the first embodiment, compression performance improvement similar to that obtained with conventional LFNST was observed. That is, in the case of ISP, the contribution of the application of 2×N and N×2 LFNST to actual compression performance can be examined through experiments.
[0299] In the current VVC's LFNST, symmetry between intra-frame prediction modes is applied. The same LFNST set is applied to the two directional modes arranged around mode 34 (prediction in the 45-degree diagonal direction at the lower right corner). For example, the same LFNST set is applied to mode 18 (horizontal prediction mode) and mode 50 (vertical prediction mode). However, in modes 35 to 66, when applying forward LFNST, the input data is transposed and then LFNST is applied.
[0300] VVC supports wide-angle intra prediction (WAIP) mode, and taking into account the WAIP mode, the LFNST set is derived based on the modified intra prediction mode. For the modes extended by WAIP, the LFNST set is determined by using symmetry, just like in the general intra prediction direction mode. For example, because mode-1 is symmetric with mode 67, the same LFNST set is applied, and because mode-14 is symmetric with mode 80, the same LFNST set is applied. Modes 67 to 80 apply the LFNST transform after transposing the input data before applying the forward LFNST.
[0301] When LFNST is applied to the upper left M×2 (M×1) block, since the block to which LFNST is applied is non-square, symmetry to LFNST cannot be applied. Therefore, instead of applying symmetry based on the intra prediction mode, as in the LFNST of Table 2, symmetry between the M×2 (M×1) block and the 2×M (1×M) block can be applied.
[0302] Figure 18 is a diagram illustrating symmetry between an M×2 (M×1) block and a 2×M (1×M) block according to an embodiment.
[0303] like Figure 18 As shown, since mode 2 in an M×2(M×1) block can be considered symmetric to mode 66 in a 2×M(1×M) block, the same LFNST set can be applied to both 2×M(1×M) blocks and M×2(M×1) blocks.
[0304] In this case, in order to apply the LFNST set applied to the M×2 (M×1) block to the 2×M (1×M) block, the LFNST set is selected based on mode 2 instead of mode 66. That is, before applying the forward LFNST, the LFNST may be applied after transposing the input data of the 2×M (1×M) block.
[0305] Figure 19 is a diagram illustrating an example of transposing a 2×M block according to an embodiment.
[0306] Figure 19 (a) is a diagram illustrating that LFNST can be applied by reading 2×M blocks of input data in column-major order, Figure 19 (b) is a diagram illustrating that LFNST can be applied by reading input data of an M×2 (M×1) block in row-major order. A method of applying LFNST to the upper left M×2 (M×1) or 2×M (M×1) block is described below.
[0307] 1. First, if Figure 19 As shown in (a) and (b), the input data is arranged to form the input vector of the forward LFNST. Figure 18 , for an M×2 block predicted in mode 2, follow Figure 19 For a 2×M block predicted in mode 66, the input data is in the order of Figure 19 The sequential arrangement of (a) can then apply the LFNST set for mode 2.
[0308] 2. For M×2 (M×1) blocks, considering WAIP, determine the LFNST set based on the modified intra prediction mode. As mentioned above, a preset mapping relationship is established between the intra prediction mode and the LFNST set, which can be represented by the mapping table shown in Table 2.
[0309] For a 2×M (1×M) block, a symmetric pattern around the prediction mode (mode 34 in the case of the VVC standard) in a downward 45-degree diagonal direction from the modified intra prediction mode in consideration of WAIP can be obtained, and then the LFNST set is determined based on the corresponding symmetric pattern and mapping table. The symmetric pattern (y) around mode 34 can be derived as follows. The mapping table will be described in more detail below.
[0310] [Equation 11]
[0311] If 2≤x≤66, then y=68-x,
[0312] Otherwise (x≤-1 or x≥67), y=66-x
[0313] 3. When forward LFNST is applied, transform coefficients may be derived by multiplying the input data prepared in process 1 by the LFNST kernel. The LFNST kernel may be selected according to the LFNST set determined in process 2 and a predetermined LFNST index.
[0314] For example, when M=8 and a 16×16 matrix is applied as the LFNST kernel, 16 transform coefficients can be generated by multiplying the matrix by 16 input data. The generated transform coefficients can be arranged in the upper left 8×2 or 2×8 area according to the scanning order used in the VVC standard.
[0315] Figure 20 The scanning order of 8×2 or 2×8 areas according to an embodiment is illustrated.
[0316] All areas except the upper left 8×2 or 2×8 area may be filled with zero values (cleared), or the existing transform coefficients to which one transform is applied may be left as is. The predetermined LFNST index may be one of the LFNST index values (0, 1, 2) attempted when calculating the RD cost while changing the LFNST index value in the programming process.
[0317] In the case of a configuration that adjusts the worst-case computational complexity to a certain degree or lower (e.g., 8 multiplications / sample), for example, after generating only 8 transform coefficients by multiplying an 8×16 matrix by taking only the upper 8 rows of the 16×16 matrix, the transform coefficients can be calculated as Figure 20 The scan order setting of , and zeroing can be applied to the remaining coefficient regions. The worst-case complexity control will be described later.
[0318] 4. When inverse LFNST is applied, a preset number (e.g., 16) of transform coefficients are set as an input vector, and the LFNST set obtained from process 2 and an LFNST kernel (e.g., a 16×16 matrix) derived from the selected parsed LFNST index are selected, and then an output vector can be derived by multiplying the LFNST kernel with the corresponding input vector.
[0319] In the case of M×2 (M×1) blocks, the output vector can be expressed as Figure 19 (b) is set in row-first order, and in the case of 2×M (1×M) blocks, the output vector can be Figure 19 (a) Column priority setting.
[0320] The remaining areas except the area where the corresponding output vector is set within the upper left M×2 (M×1) or 2×M (M×2) area and the areas except the upper left M×2 (M×1) or 2×M (M×2) area in the partition block (M×2 areas in the partition block) can all be cleared to have zero values, or can be configured to keep the reconstructed transform coefficients as is through residual encoding and inverse quantization processing.
[0321] When constructing the input vector, as in point 3, the input data can be Figure 20 The scanning order arrangement is carried out, and in order to control the computational complexity of the worst case to a certain extent or lower, the input vector can be constructed by reducing the number of input data (for example, 8 instead of 16).
[0322] For example, when M=8, if 8 input data are used, only the left 16×8 matrix can be taken from the corresponding 16×16 matrix and multiplied to obtain 16 output data. Worst-case complexity control will be described later.
[0323] In the above embodiment, when LFNST is applied, the case of applying symmetry between M×2 (M×1) blocks and 2×M (1×M) blocks is shown, but according to another example, a different LFNST set may be applied to each of the two block shapes.
[0324] Hereinafter, various examples of a mapping method using an intra prediction mode and an LFNST set configuration of an ISP mode will be described.
[0325] In the case of ISP mode, the LFNST set configuration may be different from the existing LFNST set. In other words, a core different from the existing LFNST core may be applied, and a mapping table different from the mapping table between intra prediction mode indexes and LFNST sets applied to the current VVC standard may be applied. The mapping table applied to the current VVC standard may be the same as the mapping table in Table 2.
[0326] In Table 2, the preModeIntra value indicates the intra prediction mode value changed in consideration of WAIP, and the lfnstTrSetIdx value is an index value indicating a specific LFNST set. Each LFNST set is configured with two LFNST cores.
[0327] When the ISP prediction mode is applied, if both the horizontal length and vertical length of each partition block are equal to or greater than 4, the same kernel as the LFNST kernel applied in the current VVC standard can be applied, and the mapping table can be applied as is. A mapping table and LFNST kernel different from the current VVC standard can be applied.
[0328] When the ISP prediction mode is applied, when the horizontal length or vertical length of each partition block is less than 4, a mapping table and LFNST kernel different from the current VVC standard may be applied. Below, Tables 6 to 8 show mapping tables between intra-frame prediction mode values (intra-frame prediction mode values changed in consideration of WAIP) and LFNST sets, which can be applied to M×2 (M×1) blocks or 2×M (1×M) blocks.
[0329] [Table 6]
[0330] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0
[0331] 2<=predModeIntra<=12 1 13<=predModeIntra<=23 2 24<=predModeIntra<=34 3 35<=predModeIntra<=44 4 45<=predModeIntra<=55 5 56<=predModeIntra<=66 6 67<=predModeIntra<=80 6 81<=predModeIntra<=83 0
[0332] [Table 7]
[0333] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0 2<=predModeIntra<=23 1 24<=predModeIntra<=44 2 45<=predModeIntra<=66 3 67<=predModeIntra<=80 3 81<=predModeIntra<=83 0
[0334] [Table 8]
[0335] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0 2<=predModeIntra<=80 1 81<=predModeIntra<=83 0
[0336] The first mapping table of Table 6 is configured with seven LFNST sets, the mapping table of Table 7 is configured with four LFNST sets, and the mapping table of Table 8 is configured with two LFNST sets. As another example, when it is configured with one LFNST set, the lfnstTrSetIdx value may be fixed to 0 with respect to the preModeIntra value.
[0337] Hereinafter, a method of maintaining the worst-case computational complexity when applying LFNST to the ISP mode will be described.
[0338] In the case of ISP mode, when LFNST is applied, in order to keep the number of multiplications per sample (or per coefficient, per position) to a certain value or less, the application of LFNST may be limited. Depending on the size of the partition block, the number of multiplications per sample (or per coefficient, per position) can be kept to 8 or less by applying LFNST as follows.
[0339] 1. When both the horizontal length and the vertical length of the partition block are 4 or greater, the same method as the worst-case computational complexity control method for LFNST in the current VVC standard can be applied.
[0340] That is, when the partition block is a 4×4 block, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix may be applied in the forward direction instead of the 16×16 matrix, and a 16×8 matrix obtained by sampling the left 8 columns from the 16×16 matrix may be applied in the reverse direction. Furthermore, when the partition block is an 8×8 block, in the forward direction, an 8×48 matrix obtained by sampling the upper 8 rows from a 16×48 matrix may be applied instead of the 16×48 matrix, and in the reverse direction, a 48×8 matrix obtained by sampling the left 8 columns from a 48×16 matrix may be applied instead of the 48×16 matrix.
[0341] In the case of a 4×N or N×4 (N>4) block, when performing forward transform, the 16 coefficients generated after applying the 16×16 matrix only to the upper left 4×4 block can be set in the upper left 4×4 area, and the other areas can be filled with a value of 0. In addition, when performing inverse transform, the 16 coefficients located in the upper left 4×4 block are arranged in scan order to form an input vector, and then 16 output data can be generated by multiplying the 16×16 matrix. The generated output data can be set in the upper left 4×4 area, and the remaining areas except the upper left 4×4 area can be filled with a value of 0.
[0342] In the case of an 8×N or N×8 (N>8) block, when performing forward transform, the 16×48 matrix is applied to the ROI region within only the upper left 8×8 block (except for the remaining regions other than the lower right 4×4 block in the upper left 8×8 block). The generated 16 coefficients can be set in the upper left 4×4 region, and all other regions can be filled with a value of 0. In addition, when performing inverse transform, the 16 coefficients located in the upper left 4×4 region are arranged in scan order to form an input vector, and then 48 output data can be generated by multiplying the 48×16 matrix. The generated output data can be filled in the ROI region, and all other regions can be filled with a value of 0.
[0343] 2. When the size of the partition block is N×2 or 2×N and LFNST is applied to the upper left M×2 or 2×M area (M≤N), a matrix sampled according to the N value can be applied.
[0344] In the case of M=8, for a partition block of N=8, i.e., an 8×2 or 2×8 block, in the case of forward transform, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix may be applied instead of a 16×16 matrix, and in the case of inverse transform, a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix may be applied instead of a 16×16 matrix.
[0345] When N is greater than 8, in the case of forward transform, the 16 output data generated after applying the 16×16 matrix to the upper left 8×2 or 2×8 block is set in the upper left 8×2 or 2×8 block, and the remaining area can be filled with a value of 0. In the case of inverse transform, the 16 coefficients located in the upper left 8×2 or 2×8 block are arranged in scan order to form an input vector, and then 16 output data can be generated by multiplying the 16×16 matrix. The generated output data can be placed in the upper left 8×2 or 2×8 block, and all remaining areas can be filled with a value of 0.
[0346] 3. When the size of the partition block is N×1 or 1×N and LFNST is applied to the upper left M×1 or 1×M area (M≤N), a matrix sampled according to the N value can be applied.
[0347] When M=16, for a partition block with N=16, i.e., a 16×1 or 1×16 block, in the case of forward transform, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix may be applied instead of a 16×16 matrix, and in the case of inverse transform, a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix may be applied instead of a 16×16 matrix.
[0348] When N is greater than 16, in the case of forward transform, 16 output data generated after applying the 16×16 matrix to the upper left 16×1 or 1×16 block can be set in the upper left 16×1 or 1×16 block, and the remaining area can be filled with a value of 0. In the case of inverse transform, the 16 coefficients located in the upper left 16×1 or 1×16 block can be arranged in scan order to form an input vector, and then 16 output data can be generated by multiplying the 16×16 matrix. The generated output data can be set in the upper left 16×1 or 1×16 block, and all remaining areas can be filled with a value of 0.
[0349] As another example, in order to keep the number of multiplications for each sample (or each coefficient, each position) at a certain value or less, the number of multiplications for each sample (or each coefficient, each position) based on the ISP coding unit size rather than the size of the ISP partition block can be kept at 8 or less. When only one block among the ISP partition blocks meets the conditions for applying LFNST, the worst-case complexity calculation of LFNST can be applied based on the corresponding coding unit size rather than the size of the partition block. For example, when the luminance coding block of a certain coding unit is divided into four partition blocks of size 4×4 and encoded by ISP, and there are no non-zero transform coefficients for two of the four partition blocks, it can be configured so that instead of eight transform coefficients, 16 transform coefficients are generated for the other two partition blocks (based on the encoder).
[0350] Hereinafter, a method of signaling an LFNST index in the ISP mode will be described.
[0351] As described above, the LFNST index can have any value of 0, 1, or 2, where a value of 0 indicates that LFNST is not applied, and values 1 and 2 indicate the two LFNST kernel matrices included in the selected LFNST set, respectively. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The method for transmitting the LFNST index in the current VVC standard will be described below.
[0352] 1. LFNST index can be sent once for each coding unit (CU), and in the case of dual-tree type, separate LFNST indexes can be signaled for luma blocks and chroma blocks.
[0353] 2. When the LFNST index is not signaled, the value of the LFNST index is set (inferred) to a default value of 0. The case where the value of the LFNST index is inferred to be 0 is as follows:
[0354] A. A case where a mode that does not apply transform (eg, transform skip, BDPCM, lossless coding, etc.) is used.
[0355] B. A case where one transform is not DCT-2 (DCT7 or DCT8), that is, a case where the transform in the horizontal direction or the transform in the vertical direction is not DCT-2.
[0356] C. When the horizontal length or vertical length of the luma block of a coding unit exceeds the maximum luma transformable size, for example, when the maximum luma transform size is 64, LFNST is not applied when the luma block size of the coding block is 128×16.
[0357] In the case of the dual-tree type, for each of the coding units of the luma component and the chroma component, it is determined whether the maximum luma transformable size is exceeded. That is, for the luma block, it is checked whether the maximum luma transformable size is exceeded, and for the chroma block, it is checked whether the horizontal / vertical length and the maximum luma transformable size of the corresponding luma block for the color format are exceeded. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luma block is twice the horizontal / vertical length of the chroma block, and the transform size of the corresponding luma block is twice the transform size of the chroma block. In another example, when the color format is 4:4:4, the horizontal / vertical length and transform size of the corresponding luma block are the same as those of the chroma block.
[0358] A 64-length transform or a 32-length transform refers to a transform applied horizontally or vertically with a length of 64 or 32, respectively, and the “transform size” may refer to the corresponding length 64 or 32.
[0359] In the case of a single tree type, it may be checked whether the horizontal length or vertical length of the luminance block exceeds the maximum transformable size of the luminance transform block, and when the horizontal length or vertical length of the luminance block exceeds the maximum transformable size of the luminance transform block, signaling of the LFNST index may be omitted.
[0360] D. The LFNST index may be transmitted only when both the horizontal length and the vertical length of the coding unit are greater than or equal to 4.
[0361] In case of the dual-tree type, the LFNST index may be signaled only when both the horizontal length and the vertical length of the corresponding component (ie, the luma component or the chroma component) are each greater than or equal to 4.
[0362] In case of a single tree type, the LFNST index may be signaled when both the horizontal length and the vertical length of the luma component are greater than or equal to 4, respectively.
[0363] E. If the last non-zero coefficient position is not the DC position (which is the position located at the upper left of the block), if the last non-zero coefficient position is not the DC position for the dual-tree type luma block, the LFNST index is transmitted. In the case of a dual-tree type chroma block, if either the last non-zero coefficient position of Cb or the last non-zero coefficient position of Cr is not the DC position, the corresponding LNFST index is transmitted.
[0364] In case of a single tree type, when the last non-zero coefficient position of any one of the luma component, the Cb component, and the Cr component is not the DC position, the LFNST index is transmitted.
[0365] Here, when the coded block flag (CBF) value indicating whether a transform coefficient of a transform block exists is 0, the position of the last non-zero coefficient of the corresponding transform block is not checked to determine whether to signal the LFNST index. That is, when the corresponding CBF value is 0, the transform is not applied to the corresponding block, and therefore the position of the last non-zero coefficient may not be considered when checking the conditions for LFNST index signaling.
[0366] For example, 1) in the case of a dual-tree type luma component, if the corresponding CBF value is 0, the LFNST index is not signaled; 2) in the case of a dual-tree type chroma component, if the CBF value of Cb is 0 and the CBF value of Cr is 1, only the last non-zero coefficient position of Cr is checked and the corresponding LFNST index is sent; 3) in the case of a single-tree type, the last non-zero coefficient position of any component whose CBF value is 1 among the luma component, Cb component, and Cr component is checked.
[0367] F. When a transform coefficient is found to exist at a position other than the position where the LFNST transform coefficient is allowed to exist, the signaling of the LFNST index can be omitted. In the case of 4×4 transform blocks and 8×8 transform blocks, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient can exist in 8 positions starting from the DC position and all the remaining positions are filled with 0. In addition, in cases other than 4×4 transform blocks and 8×8 transform blocks, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient can exist in 16 positions starting from the DC position and the remaining positions are all filled with 0.
[0368] Therefore, when a non-zero transform coefficient exists in a region that should be filled with 0 after residual encoding is performed, LFNST index signaling may be omitted.
[0369] In addition, the ISP mode can be applied only to the luminance block, or to both the luminance block and the chrominance block. As described above, when ISP prediction is applied, the corresponding coding unit is divided into two or four partition blocks and predicted, and a transform can be applied to each corresponding partition block. Therefore, when determining the conditions for signaling the LFNST index in units of coding units, it is necessary to consider the fact that LFNST is applicable to each of the corresponding partition blocks. In addition, when the ISP prediction mode is applied only to a specific component (e.g., a luminance block), the LFNST index should be signaled taking into account the fact that only the corresponding component is divided into partition blocks. The LFNST index signaling methods available in ISP mode can be summarized as follows.
[0370] 1. The LFNST index may be transmitted once for each coding unit (CU), and in the case of a dual-tree type, separate LFNST indexes may be signaled for luma blocks and chroma blocks, respectively.
[0371] 2. When the LFNST index is signaled, the value of the LFNST index is set (inferred) to a default value of 0. The case where the LFNST index value is inferred to be 0 is as follows.
[0372] A. A case where a mode that does not apply transform (eg, transform skip, BDPCM, lossless coding, etc.) is used.
[0373] B. When the horizontal length or vertical length of the luminance block of the coding unit exceeds the maximum luminance transformable size, for example, when the maximum transformable luminance size is 64, LFNST cannot be applied when the size of the luminance block of the coding block is 128×16.
[0374] Whether to signal the LFNST index can be determined based on the size of the partition block rather than the size of the coding unit. That is, when the horizontal length or vertical length of the partition block corresponding to the luma block exceeds the maximum luma transformable size, LFNST index signaling can be omitted and the value of the LFNST index can be inferred to be 0.
[0375] In the case of the dual tree type, it is determined whether the maximum transform block size is exceeded for each of the coding unit or partition block of the luma component and the coding unit or partition block of the chroma component. That is, the horizontal length and vertical length of the coding unit or partition block for luma are compared with the maximum luma transformable size, and if either of the horizontal length and vertical length is greater than the maximum luma transformable size, LFNST is not applied, and in the case of the coding unit or partition block for chroma, the horizontal / vertical length of the corresponding luma block for the color format is compared with the maximum luma transformable size. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luma block is twice the horizontal / vertical length of the chroma block, and the transform size of the corresponding luma block is twice the transform size of the chroma block. In another example, when the color format is 4:4:4, the horizontal / vertical length and transform size of the corresponding luma block are the same as those of the chroma block.
[0376] In case of a single tree type, it may be checked whether the horizontal length or vertical length of a luma block (coding unit or partition block) exceeds the maximum transformable size of the luma transform block, and if so, LFNST index signaling may be omitted.
[0377] C. When LFNST included in the current VVC standard is applied, the LFNST index may be transmitted only when both the horizontal length and the vertical length of the partition block are greater than or equal to 4.
[0378] When LFNST for 2×M (1×M) or M×2 (M×1) blocks is applied in addition to LFNST included in the current VVC standard, the LFNST index may be transmitted only when the size of the partition block is equal to or larger than 2×M (1×M) or M×2 (M×1) blocks. Here, if a P×Q block is equal to or larger than an R×S block, this means P ≥ R and Q ≥ S.
[0379] In summary, the LFNST index can be sent only when the partition block is equal to or larger than the minimum size for which LFNST can be applied. In the case of a dual-tree type, the LFNST index can be signaled only when the partition block of the luma or chroma component is equal to or larger than the minimum size for which LFNST can be applied. In the case of a single-tree type, the LFNST index can be signaled only when the partition block of the luma component is equal to or larger than the minimum size for which LFNST can be applied.
[0380] In this document, if an M×N block is equal to or larger than a K×L block, this means M is equal to or larger than K and N is equal to or larger than L. If an M×N block is larger than a K×L block, this means M is equal to or larger than K, N is equal to or larger than L, and either M is larger than K or N is larger than L. If an M×N block is smaller than or equal to a K×L block, this means M is smaller than or equal to K and N is smaller than or equal to L, and if an M×N block is smaller than a K×L block, this means M is smaller than or equal to K, N is smaller than or equal to L, and either M is smaller than K or N is smaller than L.
[0381] D. In the case where the last non-zero coefficient position is not the DC position (which is the position located at the upper left of the block), if the last non-zero coefficient position in any one of all partition blocks is not the DC position for the dual-tree type luminance block, the LFNST index may be transmitted. In the case of the dual-tree type chrominance block, if any of the last non-zero coefficient positions of all partition blocks for Cb (when the ISP mode is not applied to the chrominance component, the number of partition blocks is considered to be 1) and the last non-zero coefficient position of all partition blocks for Cr (when the ISP mode is not applied to the chrominance component, the number of partition blocks is considered to be 1) is not the DC position, the corresponding LNFST index may be transmitted.
[0382] In case of a single tree type, when the last non-zero coefficient position of any one of all partition blocks of luma components, Cb components, and Cr components is not a DC position, a corresponding LFNST index may be transmitted.
[0383] Here, when the coded block flag (CBF) value indicating whether a transform coefficient exists for each partition block is 0, the last non-zero coefficient position of the corresponding partition block is not checked to determine whether to signal the LFNST index. That is, when the corresponding CBF value is 0, the transform is not applied to the corresponding block, and therefore the last non-zero coefficient position of the corresponding partition block is not considered when checking the conditions for LFNST index signaling.
[0384] For example, 1) in the case of the luminance component of the dual-tree type, when the corresponding CBF value of each partition block is 0, the corresponding partition block is excluded when determining whether to signal the LFNST index; 2) in the case of the chrominance component of the dual-tree type, when the CBF value of Cb of each partition block is 0 and the CBF value of Cr is 1, only the last non-zero coefficient position of Cr is checked when determining whether to signal the LFNST index; and 3) in the case of the single-tree type, whether to signal the LFNST index can be determined by checking only the last non-zero coefficient position of the block with a CBF value of 1 for all partition blocks of the luminance component, Cb component, and Cr component.
[0385] In the case of the ISP mode, image information may be configured such that the last non-zero coefficient position is not checked, and its implementation is as follows.
[0386] i. In ISP mode, LFNST index signaling can be allowed without checking the last non-zero coefficient position of both the luma block and the chroma block. That is, even when the last non-zero coefficient position of all partition blocks is at the DC position or the corresponding CBF value is 0, the corresponding LFNST index signaling can be allowed.
[0387] ii. In the case of ISP mode, the check of the last non-zero coefficient position can be omitted only for the luma block, and the last non-zero coefficient position of the chroma block can be checked in the above manner. For example, in the case of a dual-tree type luma block, LFNST index signaling is allowed without checking the last non-zero coefficient position, while in the case of a dual-tree type chroma block, whether to signal the corresponding LFNST index can be determined by checking whether the DC position of the last non-zero coefficient position exists in the above manner.
[0388] iii. In the case of ISP mode and single tree type, the above method i or the above method ii can be applied. That is, when method i is applied to the single tree type in ISP mode, the check of the last non-zero coefficient position can be omitted for both the luminance block and the chrominance block, and LFNST index signaling is allowed. Alternatively, when method ii is applied, the check of the last non-zero coefficient position can be omitted for the partition block of the luminance component, and for the partition block of the chrominance component (when ISP is not applied to the chrominance component, the number can be considered to be 1), the last non-zero coefficient position can be checked in the above manner to determine whether to signal the corresponding LFNST index.
[0389] E. When a transform coefficient is found to exist at a position other than a position where the LFNST transform coefficient may exist for even one partition block among all partition blocks, LFNST index signaling may be omitted.
[0390] For example, in the case of a 4×4 partition block and an 8×8 partition block, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient may be present in 8 positions starting from the DC position, and all remaining positions are filled with 0. In addition, when the partition block is equal to or larger than 4×4 and is neither a 4×4 partition block nor an 8×8 partition block, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient may be present in 16 positions starting from the DC position, and all remaining positions are filled with 0.
[0391] Therefore, when there are non-zero transform coefficients in a region to which 0 padding is applied after residual encoding is performed, LFNST index signaling may be omitted.
[0392] When LFNST is applicable even for 2×M (1×M) or M×2 (M×1) partition blocks, the region where the LFNST transform coefficient is allowed to be located can be specified as follows. The region outside the region where the transform coefficient is allowed to be located can be padded with 0, and under the assumption that LFNST is applied, when there is a non-zero transform coefficient in the region that should be padded with 0, LFNST index signaling can be omitted.
[0393] i. When LFNST is applicable to 2×M or M×2 blocks and M=8, only 8 LFNST transform coefficients can be generated for 2×8 or 8×2 partition blocks. Figure 20 When arranged in the scan order shown, the 8 transform coefficients may be arranged in scan order starting from the DC position, and the remaining 8 positions may be filled with zeros.
[0394] For a 2×N or N×2 (N>8) partition block, 16 LFNST transform coefficients can be generated, and when the transform coefficients are expressed as Figure 20When the scan order is arranged as shown, the 16 transform coefficients can be arranged in a scan order starting from the DC position, and the remaining area can be filled with 0. That is, the area other than the upper left 2×8 or 8×2 block in the 2×N or N×2 (N>8) partition block can be filled with 0. Even for a 2×8 or 8×2 partition block, 16 transform coefficients can be generated instead of 8 LFNST transform coefficients, and in this case, there will be no area that must be filled with 0. As described above, in the case of applying LFNST, when it is detected that there is a non-zero transform coefficient in the area set to be filled with 0 even for one partition block, the LFNST index signaling can be omitted and the LFNST index can be inferred to be 0.
[0395] ii. When LFNST is applicable to 1×M or M×1 blocks and M=16, only 8 LFNST transform coefficients may be generated for a 1×16 or 16×1 partitioned block. When the transform coefficients are arranged in a left-to-right or top-to-bottom scan order, the 8 transform coefficients may be arranged starting from the DC position in the corresponding scan order, and the remaining 8 positions may be padded with 0s.
[0396] For a 1×N or N×1 (N>16) partition block, 16 LFNST transform coefficients may be generated, and when the transform coefficients are arranged in a left-to-right or top-to-bottom scan order, the 16 transform coefficients may be arranged in the corresponding scan order starting from the DC position, and the remaining area may be padded with 0. That is, the area other than the top-left 1×16 or 16×1 block in the 1×N or N×1 (N>16) partition block may be padded with 0.
[0397] Even for a 1×16 or 16×1 partition block, 16 transform coefficients can be generated instead of 8 LFNST transform coefficients, and in this case, there is no area that must be padded with 0. As described above, when LFNST is applied, when it is detected that there is a non-zero transform coefficient in an area set to be padded with 0 for even one partition block, LFNST index signaling can be omitted and the LFNST index can be inferred to be 0.
[0398] Furthermore, in ISP mode, according to the current VVC standard, the horizontal and vertical directions are independently considered length conditions, and DST-7 is applied instead of DCT-2 without signaling the MTS index. A determination is made as to whether the horizontal or vertical length is greater than or equal to 4 or greater than or equal to 16, and a primary transform kernel is determined based on the determination result. Therefore, for cases where LFNST can be applied in ISP mode, the following transform combinations are possible.
[0399] 1. When the LFNST index is 0 (including the case where the LFNST index is inferred to be 0), the one-time transform determination condition in the ISP mode can be followed, which is included in the current VVC standard. That is, if the length condition (which is equal to or greater than 4 and equal to or less than 16) is satisfied separately and independently for the horizontal direction and the vertical direction; if so, DST-7 can be applied instead of DCT-2, and if not, DCT-2 can be applied.
[0400] 2. When the LFNST index is greater than 0, the following two configurations can be used as a single transformation.
[0401] A. DCT-2 can be applied to both horizontal and vertical directions.
[0402] B. The conditions for determining a primary transform in ISP mode, which are included in the current VVC standard, may be followed. In other words, the horizontal and vertical directions are checked to see whether the length condition (which is equal to or greater than 4 and equal to or less than 16) is satisfied separately and independently; if so, DST-7 may be applied instead of DCT-2, and if not, DCT-2 may be applied.
[0403] In ISP mode, image information may be configured so that an LFNST index is transmitted for each partition block rather than for each coding unit. In this case, in the above-described LFNST index signaling method, whether to signal the LFNST index may be determined considering that there is only one partition block in the unit in which the LFNST index is transmitted.
[0404] Furthermore, according to one embodiment, in the ISP mode as shown in Table 9, an LFNST core consisting of three groups may be applied.
[0405] In cases other than the ISP mode, a mapping table as shown in the following table may be applied, and even when LFNST is applied in the ISP mode, LFNST set as shown in Table 9 may be applied.
[0406] [Table 9]
[0407] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0 2<=predModeIntra<=18 1 19<=predModeIntra<=49 3 50<=predModeIntra<=80 1 81<=predModeIntra<=83 0
[0408] The value of preModeIntra in Table 9 indicates the intra prediction mode value after WAIP has been changed, and the value of lfnstTrSetIdx is an index value indicating a specific LFNST set, and each LFNST set consists of two LFNST kernels. Each index (lfnstTrSetIdx) can specify a corresponding LFNST set in the VVC standard document (JVET-O2001-vE.docx) (which refers to the LFNST set specified by the lfnstTrSetIdx variable in the "Low frequency non-separable transformation matrix derivation process" section in the corresponding document).
[0409] In the case of applying LFNST in the ISP mode, the mapping table shown in Table 9 may be applied only when both the horizontal length and the vertical length of the ISP partition block are greater than 4.
[0410] The implementation of applying LFNST in the above ISP mode is summarized as follows.
[0411] (1) When LFNST is applied in ISP mode, the transform unit to be divided must have a size of at least 4×4.
[0412] (2) The same LFNST core as the existing LFNST core applied to the coding unit to which the ISP mode is not applied may be used.
[0413] (3) Each transform unit must satisfy the maximum last position value condition (which is a condition for the last non-zero coefficient position). When one or more transform units do not satisfy the maximum last position value condition, LFNST is not used and the LFNST index is not parsed.
[0414] (4) When the ISP mode is applied, the setting that the LFNST is applicable only when there are valid coefficients at positions other than the DC position may be ignored.
[0415] (5) When LFNST is applied, DCT-2 can be used for one transform of the transform unit to which ISP is applied.
[0416] Table 10 below shows syntax elements including the above content.
[0417] [Table 10]
[0418]
[0419] Table 10 sets the width and height of the area where LFNST is applied according to the tree type, and shows the conditions for transmitting the LFNST index. The syntax elements of Table 10 can be signaled at the coding unit (CU) level. In the case of a dual tree type, separate LFNST indices for luma blocks and chroma blocks can be signaled.
[0420] First, when the tree type of a coding unit is dual-tree chroma, the width of the region where LFNST is applied (lfnstWidth) can be set to reflect the width of the color format in the coding unit ((treeType==DUAL_TREE_CHROMA)·cbWidth / SubWidthC).
[0421] On the other hand, when the tree type of the coding unit is not dual-tree chroma but dual-tree or single-tree luminance, the width of the area to which LFNST is applied (lfnstWidth) can be set to a value obtained by dividing the coding unit by the number of sub-partitions or the width of the coding unit ((IntraSubPartitionsSplitType==ISP_VER_SPLIT)?cbWidth / NumIntraSubPartitions:cbWidth) according to whether the coding unit is ISP-split. That is, if the coding unit is vertically split by the ISP (IntraSubPartitionsSplitType==ISP_VER_SPLIT), the width of the area to which LFNST is applied can be set to a value obtained by dividing the coding unit by the number of sub-partitions (cbWidth / NumIntraSubPartitions); if not, the width of the area to which LFNST is applied can be set to the width of the coding unit (cbWidth).
[0422] Similarly, when the tree type of the coding unit is dual-tree chroma, the height of the area where LFNST is applied (lfnstHeight) can be set to reflect the height of the color format in the height of the coding unit ((treeType==DUAL_TREE_CHROMA)?cbHeight / SubHeightC).
[0423] On the other hand, when the tree type of the coding unit is not dual-tree chroma but dual-tree or single-tree luminance, the height of the area to which LFNST is applied (lfnstHeight) can be set to the value obtained by dividing the coding unit by the number of sub-partitions or the height of the coding unit ((IntraSubPartitionsSplitType==ISP_HOR_SPLIT)?cbHeight / NumIntraSubPartitions:cbHeight) according to whether the coding unit is split by the ISP. That is, if the coding unit is split in the horizontal direction by the ISP (IntraSubPartitionsSplitType==ISP_HOR_SPLIT), the height of the area to which LFNST is applied can be set to the value obtained by dividing the coding unit by the number of sub-partitions (cbHeight / NumIntraSubPartitions); if not, it can be set to the height of the coding unit (cbHeight).
[0424] In order to apply LFNST in this manner, the width and height of the area to which LFNST is applied must be greater than or equal to 4. That is, in the case of encoding a dual-tree type, the LFNST index may be signaled only when both the horizontal length and the vertical length of the corresponding component (ie, the luma or chroma component) are greater than or equal to 4, and in the case of a single-tree type, the LFNST index may be signaled when both the horizontal length and the vertical length of the luma component are greater than or equal to 4.
[0425] When ISP is applied to a coding unit, the LFNST index may be transmitted only when both the horizontal length and the vertical length of the partition block are greater than or equal to 4.
[0426] In addition, when the horizontal length or vertical length of the luma block of the coding unit exceeds the maximum luma block transformable size (when the condition of Max(cbWidth, cbHeight)<=MaxTbSizeY is not satisfied), LFNST may be applied and the LFNST index may not be transmitted.
[0427] Furthermore, the LFNST index may be signaled only when the last non-zero coefficient position is not the DC position (which is a position located at the top left of the block).
[0428] In the case of a dual-tree type luma block, when the last non-zero coefficient position is not the DC position, an LFNST index is transmitted. In the case of a dual-tree type chroma block, when either the last non-zero coefficient position of Cb or the last non-zero coefficient position of Cr is not the DC position, the corresponding LFNST index is transmitted. In the case of a single-tree type, when the last non-zero coefficient position of any one of the luma component, Cb component, and Cr component is not the DC position, an LFNST index may be transmitted.
[0429] In addition, when ISP is applied to a coding unit, the LFNST index can be signaled without checking the last non-zero coefficient position (IntraSubPartitionsSplitType!=ISP_NO_SPLIT||LfnstDcOnly==0). That is, even when the last non-zero coefficient position of all partition blocks is at the DC position, LFNST index signaling can be allowed. The DC position indicates the position at the upper left of the corresponding block.
[0430] Finally, when a transform coefficient is found to exist at a position other than a position where the LFNST transform coefficient is allowed to exist, LFNST index signaling may be omitted (LfnstZeroOutSigCoeffFlag==1).
[0431] In the case of applying ISP to a coding unit, when it is confirmed that a transform coefficient exists at a position other than a position where the LFNST transform coefficient is assumed to exist even for one partition block among all partition blocks, LFNST index signaling may be omitted.
[0432] The following figures are provided to describe specific examples of the present disclosure. Since the specific names of the devices or the names of the specific signals / messages / fields shown in the figures are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following figures.
[0433] Figure 21 is a flowchart illustrating the operation of a video decoding device according to an embodiment of this document.
[0434] Figure 21 Each step disclosed in Figures 4 to 20 Therefore, some of the contents in the above description will be omitted or simplified. Figures 3 to 20 The same detailed description as above.
[0435] The decoding apparatus 300 according to the embodiment may receive residual information from a bitstream ( S2110 ).
[0436] More specifically, the decoding device 300 can decode information about the quantized transform coefficients of the target block from the bitstream, and can derive the quantized transform coefficients of the current block based on the information about the quantized transform coefficients of the current block. The information about the quantized transform coefficients of the target block may be included in a sequence parameter set (SPS) or a slice header, and may include information about whether a reduced transform (RST) is applied, information about a simplification factor, information about a minimum transform size for applying the reduced transform, information about a maximum transform size for applying the reduced transform, a reduced inverse transform size, and at least one of information indicating a transform index of any one transform kernel matrix included in a transform set.
[0437] In addition, the decoding device may also receive information regarding the intra-frame prediction mode of the current block and information regarding whether ISP coding is applied to the current block. The decoding device may infer whether the current block is divided into a predetermined number of sub-partition blocks by receiving and decoding flag information indicating whether ISP coding or IPS mode is applied. Here, the current block may be a coding block. Furthermore, the decoding device may infer the size and number of sub-partition blocks divided by flag information indicating the direction in which the current block is divided.
[0438] For example, Figure 16 As shown, when the size (width×height) of the current block is 8×4, the current block can be vertically divided and divided into two sub-blocks, and when the size (width×height) of the current block is 4×8, the current block can be horizontally divided and divided into two sub-blocks. Alternatively, as Figure 17 As shown, when the size (width × height) of the current block is greater than 4×8 or 8×4, that is, when the size of the current block is: 1) 4×N or N×4 (N≥16) or 2) M×N (M≥8, N≥8), the current block can be divided into 4 sub-blocks in the horizontal direction or the vertical direction.
[0439] The same intra prediction mode can be applied to the sub-partition blocks divided from the current block, and the decoding device can derive prediction samples for the corresponding sub-partition blocks. That is, the decoding device performs intra prediction sequentially from left to right or from top to bottom (e.g., horizontally or vertically) according to the form in which the sub-partition blocks are divided. For the leftmost or topmost sub-block, the reconstructed pixels of the already encoded coding block are referenced as in the conventional intra prediction method. In addition, when each side of the subsequent internal sub-partition block is not adjacent to the previous sub-partition block, the reconstructed pixels of the previously encoded adjacent coding block are referenced as in the conventional intra prediction method to derive the reference pixels adjacent to the corresponding side.
[0440] The decoding apparatus 300 may induce a transform coefficient by dequantizing residual information (ie, a quantized transform coefficient) about the current block ( S2120 ).
[0441] The derived transform coefficients can be arranged in a reverse diagonal scan order in units of 4×4 blocks, and the transform coefficients in a 4×4 block can also be arranged in a reverse diagonal scan order. That is, the dequantized transform coefficients can be arranged in a reverse scan order used in a video codec such as VVC or HEVC.
[0442] The transform coefficient derived based on the residual information may be a dequantized transform coefficient as described above, or may be a quantized transform coefficient. That is, the transform coefficient may be data that can be checked whether it is non-zero data in the current block regardless of quantization.
[0443] In order to derive the modified transform coefficient by applying LFNST, the decoding apparatus may derive a first variable indicating whether a transform coefficient (ie, a significant coefficient) exists in a region other than a DC position of the current block ( S2130 ).
[0444] The first variable may be a variable LfnstDcOnly that can be derived during the residual coding process. When the index of the subblock including the last significant coefficient in the current block is 0 and the position of the last significant coefficient in the subblock is greater than 0, the first variable may be derived as 0, and when the first variable is 0, the LFNST index may be parsed. A subblock refers to a unit block used as a coding unit in residual coding, and when the current block is greater than or equal to 4×4, the corresponding unit block is a 4×4 block and may be referred to as a coefficient group (CG). When the current block is greater than or equal to 4×4, a subblock index of 0 indicates a 4×4 subblock on the upper left.
[0445] The first variable may be initially set to 1, and may remain at 1 or may be changed to 0 depending on whether a significant coefficient exists in a region other than a DC position.
[0446] The variable LfnstDcOnly indicates whether there is a non-zero coefficient at a position other than the DC component for at least one transform block in one coding unit, and when there is a non-zero coefficient at a position other than the DC component for at least one transform block in one coding unit, the variable LfnstDcOnly may be 0, and when there is no non-zero coefficient at a position other than the DC component for all transform blocks in one coding unit, the variable LfnstDcOnly may be 1.
[0447] The decoding apparatus may parse the LFNST index based on the derivation result, and parse the LFNST index without deriving the first variable based on the current block being divided into a plurality of sub-partition blocks ( S2140 ).
[0448] According to an example, the decoding device can parse the LFNST index when the current block is not divided into a plurality of sub-partition blocks and the first variable indicates that there is a transform coefficient in an area other than the DC position, and when the first block is divided into a plurality of sub-partition blocks, the decoding device can parse the LFNST index without checking the first variable or ignoring the first variable value.
[0449] That is, when ISP is applied to the current block, LFNST index signaling may be allowed even if the last non-zero coefficient of each sub-partition block is located at the DC position.
[0450] In addition, according to an example, the decoding apparatus may also determine whether there is any transform coefficient in a second region other than the first region located at the upper left of the current block, and when there is no transform coefficient in the second region, may parse the LFNST index.
[0451] By deriving a second variable indicating whether a significant coefficient exists in a second region other than the first region located at the upper left of the current block, it may be checked whether clearing has been performed on the second region.
[0452] The second variable may be a variable LfnstZeroOutSigCoeffFlag, which indicates that clearing is performed when LFNST is applied. The second variable may be initially set to 1, and when there is a valid coefficient in the second region, the second variable may be changed to 0.
[0453] When the index of the sub-block where the last non-zero coefficient exists is greater than 0 and both the width and height of the transform block are equal to or greater than 4, or when the last position of the non-zero coefficient in the sub-block where the last non-zero coefficient exists is greater than 7 and the size of the transform block is 4×4 or 8×8, the variable LfnstZeroOutSigCoeffFlag can be derived to 0.
[0454] That is, when a non-zero coefficient is derived in an area other than the upper left area where the LFNST transform coefficient can exist, or when a non-zero coefficient exists other than the eighth position in the scanning order for 4×4 blocks and 8×8 blocks, the variable LfnstZeroOutSigCoeffFlag is set to 0.
[0455] According to an example, when ISP is applied to a coding unit, if a transform coefficient is found to exist in a position other than a position where the LFNST transform coefficient can exist for one partition block among all sub-partition blocks, LFNST index signaling can be omitted. That is, if zeroing is not performed on one sub-partition block and a significant coefficient exists in the second region, the LFNST index is not signaled.
[0456] Furthermore, the first region may be derived based on the size of the current block.
[0457] For example, when the size of the current block is 4×4 or 8×8, the first area may be an area from the upper left to the eighth sample position of the current block in the scanning direction. When the current block is split and the size of the sub-partition block is 4×4 or 8×8, the first area may be an area from the upper left to the eighth sample position of the sub-partition block in the scanning direction.
[0458] When the size of the current block is 4×4 or 8×8, 8 data are output by the forward LFNST, and thus the 8 transform coefficients received by the decoding device can be arranged from the upper left to the 8th sample position of the current block in the scanning direction, as shown in FIG. Figure 13 (a) and Figure 14 As shown in (a).
[0459] In addition, when the size of the current block is not 4×4 or 8×8, the first area may be a 4×4 area in the upper left of the current block. When the size of the current block is not 4×4 or 8×8, 16 data are output by the forward LFNST, and thus the 16 transform coefficients received by the decoding device may be arranged in the upper left 4×4 area in the current block, as shown in FIG. Figure 13 (b) to (d) and Figure 14 As shown in (b).
[0460] Furthermore, the transform coefficients that may be arranged in the first region may be arranged along a diagonal scan direction, such as Figure 8 shown.
[0461] As described above, when the current block is divided into sub-partition blocks, when no transform coefficient exists in all corresponding second regions of the plurality of sub-partition blocks, the decoding device may parse the LFNST index. When a transform coefficient exists in the second region of any sub-partition block, the LFNST index is not parsed.
[0462] As described above, LFNST may be applied to a sub-partition block in which a width and a height are greater than or equal to 4, and an LFNST index of a current block as a coding block may be applied to a plurality of sub-partition blocks.
[0463] In addition, since the clearing reflecting the LFNST (including all clearing that may accompany the application of the LFNST) is also applied to the sub-partitioned block, the first area is also applied to the sub-partitioned block. That is, when the divided sub-partitioned block is a 4×4 block or an 8×8 block, the LFNST is applied to the 8th transform coefficient in the scanning direction starting from the upper left position of the sub-partitioned block, and when the sub-partitioned block is not a 4×4 block or an 8×8 block, the LFNST may be applied to the transform coefficient in the upper left 4×4 area of the sub-partitioned block.
[0464] Then, the decoding apparatus may derive a modified transform coefficient from the transform coefficient based on the LFNST index and the LFNST matrix used for LFNST ( S2150 ).
[0465] The decoding apparatus may determine an LFNST set including an LFNST matrix based on the intra prediction mode derived from the intra prediction mode information, and select any one of a plurality of LFNST matrices based on the LFNST set and the LFNST index.
[0466] In this case, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks split from the current block. That is, because the same intra prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra prediction mode can be applied equally to all sub-partition transform blocks. In addition, because the LFNST index is signaled at the coding unit level, the same LFNST matrix can be applied to the sub-partition transform blocks split from the current block.
[0467] As described above, a transform set can be determined according to the intra prediction mode of the transform block to be transformed, and inverse LFNST can be performed based on any one of the transform kernel matrices (i.e., LFNST matrices) included in the transform set indicated by the LFNST index. The matrix applied to inverse LFNST can be referred to as an inverse LFNST matrix or an LFNST matrix, and such a matrix can have any name as long as it has a transposed relationship with the matrix used for forward LFNST.
[0468] In one example, the inverse LFNST matrix may be a non-square matrix in which the number of columns is less than the number of rows.
[0469] The predetermined number of transform coefficients as output data of the LFNST may be derived based on the size of the current block or the sub-partitioned transform block. For example, when the height and width of the current block or the sub-partitioned transform block are 8 or more, 48 transform coefficients may be derived, as shown in FIG. Figure 7 As shown on the left side of , and when the width and height of the sub-partition transform block are not 8 or more, that is, when the width or height of the sub-partition transform block is 4 or more and less than 8, 16 transform coefficients can be derived, as shown in Figure 7 as shown on the right side of .
[0470] like Figure 7 As shown, 48 transform coefficients may be arranged in the upper left, upper right, and lower left 4×4 regions of the upper left 8×8 region of the subpartition transform block, and 16 transform coefficients may be arranged in the upper left 4×4 region of the subpartition transform block.
[0471] 48 transform coefficients and 16 transform coefficients are arranged in the vertical direction or the horizontal direction according to the intra prediction mode of the sub-partition transform block. For example, when the intra prediction mode is based on the diagonal direction ( Figure 9 The horizontal direction of mode 34) ( Figure 9 2 to 34 in the example), the transform coefficients may be arranged in the horizontal direction, that is, in row-first order, as shown in FIG. Figure 7 (a), and when the intra prediction mode is based on the vertical direction of the diagonal direction ( Figure 9 35 to 66 in the example), the transform coefficients may be arranged in the horizontal direction, that is, in column-first order, as in Figure 7 As shown in (b).
[0472] The decoding apparatus may induce residual samples of the current block based on one inverse transform of the modified transform coefficient ( S2160 ).
[0473] In this case, a simplified inverse transform may be applied to the inverse primary transform, or a conventional separable transform may be used. Alternatively, MTS may be used as the inverse primary transform.
[0474] Subsequently, the decoding apparatus 300 may generate reconstructed samples based on the residual samples of the current block and the predicted samples of the current block.
[0475] The following figures are provided to describe specific examples of the present disclosure. Since the specific names of the devices or the names of the specific signals / messages / fields shown in the figures are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following figures.
[0476] Figure 22 is a flowchart illustrating the operation of a video encoding apparatus according to an embodiment of this document.
[0477] Figure 22 Each step disclosed in Figures 4 to 20 Therefore, some of the contents described in the above will be omitted or simplified. Figure 2 and Figures 4 to 20 Detailed description of those contents repeated in .
[0478] The encoding apparatus 200 according to an embodiment may induce a prediction sample of a current block based on an intra prediction mode applied to the current block ( S2210 ).
[0479] When applying the ISP to the current block, the encoding apparatus may perform prediction on each sub-partitioned transform block.
[0480] The encoding device can determine whether ISP encoding or ISP mode is applied to the current block (i.e., the encoding block), determine the direction in which the current block will be divided based on the determination result, and derive the size and number of divided sub-blocks.
[0481] For example, when the size (width × height) of the current block is 8 × 4, Figure 16 As shown, the current block can be vertically divided into two sub-blocks, and when the size (width×height) of the current block is 4×8, the current block can be horizontally divided into two sub-blocks. Figure 17 As shown, when the size (width × height) of the current block is greater than 4×8 or 8×4, that is, when the size of the current block is: 1) 4×N or N×4 (N≥16) or 2) M×N (M≥8, N≥8), the current block can be divided into 4 sub-blocks in the horizontal direction or the vertical direction.
[0482] The same intra prediction mode can be applied to the sub-partition transform blocks divided from the current block, and the encoding device can derive prediction samples for each sub-partition transform block. That is, the encoding device performs intra prediction sequentially from left to right or from top to bottom, for example, horizontally or vertically, according to the division form of the sub-partition transform blocks. For the leftmost or topmost sub-block, the reconstructed pixels of the already encoded coding block are referenced, as in the traditional intra prediction method. In addition, for each side of the subsequent internal sub-partition transform block, when it is not adjacent to the previous sub-partition transform block, in order to derive the reference pixels adjacent to the corresponding side, the reconstructed pixels of the already encoded adjacent coding block are referenced, as in the traditional intra prediction method.
[0483] The encoding apparatus 200 may induce residual samples of the current block based on the prediction samples ( S2220 ).
[0484] In addition, the encoding apparatus 200 may induce a transformation coefficient of the current block based on one transformation of the residual sample ( S2230 ).
[0485] One transform may be performed by a plurality of transform kernels, and in this case, the transform kernel may be selected based on the intra prediction mode.
[0486] The encoding apparatus 200 may determine whether to perform a secondary transform or an inseparable transform, particularly LFNST, on a transform coefficient of a current block, and may apply LFNST to the transform coefficient to derive a modified transform coefficient.
[0487] Unlike a single transform that separates and transforms the target coefficients in the vertical or horizontal direction, LFNST is an inseparable transform that applies the transform without separating the coefficients in a specific direction. An inseparable transform can be a low-frequency inseparable transform that applies the transform only to the low-frequency region rather than the entire target block to be transformed.
[0488] According to an example, the encoding apparatus may determine whether LFNST is applicable to the height and width of the current block based on the tree type and color format of the current block.
[0489] According to an example, in a case where the tree type of the current block is dual-tree chroma, when a height and a width corresponding to a chroma component block of the current block are greater than or equal to 4, the encoding apparatus may determine that LFNST is applicable.
[0490] In addition, according to an example, in a case where the tree type of the current block is single-tree or dual-tree luma, when a height and a width corresponding to a luma component block of the current block are greater than or equal to 4, the encoding apparatus may determine that LFNST is applicable.
[0491] In addition, when ISP is applied to the current block, that is, when the current block is divided into sub-partition transform blocks, the encoding device can determine whether LFNST is applicable to the height and width of the divided sub-partitions. In this case, when the height and width of the sub-partition block are greater than or equal to 4, the encoding device can determine that LFNST is applicable.
[0492] For example, when the tree type of the current block is dual-tree chroma, ISP may not be applied, and in this case, when the height and width corresponding to the chroma component block of the current block are greater than or equal to 4, the encoding device may determine that LFNST is applicable.
[0493] On the other hand, when the tree type of the current block is not dual-tree chroma but dual-tree or single-tree luminance, when the height and width of the sub-partition block of the luminance component block of the current block or the height and width of the current block are greater than or equal to 4, the encoding device can determine whether LFNST is applicable based on whether ISP is applied to the current block.
[0494] In addition, according to an example, when the current block is a coding unit and the width and height of the coding unit are less than or equal to the maximum luma transformable size, the encoding apparatus may determine that LFNST is applicable.
[0495] When it is determined to perform LFNST, the encoding apparatus 200 may induce a modified transform coefficient of the current block or the sub-partitioned transform block based on the LFNST set mapped to the intra prediction mode and the LFNST matrix included in the LFNST set ( S2240 ).
[0496] The encoding apparatus 200 may determine an LFNST set based on a mapping relationship according to an intra prediction mode applied to a current block, and perform LFNST, ie, inseparable transform, based on one of two LFNST matrices included in the LFNST set.
[0497] In this case, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks split from the current block. That is, because the same intra prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra prediction mode can also be applied equally to all sub-partition transform blocks. In addition, because the LFNST index is encoded in units of coding units, the same LFNST matrix can be applied to the sub-partition transform blocks split from the current block.
[0498] As described above, the transform set may be determined according to the intra prediction mode of the transform block to be transformed.The matrix applied to LFNST has a transposed relationship with the matrix used for inverse LFNST.
[0499] In one example, the LFNST matrix may be a non-square matrix in which the number of rows is less than the number of columns.
[0500] The area where the transform coefficients used as input data for LFNST are located can be derived based on the size of the sub-partition transform block. For example, when the height and width of the sub-partition transform block are 8 or more, the area is the upper left, upper right, and lower left 4×4 areas of the upper left 8×8 area of the sub-partition transform block, as shown in FIG. Figure 7 As shown on the left side of , and when the height and width of the sub-partitioned transform block are not equal to or greater than 8, the area can be the upper left 4×4 area of the current block, as shown Figure 7 as shown on the right side of .
[0501] Transform coefficients in an area may be read in a vertical direction or a horizontal direction according to an intra prediction mode of a subpartitioned transform block to construct a one-dimensional vector for a multiplication operation with an LFNST matrix.
[0502] 48 modified transform coefficients or 16 modified transform coefficients may be read in the vertical direction or the horizontal direction according to the intra prediction mode of the sub-partition transform block and arranged in one dimension. For example, when the intra prediction mode is based on the diagonal direction ( Figure 9 The horizontal direction of mode 34) ( Figure 9 2 to 34 in the example), the transform coefficients may be arranged in the horizontal direction, that is, in row-first order, as shown in FIG. Figure 7 (a), and when the intra prediction mode is based on the vertical direction of the diagonal direction ( Figure 9 35 to 66 in the example), the transform coefficients may be arranged in the horizontal direction, that is, in column-first order, as in Figure 7 As shown in (b).
[0503] In one embodiment, the encoding device may include the following steps: determining whether the encoding device is in a condition to apply LFNST, generating and encoding an LFNST index based on the determination, selecting a transform kernel matrix, and applying LFNST to the residual sample based on the selected transform kernel matrix and / or simplification factor when the encoding device is in a condition to apply LFNST. In this case, the size of the simplified transform kernel matrix may be determined based on the simplification factor.
[0504] Furthermore, according to an example, the encoding apparatus may clear a second region of the current block in which the modified transform coefficient does not exist.
[0505] like Figure 13 and Figure 14 As shown, all remaining regions in the current block where no modified transform coefficients exist can be treated as zeros. Due to this clearing to zeros, the amount of computation required to perform the entire transform process is reduced, thereby reducing the power consumption required to perform the transform. In addition, image coding efficiency can be increased by reducing the latency involved in the transform process.
[0506] According to the example, the encoding device may configure the image information so as to signal the LFNST index indicating the LFNST matrix based on the presence of a transform coefficient in an area other than the DC position of the current block, and the encoding device may configure the image information so as to signal the LFNST index regardless of whether the transform coefficient exists in an area other than the DC position based on the current block being divided into a plurality of sub-partition blocks (S2250).
[0507] The encoding device may configure the image information so that the image information shown in Table 10 is parsed by the decoding device.
[0508] That is, when the current block is not divided into a plurality of sub-partition blocks and there are transform coefficients in an area other than the DC position of the current block, the encoding device may configure the image information so that the LFNST index is parsed, and when the current block is divided into a plurality of sub-partition blocks, the encoding device may configure the image information so that the LFNST index is parsed without checking whether there are transform coefficients in an area other than the DC position.
[0509] That is, when the ISP is applied to the current block, the encoding apparatus may configure the image information so that the LFNST index is signaled even if the last non-zero coefficients of all sub-partition blocks are located at the DC position.
[0510] According to an example, when the index of the subblock containing the last significant coefficient in the current block (or sub-partition block) is 0 and the position of the last significant coefficient in the subblock is greater than 0, the encoding device may determine that there are significant coefficients in an area other than the DC position, and may configure the image information so that the LFNST index is signaled. In the present disclosure, the first position in the scanning order may be 0.
[0511] In addition, according to the example, when the index of the sub-block containing the last significant coefficient in the current block (or sub-partition block) is greater than 0 and the width and height of the current block are greater than or equal to 4, the encoding device can determine that LFNST is definitely not applied and can configure the image information so that the LFSNT index is not signaled.
[0512] In addition, according to the example, when the size of the current block (or sub-partition block) is 4×4 or 8×8, and when the position of the last significant coefficient is greater than 7 when the start of the position in the scanning order is 0, the encoding device can determine that LFNST is definitely not applied, and can configure the image information so that the LFNST index is not signaled.
[0513] That is, after the variable LfnstDcOnly and the variable LfnstZeroOutSigCoeffFlag are derived in the decoding device, the encoding device may configure the image information so that the LFNST index is parsed according to the derived variable values.
[0514] The encoding apparatus may induce a quantized transform coefficient by performing quantization based on the modified transform coefficient of the current block, and may encode information about the quantized transform coefficient and an LFNST index (if LFNST is applicable) ( S2260 ).
[0515] That is, the encoding device can generate residual information including information about the quantized transform coefficients. The residual information can include the above-mentioned transform-related information / syntax elements. The encoding device can encode the image / video information including the residual information and output the encoded image / video information in the form of a bitstream.
[0516] More specifically, the encoding apparatus 200 may generate information about the quantized transform coefficient and encode the generated information about the quantized transform coefficient.
[0517] The syntax element of the LFNST index according to this embodiment may indicate whether (inverse) LFNST is applied and any one LFNST matrix included in the LFNST set, and when the LFNST set includes two transform kernel matrices, the syntax element of the LFNST index may have three values.
[0518] According to an embodiment, when the partition tree structure of the current block is a dual tree type, an LFNST index may be encoded for each of the luma block and the chroma block.
[0519] According to an embodiment, the syntax element value of the transform index may be derived as 0, 1, and 2, 0 indicating that (inverse) LFNST is not applied to the current block, 1 indicating the first LFNST matrix in the LFNST matrix, and 2 indicating the second LFNST matrix in the LFNST matrix.
[0520] In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency.
[0521] In addition, in the present disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled through residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.
[0522] In the above embodiments, the method is explained based on a flowchart with the aid of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may be performed in an order or step different from the above order or step, or a certain step may be performed concurrently with other steps. In addition, it will be understood by those skilled in the art that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0523] The above-mentioned method according to the present disclosure may be implemented in a software form, and the encoding device and / or decoding device according to the present disclosure may be included in a device for image processing such as a television, a computer, a smart phone, a set-top box, and a display device.
[0524] When the embodiments in the present disclosure are implemented by software, the above methods can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium and / or other storage devices. In other words, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units shown in each of the accompanying drawings can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip.
[0525] In addition, the decoding device and encoding device to which the present disclosure is applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device (such as video communication), a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0526] In addition, the processing method of the present invention can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network. In addition, the embodiments of the present invention can be implemented as a computer program product through program code, and the program code can be executed on a computer according to the embodiments of the present invention. The program code can be stored on a computer-readable carrier.
[0527] Figure 23The structure of the content streaming system to which the present disclosure is applied is illustrated.
[0528] Furthermore, a content streaming system to which the present disclosure is applied may generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0529] The encoding server is used to compress content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream, and then transmit the bitstream to the streaming server. As another example, if the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. The streaming server can also temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0530] The streaming server transmits multimedia data to user devices via a web server based on user requests. The web server serves as a tool for notifying users of available services. When a user requests a desired service, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case controls commands and responses between the various devices in the content streaming system.
[0531] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to smoothly provide a streaming service, the streaming server can store the bitstream for a predetermined time.
[0532] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system may operate as a distributed server, and in this case, data received by each server may be processed in a distributed manner.
[0533] The claims disclosed herein may be combined in various ways. For example, the technical features of the method claims of the present disclosure may be combined to be implemented or performed in a device, and the technical features of the device claims may be combined to be implemented or performed in a method. Furthermore, the technical features of method claims and device claims may be combined to be implemented or performed in a device, and the technical features of method claims and device claims may be combined to be implemented or performed in a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtain residual information from the bitstream; deriving transform coefficients of the current block based on the residual information; deriving modified transform coefficients by applying an LFNST based on an LFNST matrix associated with a low frequency inseparable transform LFNST index to the transform coefficients; deriving residual samples of the current block based on an inverse primary transform of the modified transform coefficients; as well as Generate a reconstructed picture based on the residual samples, wherein, for the current block to which the intra sub-partition ISP mode is not applied, the LFNST index is parsed based on the presence of non-zero transform coefficients in a region other than a DC position of the current block, and Wherein, for the current block to which the ISP mode is applied, the LFNST index is parsed without considering whether the non-zero transform coefficient exists in the region except the DC position of the current block.
2. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: deriving transform coefficients of the current block based on a transform of residual samples of the current block; deriving modified transform coefficients from the transform coefficients based on a LFNST matrix for a low frequency non-separable transform LFNST; generating an LFNST index associated with the LFNST matrix; and encoding image information, the image information including residual information associated with the modified transform coefficient and the LFNST index, wherein, for the current block to which the intra sub-partition ISP mode is not applied, the image information is configured such that the LFNST index is parsed based on the presence of a non-zero transform coefficient in a region other than a DC position of the current block, and Here, for the current block to which the ISP mode is applied, the image information is configured such that the LFNST index is signaled without considering whether the non-zero transform coefficient exists in the region other than the DC position of the current block.
3. A method for transmitting image data, the method comprising the following steps: Obtaining a bitstream of the image, wherein the bitstream is generated based on the following: deriving transform coefficients of the current block based on a primary transform of residual samples of the current block; deriving modified transform coefficients from the transform coefficients based on an LFNST matrix for a low-frequency non-separable transform (LFNST); generating LFNST indices associated with the LFNST matrix; and encoding image information, the image information including residual information associated with the modified transform coefficients and the LFNST indices; and sending said data comprising said bitstream, wherein, for the current block to which the intra sub-partition ISP mode is not applied, the image information is configured such that the LFNST index is parsed based on the presence of a non-zero transform coefficient in a region other than a DC position of the current block, and Here, for the current block to which the ISP mode is applied, the image information is configured such that the LFNST index is signaled without considering whether the non-zero transform coefficient exists in the region other than the DC position of the current block.
Citation Information
Patent Citations
Method and apparatus for transform-based image encoding / decoding
CN109417636A
Image encoding / decoding method
CN109644276A