Image decoding apparatus, image encoding apparatus, and image data transmission apparatus
By adopting the cross-component linear model (CCLM) mode and LFNST matrix in image coding and optimizing the intra-frame prediction mode of chroma blocks, the problem of efficient compression of high-resolution images/videos is solved, and the coding efficiency and the effect of secondary transformation are improved.
Patent Information
- Application Number
- CN202511057676.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-29
- Filing Date
- 2020-10-29
- Publication Date
- 2025-09-16
AI Technical Summary
When transmitting and storing high-resolution, high-quality images/videos, existing technologies increase the amount of information, resulting in high costs and a lack of effective compression and encoding methods. This is especially true in the broadcasting of immersive media and high-frame-rate images/videos.
The Cross Component Linear Model (CCLM) mode is adopted to update the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block, and the LFNST set is derived through the LFNST matrix. The intra DC mode and intra block copy mode are used to optimize the encoding process.
Improves image/video compression efficiency, increases the efficiency of encoding LFNST index, and optimizes the encoding process through secondary transformation, thereby improving overall encoding efficiency.
Smart Images

Figure CN120658884A_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application number 202080090729.6 (International application number: PCT / KR2020 / 014915, application date: October 29, 2020, invention name: Transformation-based image coding method and device thereof). Technical Field
[0002] The present disclosure relates to image coding technology, and more particularly, to a method and apparatus for transform-based encoding of an image in an image coding system. Background Art
[0003] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K or higher ultra-high-definition (UHD) images / videos has been growing in various fields. As image / video data becomes higher resolution and higher quality, the amount of information or bit volume transmitted increases compared to traditional image data. Therefore, when using a medium such as a traditional wired / wireless broadband line to transmit image data or using an existing storage medium to store image / video data, its transmission cost and storage cost increase.
[0004] In addition, today, interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and broadcasting of images / videos having image characteristics different from real images such as game images is increasing.
[0005] Therefore, there is a need for an efficient image / video compression technology that effectively compresses and transmits or stores and reproduces information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention
[0006] Technical Purpose
[0007] A technical aspect of the present disclosure is to provide a method and apparatus for increasing image encoding efficiency.
[0008] Another technical aspect of the present disclosure is to provide a method and apparatus for increasing the efficiency of encoding LFNST indexes.
[0009] Yet another technical aspect of the present disclosure is to provide a method and apparatus for improving the efficiency of secondary transform by encoding an LFNST index.
[0010] Yet another technical aspect of the present disclosure is to provide an image encoding method and an image encoding apparatus for deriving an LFNST transform set using an intra mode of a luma block in a CCLM mode.
[0011] Technical Solution
[0012] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method may include the following steps: based on the intra-frame prediction mode of the chroma block being a cross-component linear model (CCLM) mode, updating the intra-frame prediction mode of the chroma block based on the intra-frame prediction mode of the luminance block corresponding to the chroma block; determining an LFNST set including an LFNST matrix based on the updated intra-frame prediction mode; and performing LFNST on the chroma block based on the LFNST matrix derived from the LFNST set, wherein the updated intra-frame prediction mode is derived as an intra-frame prediction mode corresponding to a specific position in the luminance block, and wherein, based on the intra-frame prediction mode corresponding to the specific position being an intra-frame block copy (IBC) mode, updating the updated intra-frame prediction mode to an intra-frame DC mode.
[0013] The specific position is set based on the color format of the chroma block.
[0014] The specific position is the center position of the luminance block.
[0015] The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)), xTbY and yTbY represent the upper left coordinate of the luminance block, nTbW and nTbH represent the width and height of the chrominance block, and SubWidthC and SubHeightC represent variables corresponding to the color format.
[0016] When the color format is 4:2:0, SubWidthC and SubHeightC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0017] When the intra prediction mode corresponding to the specific position is the MIP mode, the updated intra prediction mode is the intra planar mode.
[0018] When the intra prediction mode corresponding to the specific position is the palette mode, the updated intra prediction mode is the intra DC mode.
[0019] According to another embodiment of the present disclosure, an image encoding method performed by an encoding device is provided. The method may include the following steps: deriving prediction samples of the chroma block based on the intra-frame prediction mode of the chroma block being a cross-component linear model (CCLM); deriving residual samples of the chroma block based on the prediction samples, wherein the updated intra-frame prediction mode is derived as an intra-frame prediction mode corresponding to a specific position in the luminance block, and updating the updated intra-frame prediction mode to an intra-frame DC mode based on the intra-frame prediction mode corresponding to the specific position being an intra-frame block copy (IBC) mode.
[0020] According to still another embodiment of the present disclosure, a digital storage medium storing image data including a bit stream generated according to an image encoding method performed by an encoding device and encoded image information may be provided.
[0021] According to yet another embodiment of the present disclosure, a digital storage medium storing image data including encoded image information and a bit stream so that a decoding device performs an image decoding method may be provided.
[0022] Technical Effects
[0023] According to the present disclosure, the overall image / video compression efficiency can be increased.
[0024] According to the present disclosure, the efficiency of encoding LFNST indexes can be increased.
[0025] According to the present disclosure, the efficiency of secondary transformation can be increased by encoding the LFNST index.
[0026] According to the present disclosure, an image encoding method and an image encoding apparatus for deriving an LFNST transform set using an intra mode of a luma block in a CCLM mode may be provided.
[0027] The effects that can be obtained through the specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood or derived from the present disclosure by a person of ordinary skill in the relevant field. Therefore, the specific effects of the present disclosure are not limited to those explicitly described in the present disclosure, and may include various effects that can be understood or derived based on the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 An example of a video / image encoding system to which the present disclosure is applicable is schematically illustrated.
[0029] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which the present disclosure is applicable.
[0030] Figure 3 is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure is applicable.
[0031] Figure 4 Schematically illustrates multiple transformation schemes according to the embodiments of this document.
[0032] Figure 5 The intra directional mode for 65 prediction directions is schematically shown.
[0033] Figure 6 is a diagram for explaining RST according to an embodiment of this document.
[0034] Figure 7 is a diagram illustrating an order of arranging output data of a forward primary transform into a one-dimensional vector according to an example.
[0035] Figure 8 is a diagram illustrating an order of arranging output data of a forward quadratic transform into two-dimensional blocks according to an example.
[0036] Figure 9 is a diagram illustrating a wide-angle intra prediction mode according to an embodiment of this document.
[0037] Figure 10 is a diagram illustrating a block shape to which LFNST is applied.
[0038] Figure 11 is a diagram illustrating arrangement of output data of a forward LFNST according to an embodiment.
[0039] Figure 12 is a diagram illustrating that the number of output data of the forward LFNST according to an example is limited to a maximum of 16.
[0040] Figure 13 is a diagram illustrating clearing of zeros in a block to which 4×4 LFNST is applied according to an example.
[0041] Figure 14 is a diagram illustrating clearing of zeros in a block to which 8×8 LFNST is applied according to an example.
[0042] Figure 15 is a diagram illustrating a CCLM applicable when deriving an intra prediction mode of a chroma block according to an embodiment.
[0043] Figure 16 is a flowchart for explaining an image decoding method according to an example.
[0044] Figure 17 is a flowchart for explaining an image encoding method according to an example.
[0045] Figure 18 The structure of a content streaming system to which the present disclosure is applied is illustrated. DETAILED DESCRIPTION
[0046] Although the present disclosure may be susceptible to various modifications and includes various embodiments, its specific embodiments have been shown by way of example in the accompanying drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the technical ideas of the present disclosure. Unless the context clearly indicates otherwise, the singular form may include the plural form. Terms such as "including" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components, or combinations thereof used in the following description and should therefore not be understood as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0047] In addition, for the convenience of describing different characteristic functions, each component in the drawings described herein is illustrated independently, however, it is not intended that each component is implemented by separate hardware or software. For example, any two or more of these components can be combined to form a single component, and any single component can be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent rights of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0048] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. In addition, in the accompanying drawings, the same reference numerals are used for the same components, and repeated description of the same components will be omitted.
[0049] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).
[0050] In this document, various embodiments related to video / image encoding may be provided, and unless otherwise specified, these embodiments may be combined with each other and performed.
[0051] In this document, video can refer to a collection of images over a period of time. Generally, a picture refers to a unit that represents an image in a specific time region, and a slice / tile is a unit that constitutes a part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.
[0052] A pixel or a picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. Alternatively, a sample may refer to a pixel value in a spatial domain, or when the pixel value is transformed into a frequency domain, it may refer to a transform coefficient in the frequency domain.
[0053] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region and information related to the region. A unit may include a luminance block and two chrominance (e.g., CB, CR) blocks. Depending on the situation, terms such as unit and block, region, etc. may be used interchangeably. In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0054] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". In addition, "A / B / C" may mean "at least one of A, B, and / or C".
[0055] Additionally, in this document, the term "or" should be interpreted as meaning "and / or." For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as meaning "additionally or alternatively."
[0056] In the present disclosure, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in the present disclosure, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as “at least one of A and B”.
[0057] Furthermore, in the present disclosure, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” Furthermore, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”
[0058] In addition, the brackets used in this disclosure may indicate "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, it may mean that "intra-frame prediction" is proposed as an example of "prediction." In other words, "prediction" in this disclosure is not limited to "intra-frame prediction," and "intra-frame prediction" is proposed as an example of "prediction." In addition, when "prediction (i.e., intra-frame prediction)" is indicated, it may also mean that "intra-frame prediction" is proposed as an example of "prediction."
[0059] Technical features described separately in one drawing in the present disclosure may be implemented separately or may be implemented simultaneously.
[0060] Figure 1 An example of a video / image encoding system to which the present disclosure is applicable is schematically illustrated.
[0061] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may transmit the encoded video / image information or data to the receive device in the form of a file or stream via a digital storage medium or a network.
[0062] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0063] The video source can obtain the video / image by capturing, synthesizing, or generating a video / image. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.
[0064] An encoding device can encode input video / images. It can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0065] The transmitter can transmit the encoded video / image information or data, output as a bitstream, to a receiver in a receiving device via a digital storage medium or network in the form of a file or stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received / extracted bitstream to a decoding device.
[0066] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0067] The renderer can render the decoded video / image, and the rendered video / image can be displayed on a display.
[0068] Figure 2 Schematically illustrates a configuration of a video / image encoding device to which the present disclosure is applicable. Hereinafter, the so-called video encoding device may include an image encoding device.
[0069] Reference Figure 2 , the encoding device 200 may include an image divider 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. Depending on the embodiment, the image divider 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be composed of one or more hardware components (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0070] The image divider 210 may divide the input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a maximum coding unit (LCU), the coding units may be recursively divided according to a quadtree, binary tree, ternary tree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, a coding unit may be divided into multiple coding units of a deeper depth. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that has not been further divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively divided into coding units of a deeper depth as needed, thereby allowing the optimally sized coding unit to be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be separated or divided from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0071] Depending on the situation, terms such as unit and block, region, etc. may be used instead of each other. In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample may be used as a term corresponding to a pixel or a picture element (pel) of a picture (or image).
[0072] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from the predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. The predictor 220 can perform prediction on the processing target block (hereinafter referred to as "current block") and can generate a prediction block including prediction samples of the current block. The predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. Information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0073] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference sample can be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0074] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate was used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, the residual signal cannot be sent. In the case of motion information prediction (motion vector prediction, MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0075] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction at the same time. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode to perform prediction on the block. The IBC prediction mode or the palette mode can be used for content image / video encoding such as games such as screen content coding (SCC). Although IBC basically performs prediction in the current block, its execution method is similar to inter-frame prediction in that it derives a reference block in the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0076] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate a reconstructed signal or a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT means a transform obtained from a curve graph when the relationship information between pixels is represented by a curve graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size rather than square blocks.
[0077] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream. The information about the quantized transform coefficients can be called residual information. The quantizer 233 can rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode information required for video / image reconstruction in addition to the quantized transform coefficients (e.g., syntax element values, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream on a unit basis of the network abstraction layer (NAL). The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. In addition, the video / image information may also include general constraint information. In the present disclosure, information and / or syntax elements sent / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded by the above-mentioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 or a memory (not shown) that stores it may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0078] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual sample) can be reconstructed. The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222, so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the target picture, and as described later, can be used for inter-frame prediction of the next picture performed by filtering.
[0079] Furthermore, in the picture encoding and / or reconstruction process, luma mapping with chroma scaling (LMCS) may be applied.
[0080] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, especially in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive ring filter, bilateral filter, etc. As discussed later in the description of each filtering method, the filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0081] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. Accordingly, the encoding device can avoid prediction mismatch in the encoding device 100 and the decoding device when applying inter-frame prediction, and can also improve encoding efficiency.
[0082] The memory 270DPB can store the modified reconstructed picture so that it can be used as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the blocks in the current picture from which the motion information has been derived (or encoded) and / or the motion information of the blocks in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 221 to be used as the motion information of the neighboring blocks or the motion information of the temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.
[0083] Figure 3is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure is applicable.
[0084] Reference Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be composed of one or more hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0085] When a bit stream including video / image information is input, the decoding device 300 can Figure 2 The image is reconstructed correspondingly to the processing of the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the information related to the block segmentation obtained from the bit stream. The decoding device 300 can perform decoding by using the processing unit applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided into a quadtree structure, a binary tree structure and / or a ternary tree structure using a coding tree unit or a maximum coding unit. One or more transformation units can be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproducer.
[0086] The decoding device 300 may receive the data from the Figure 2The received signal is output by the encoding device, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. In the present disclosure, the signaled / received information and / or syntax elements described later can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, use the decoded target syntax element information and the decoded information of the neighboring and decoded target blocks or the information of the symbol / bin decoded in the previous step to determine the context model, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bin to generate the symbol corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded by the context model for the next symbol / bin after determining the context model. The information about prediction among the information decoded in the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (i.e., quantized transform coefficient) and associated parameter information for which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. In addition, a receiver (not shown) that receives a signal output from the encoding device may also constitute the decoding device 300 as an internal / external element, and the receiver may be a component of the entropy decoder 310. In addition, the decoding device according to the present disclosure may be referred to as a video / image / picture encoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0087] The dequantizer 322 can output the transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block. In this case, the rearrangement can be performed based on the order of coefficient scanning performed in the encoding device. The dequantizer 321 can dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0088] The dequantizer 322 obtains a residual signal (residual block, residual sample array) by performing inverse transformation on the transformation coefficients.
[0089] The predictor may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and specifically may determine the intra / inter prediction mode.
[0090] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction at the same time. This can be called combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor can perform intra-block copying (IBC) for the prediction of the block. Intra-block copying can be used for content image / video coding such as games such as screen content coding (SCC). Although IBC basically performs prediction in the current block, its execution method is similar to inter-frame prediction in that it derives a reference block in the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0091] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0092] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the inter-frame prediction mode for the current block.
[0093] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor 330. When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.
[0094] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next processing target block in the current block, and as described later, may be output through filtering or used for inter prediction of the next picture.
[0095] Furthermore, in the picture decoding process, luma mapping with chroma scaling (LMCS) may be applied.
[0096] The filter 350 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can send the modified reconstructed picture to the memory 360, in particular, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive ring filter, bilateral filter, etc.
[0097] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block in the current picture from which the motion information has been derived (or decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the neighboring block or the motion information of the temporally neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.
[0098] In this specification, the examples described in the predictor 330, dequantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 may be similarly or correspondingly applied to the predictor 220, dequantizer 234, inverse transformer 235, and filter 260 of the encoding device 200, respectively.
[0099] As described above, prediction is performed in order to improve compression efficiency when performing video encoding. Accordingly, a prediction block including prediction samples for a current block as an encoding target block can be generated. Here, the prediction block includes prediction samples in a spatial domain (or a pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve image coding efficiency by signaling to the decoding device information (residual information) about the residual between the original block and the prediction block, rather than the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed picture including the reconstructed block.
[0100] Residual information can be generated through a transform process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients, so that it can signal the associated residual information to the decoding device (through a bitstream). Here, the residual information may include value information, position information, transform technology, transform kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / dequantization process based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive a residual block by dequantizing / inverse transforming the quantized transform coefficients to serve as a reference for inter-frame prediction of the next picture, and can generate a reconstructed picture based on this.
[0101] Figure 4 The multi-conversion technology according to the embodiment of the present disclosure is schematically illustrated.
[0102] Reference Figure 4 , the converter can correspond to the aforementioned Figure 2 The converter in the encoding device, and the inverse converter may correspond to the aforementioned Figure 2 The inverse transformer in the encoding device, or Figure 3 An inverse transformer in a decoding device.
[0103] The transformer may derive (primary) transform coefficients by performing a primary transform based on the residual samples (residual sample array) in the residual block (S410). This primary transform may be referred to as a core transform. In this document, the primary transform may be based on a multi-transform selection (MTS), and when multiple transforms are used as the primary transform, it may be referred to as a multi-core transform.
[0104] Multi-core transform may refer to a method for performing transforms using discrete cosine transform (DCT) type 2 and discrete sine transform (DST) type 7, DCT type 8, and / or DST type 1 in addition. In other words, multi-core transform may refer to a transform method that transforms a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this document, the primary transform coefficients may be referred to as temporary transform coefficients from the perspective of the transformer.
[0105] In other words, when a conventional transform method is applied, transform coefficients can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, when a multi-core transform is applied, transform coefficients (or primary transform coefficients) can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this document, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transform types, transform kernels, or transform cores. These DCT / DST transform types may be defined based on basis functions.
[0106] When performing multi-core transformation, a vertical transform kernel and a horizontal transform kernel for the target block can be selected from the transform kernels, a vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can indicate the transform of the horizontal component of the target block, and the vertical transform can indicate the transform of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block) including the residual block.
[0107] In addition, according to an example, if a transform is performed once by applying MTS, the mapping relationship of the transform kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transform or the horizontal transform. For example, when the horizontal transform kernel is represented as trTypeHor and the vertical transform kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT8.
[0108] In this case, the MTS index information may be encoded and signaled to the decoding device to indicate any one of a plurality of transform kernel sets. For example, an MTS index of 0 may indicate that both trTypeHor and trTypeVer values are 0, an MTS index of 1 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 2 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 3 may indicate that both trTypeHor and trTypeVer values are 1 and 2, and an MTS index of 4 may indicate that both trTypeHor and trTypeVer values are 2.
[0109] In one example, the transformation kernel set according to the MTS index information is shown in the following table.
[0110] [Table 1]
[0111] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0112] The transformer may perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients (S420). A primary transform is a transform from the spatial domain to the frequency domain, while a secondary transform refers to a transform into a more compact representation using the correlation existing between the (primary) transform coefficients. The secondary transform may include an inseparable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a pattern-dependent non-separable secondary transform (MDNSST). NSST may represent a transform that performs a secondary transform on the (primary) transform coefficients derived by the primary transform based on a non-separable transform matrix to generate modified transform coefficients (or secondary transform coefficients) for the residual signal. Here, based on the non-separable transform matrix, the transform may be applied once to the (primary) transform coefficients without separating the vertical transform and the horizontal transform (or applying the horizontal / vertical transform independently). In other words, NSST is not applied separately to (primary) transform coefficients in the vertical and horizontal directions, and can represent, for example, a transform method in which a two-dimensional signal (transform coefficient) is rearranged into a one-dimensional signal through a specific predetermined direction (e.g., a row-first direction or a column-first direction) and then the modified transform coefficients (or secondary transform coefficients) are generated based on an inseparable transform matrix. For example, the row-first order is to arrange the M×N blocks in the order of the first row, the second row, ..., and the Nth row, while the column-first order is to arrange the M×N blocks in the order of the first column, the second column, ..., and the Mth column. NSST can be applied to the upper left area of a block (hereinafter, referred to as a transform coefficient block) configured with (primary) transform coefficients. For example, when the width W and the height H of the transform coefficient block are both 8 or larger, 8×8 NSST can be applied to the upper left 8×8 area of the transform coefficient block. In addition, when both the width (W) and the height (H) of the transform coefficient block are 4 or more, when the width (W) or the height (H) of the transform coefficient block is less than 8, the 4×4 NSST may be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block. However, the embodiment is not limited thereto, and for example, even if only the condition that the width W or the height H of the transform coefficient block is 4 or more is satisfied, the 4×4 NSST may be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block.
[0113] Specifically, for example, if a 4×4 input block is used, the non-separable secondary transform may be performed as follows.
[0114] A 4×4 input block X can be represented as follows.
[0115] [Formula 1]
[0116]
[0117] If X is represented as a vector, then the vector It can be expressed as follows.
[0118] [Formula 2]
[0119]
[0120] In Equation 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to row-major order.
[0121] In this case, the non-separable quadratic transform can be calculated as follows.
[0122] [Formula 3]
[0123]
[0124] In this formula, denotes a transform coefficient vector, and T denotes a 16x16 (non-separable) transform matrix.
[0125] By using the above formula 3, the 16×1 transform coefficient vector can be derived And the vector can be scanned in order (horizontally, vertically, diagonally, etc.) Reorganized into 4×4 blocks. However, the above calculation is an example, and Hypercube-Givens Transform (HyGT) or the like may also be used for the calculation of the inseparable secondary transform in order to reduce the computational complexity of the inseparable secondary transform.
[0126] Furthermore, in the inseparable secondary transform, the transform kernel (or transform core, transform type) may be selected to be mode-dependent. In this case, the mode may include an intra prediction mode and / or an inter prediction mode.
[0127] As described above, an inseparable secondary transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 area included in the transform coefficient block when both W and H are equal to or greater than 8, and the 8×8 area can be the upper left 8×8 area in the transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 area included in the transform coefficient block when both W and H are equal to or greater than 4, and the 4×4 area can be the upper left 4×4 area in the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0128] Here, in order to select mode-dependent transform kernels, two inseparable secondary transform kernels may be configured for each transform set used for inseparable secondary transforms for both the 8×8 transform and the 4×4 transform, and four transform sets may exist. That is, four transform sets may be configured for the 8×8 transform, and four transform sets may be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform may include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform may include two 4×4 transform kernels.
[0129] However, as the size of the transform (ie, the size of the region to which the transform is applied) may be other than 8×8 or 4×4, for example, the number of sets may be n, and the number of transform kernels in each set may be k.
[0130] The transform set may be referred to as an NSST set or a LFNST set. A specific set among the transform sets may be selected, for example, based on the intra prediction mode of the current block (CU or subblock). A low-frequency non-separable transform (LFNST) may be an example of a reduced non-separable transform, which will be described later and represents a non-separable transform for low-frequency components.
[0131] For reference, for example, the intra prediction mode may include two non-directional (or non-angle) intra prediction modes and 65 directional (or angle) intra prediction modes. The non-directional intra prediction mode may include a plane intra prediction mode No. 0 and a DC intra prediction mode No. 1, and the directional intra prediction mode may include 65 intra prediction modes No. 2 to No. 66. However, this is an example, and this document may be applied even if the number of intra prediction modes is different. In addition, in some cases, intra prediction mode No. 67 may also be used, and intra prediction mode No. 67 may represent a linear model (LM) mode.
[0132] Figure 5 The intra directional mode for 65 prediction directions is schematically shown.
[0133] Reference Figure 5 , based on the intra prediction mode 34 having the upper left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode having a horizontal directionality and an intra prediction mode having a vertical directionality. Figure 5In FIG, H and V denote horizontal and vertical directivities, respectively, and numbers -32 to 32 indicate displacements of 1 / 32 units on the sample grid position. These numbers may represent offsets for mode index values. Intra-prediction modes 2 to 33 have horizontal directivities, and intra-prediction modes 34 to 66 have vertical directivities. Strictly speaking, intra-prediction mode 34 may be considered neither horizontal nor vertical, but may be classified as belonging to horizontal directivity when determining the transform set for the secondary transform. This is because the input data is transposed for a vertically oriented mode that is symmetrical based on intra-prediction mode 34, and the input data alignment method for the horizontal mode is used for intra-prediction mode 34. Transposing the input data means switching the rows and columns of two-dimensional M×N block data to N×M data. Intra-prediction mode 18 and intra-prediction mode 50 may represent horizontal intra-prediction mode and vertical intra-prediction mode, respectively, and intra-prediction mode 2 may be referred to as an upper right diagonal intra-prediction mode because intra-prediction mode 2 has a left reference pixel and performs prediction in the upper right direction. Similarly, intra-prediction mode 34 may be referred to as a bottom-right diagonal intra-prediction mode, and intra-prediction mode 66 may be referred to as a bottom-left diagonal intra-prediction mode.
[0134] According to an example, four transform sets according to intra prediction modes may be mapped, for example, as shown in the following table.
[0135] [Table 2]
[0136] lfnstPredModeIntra lfnstIrSetIdx lfnstPredModeIntra<0 1 0<=lfnstPredModeIntra<=1 0 2<=lfnstPredModeIntra<=12 1 13<=lfnstPredModeIntra<=23 2 24<=lfnstPredModeIntra<=44 3 45<=lfnstPredModeIntra<=55 2 56<=lfnstPredModeIntra<=80 1 81<=lfnstPredModeIntra<=83 0
[0137] As shown in Table 2, any one of four transform sets, ie, lfnstTrSetIdx, may be mapped to any one of four indexes (ie, 0 to 3) according to the intra prediction mode.
[0138] When it is determined that a specific set is used for an inseparable transform, one of the k transform cores in the specific set can be selected by an inseparable secondary transform index. The encoding device can derive an inseparable secondary transform index indicating a specific transform core based on a rate-distortion (RD) check, and can signal the inseparable secondary transform index to the decoding device. The decoding device can select one of the k transform cores in the specific set based on the inseparable secondary transform index. For example, an lfnst index value 0 can refer to a first inseparable secondary transform core, an lfnst index value 1 can refer to a second inseparable secondary transform core, and an lfnst index value 2 can refer to a third inseparable secondary transform core. Alternatively, an lfnst index value 0 can indicate that the first inseparable secondary transform is not applied to the target block, and lfnst index values 1 to 3 can indicate three transform cores.
[0139] The transformer can perform a non-separable secondary transform based on the selected transform kernel and can obtain modified (secondary) transform coefficients. As described above, the modified transform coefficients can be derived as transform coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device and transmitted to the dequantizer / inverse transformer in the encoding device.
[0140] In addition, as described above, if the secondary transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device and transmitted to the dequantizer / inverse transformer in the encoding device.
[0141] The inverse transformer may perform a series of processes in the reverse order of the order already performed in the above-mentioned transformer. The inverse transformer may receive the (dequantized) transform coefficients and derive the (primary) transform coefficients by performing a secondary (inverse) transform (S450), and may obtain the residual block (residual sample) by performing a primary (inverse) transform on the (primary) transform coefficients (S460). In this regard, from the perspective of the inverse transformer, the primary transform coefficients may be referred to as modified transform coefficients. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.
[0142] The decoding device may further include a secondary inverse transform application determiner (or an element for determining whether to apply a secondary inverse transform) and a secondary inverse transform determiner (or an element for determining a secondary inverse transform). The secondary inverse transform application determiner may determine whether to apply a secondary inverse transform. For example, the secondary inverse transform may be NSST, RST, or LFNST, and the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a secondary transform flag obtained by parsing the bitstream. In another example, the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a transform coefficient of a residual block.
[0143] The secondary inverse transform determiner may determine the secondary inverse transform. In this case, the secondary inverse transform determiner may determine the secondary inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra-frame prediction mode. In an embodiment, the secondary transform determination method may be determined depending on the primary transform determination method. Various combinations of primary and secondary transforms may be determined according to the intra-frame prediction mode. In addition, in an example, the secondary inverse transform determiner may determine the area to which the secondary inverse transform is applied based on the size of the current block.
[0144] In addition, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, a (separable) inverse transform can be performed once, and a residual block (residual sample) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0145] Furthermore, in the present disclosure, reduced quadratic transform (RST) in which the size of a transformation matrix (kernel) is reduced may be applied in the concept of NSST in order to reduce the amount of calculation and storage required for an inseparable quadratic transform.
[0146] In addition, the transformation kernel, transformation matrix, and coefficients constituting the transformation kernel matrix described in the present disclosure, that is, kernel coefficients or matrix coefficients, can be represented in 8 bits. This can be a condition for implementation in decoding devices and encoding devices, and compared with existing 9 bits or 10 bits, the amount of storage required to store the transformation kernel can be reduced, and performance degradation can be reasonably adapted. In addition, representing the kernel matrix in 8 bits can allow the use of small multipliers and can be more suitable for single instruction multiple data (SIMD) instructions for optimal software implementation.
[0147] In this specification, the term "RST" may refer to a transform performed on residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. When performing a downscale transform, the amount of computation required for the transform can be reduced due to the reduction in the size of the transform matrix. In other words, RST can be used to address the computational complexity issues that arise when transforming large blocks or non-separable transforms.
[0148] RST may be referred to by various terms such as reduced transform, reduced secondary transform, downscaling transform, simplified transform, and simple transform, and the names that RST may be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in a low-frequency region including non-zero coefficients in a transform block, it may be referred to as a low-frequency non-separable transform (LFNST). The transform index may be referred to as an LFNST index.
[0149] In addition, when performing a secondary inverse transform based on an RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include: an inverse downscaled secondary transformer that derives modified transform coefficients based on an inverse RST of the transform coefficients; and an inverse primary transformer that derives residual samples of the target block based on an inverse primary transform of the modified transform coefficients. An inverse primary transform refers to an inverse transform of a primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.
[0150] Figure 6is a diagram illustrating an RST according to an embodiment of the present disclosure.
[0151] In this disclosure, a “target block” may refer to a current block to be encoded, a residual block, or a transform block.
[0152] In the RST according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, so that a reduced transformation matrix can be determined, where R is less than N. N can refer to the square of the length of the side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the reduction factor can refer to an R / N value. The reduction factor can be referred to as a reduction factor, a shrinkage factor, a simplification factor, a simple factor, or various other terms. In addition, R can be referred to as a reduction coefficient, but depending on the situation, the reduction factor can refer to R. In addition, depending on the situation, the reduction factor can refer to an N / R value.
[0153] In the example, the reduction factor or reduction coefficient may be signaled through the bitstream, but the example is not limited thereto. For example, a predetermined value for the reduction factor or reduction coefficient may be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or reduction coefficient may not be signaled separately.
[0154] The size of the reduced transform matrix according to an example may be R×N, which is smaller than N×N (the size of a conventional transform matrix), and may be defined as in Equation 4 below.
[0155] [Formula 4]
[0156]
[0157] Figure 6 The matrix T in the reduced transform block shown in (a) may refer to the matrix T of Equation 4. R×N .like Figure 6 As shown in (a), when the reduced transformation matrix T R×N When multiplied by the residual samples of the target block, the transform coefficients of the current block can be derived.
[0158] In an example, if the size of the block to which the transform is applied is 8×8 and R=16 (ie, R / N=16 / 64=1 / 4), then according to Figure 6 The RST of (a) can be expressed as a matrix operation shown in the following Equation 5. In this case, the storage and multiplication calculations can be reduced to about 1 / 4 by a reduction factor.
[0159] In the present disclosure, a matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix provided on the left side of the column vector.
[0160] [Formula 5]
[0161]
[0162] In formula 5, r1 to r 64 ∫ may represent the residual sample of the target block, and specifically may be a transform coefficient generated by applying one transform. As a result of the calculation of Equation 5, the transform coefficient c of the target block may be derived. i , and derive c i The process can be shown as Equation 6.
[0163] [Formula 6]
[0164]
[0165] As a result of the calculation of Equation 6, the transform coefficients c1 to c R That is, when R=16, the transform coefficients c1 to c 16 If a normal transform is applied instead of RST and a transform matrix of 64×64 (N×N) size is multiplied by a residual sample of 64×1 (N×1) size, only 16 (R) transform coefficients are derived for the target block because RST is applied, even though 64 (N) transform coefficients are derived for the target block. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data transmitted by the encoding device 200 to the decoding device 300 is reduced, and thus the transmission efficiency between the encoding device 200 and the decoding device 300 can be improved.
[0166] When considering the size of the transformation matrix, the size of the conventional transformation matrix is 64×64 (N×N), but the size of the reduced transformation matrix is reduced to 16×64 (R×N). Therefore, compared with the case of performing the conventional transformation, the memory usage when performing RST can be reduced by the R / N ratio. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional transformation matrix, the number of multiplication calculations (R×N) can be reduced by the R / N ratio when using the reduced transformation matrix.
[0167] In an example, the transformer 232 of the encoding device 200 may derive transform coefficients of the target block by performing a primary transform and a secondary transform based on an RST on the residual samples of the target block. These transform coefficients may be transmitted to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 may derive modified transform coefficients based on an inverse reduced secondary transform (RST) for the transform coefficients, and may derive residual samples of the target block based on an inverse primary transform for the modified transform coefficients.
[0168] According to the example, the inverse RST matrix T N×RThe size of is N×R which is smaller than the size of the conventional inverse transform matrix N×N and is similar to the reduced transform matrix T shown in Equation 4. R×N Has a transposition relationship.
[0169] Figure 6 The matrix T in the reduced inverse transform block shown in (b) t Can refer to the inverse RST matrix T N×R T (The superscript T refers to transposition). Figure 6 As shown in (b), when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. R×N T It can be expressed as (T R×N ) T N×R .
[0170] More specifically, when the inverse RST is used as the secondary inverse transform, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the residual samples of the target block can be derived.
[0171] In an example, if the size of the block to which the inverse transform is applied is 8×8 and R=16 (ie, R / N=16 / 64=1 / 4), then according to Figure 6 The RST of (b) can be expressed as the matrix operation shown in the following Equation 7.
[0172] [Formula 7]
[0173]
[0174] In formula 7, c1 to c 16 As a result of the calculation of Equation 7, r representing the modified transform coefficient of the target block or the residual sample of the target block can be derived. j , and derive r j The process can be shown as formula 8.
[0175] [Formula 8]
[0176]
[0177] As a result of the calculation of Equation 8, r1 to r2 representing the modified transform coefficients of the target block or the residual samples of the target block can be derived. N . From the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduced transform matrix is reduced to 64×16 (R×N), so the storage usage in the case of performing inverse RST can be reduced by the R / N ratio compared to the case of performing the conventional inverse transform. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional inverse transform matrix, the use of the inverse reduced transform matrix can reduce the number of multiplication calculations (N×R) by the R / N ratio.
[0178] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied according to the transform set in Table 2. Since a transform set includes two or three transforms (kernels) according to the intra prediction mode, it can be configured to select one of up to four transforms, including the case where the secondary transform is not applied. In the transform where the secondary transform is not applied, the application of the identity matrix can be considered. Assuming that indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case where the identity matrix is applied, that is, the case where the secondary transform is not applied), the transform index or LFNST index as a syntax element can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the upper left 8×8 block, the 8×8 NSST in the RST configuration can be specified by the transform index, or the 8×8 LFNST can be specified when LFNST is applied. 8×8lfnst and 8×8RST refer to transforms that can be applied to an 8×8 region included in a transform coefficient block when both W and H of a target block to be transformed are equal to or greater than 8, and the 8×8 region may be the upper left 8×8 region in the transform coefficient block. Similarly, 4×4lfnst and 4×4RST refer to transforms that can be applied to a 4×4 region included in a transform coefficient block when both W and H of a target block are equal to or greater than 4, and the 4×4 region may be the upper left 4×4 region in the transform coefficient block.
[0179] According to an embodiment of the present disclosure, for the transformation in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 area. Here, “maximum” means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is, when RST is performed by applying an m×48 transform kernel matrix (m≤16) to an 8×8 area, 48 pieces of data are input and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming an 8×8 area can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on the 48 pieces of data constituting the area other than the lower right 4×4 area among the 8×8 areas. Here, when the matrix operation is performed by applying the maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper left 4×4 area according to the scanning order, and the upper right 4×4 area and the lower left 4×4 area can be filled with zeros.
[0180] For the inverse transform in the decoding process, a transposed matrix of the aforementioned transform kernel matrix may be used. That is, when inverse RST or LFNST is performed in the inverse transform process performed by the decoding device, input coefficient data to which inverse RST is applied is arranged in a one-dimensional vector according to a predetermined arrangement order, and modified coefficient vectors obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector may be arranged in a two-dimensional block according to a predetermined arrangement order.
[0181] In summary, during the transform process, when RST or LFNST is applied to an 8×8 region, the 48 transform coefficients in the upper left, upper right, and lower left regions of the 8×8 region, excluding the lower right region, are subjected to a matrix operation with the 16×48 transform kernel matrix. For the matrix operation, the 48 transform coefficients are input as a one-dimensional array. When the matrix operation is performed, 16 modified transform coefficients are derived and arranged in the upper left region of the 8×8 region.
[0182] On the contrary, in the inverse transform process, when the inverse RST or LFNST is applied to an 8×8 area, 16 transform coefficients corresponding to the upper left area of the 8×8 area among the transform coefficients in the 8×8 area can be input in a one-dimensional array according to the scanning order, and can undergo a matrix operation with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix)*(16×1 transform coefficient vector)=(48×1 modified transform coefficient vector). Here, the n×1 vector can be interpreted as having the same meaning as the n×1 matrix, and can therefore be expressed as an n×1 column vector. In addition, * represents matrix multiplication. When the matrix operation is performed, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left area, the upper right area, and the lower left area of the 8×8 area except the lower right area.
[0183] When the secondary inverse transform is based on the RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include an inverse downscaling secondary transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients and an inverse primary transformer for deriving residual samples of the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.
[0184] The non-separable transform (LFNST) described above will be described in detail as follows: LFNST may include a forward transform performed by an encoding device and an inverse transform performed by a decoding device.
[0185] The encoding device receives as input a result (or a portion of the result) derived after applying a primary (core) transform, and applies a forward secondary transform (secondary transform).
[0186] [Formula 9]
[0187] V=G T x
[0188] In Equation 9, x and y are the input and output of the quadratic transform, respectively, G is a matrix representing the quadratic transform, and the transform basis vectors consist of column vectors. In the case of inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the transpose of the matrix G becomes G T dimension.
[0189] For inverse LFNST, the dimensions of the matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of 8 transformation basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.
[0190] On the other hand, for forward LFNST, the matrix G T The dimensions are [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transformation basis vectors from the upper parts of the [16×48] matrix and the [16×16] matrix, respectively.
[0191] Therefore, in the case of forward LFNST, a [48×1] vector or a [16×1] vector can be used as input x, and a [16×1] vector or a [8×1] vector can be used as output y. In video encoding and decoding, the output of the forward primary transform is two-dimensional (2D) data, so in order to construct a [48×1] vector or a [16×1] vector as input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data as the output of the forward transform.
[0192] Figure 7 is a diagram illustrating an order of arranging output data of a forward primary transform into a one-dimensional vector according to an example. Figure 7 The left figures of (a) and (b) show the order for constructing a [48×1] vector, and Figure 7 The right figures of (a) and (b) show the order for constructing the [16×1] vector. In the case of LFNST, the 2D data can be constructed by Figure 7 The same order as in (a) and (b) is sequentially arranged to obtain a one-dimensional vector x.
[0193] The arrangement direction of the output data of the forward primary transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in the horizontal direction relative to the diagonal direction, the output data of the forward primary transform can be arranged in the following manner: Figure 7 The output data of the forward primary transform are arranged in the order of (a) and when the intra prediction mode of the current block is in a vertical direction relative to the diagonal direction, the output data of the forward primary transform can be arranged in the order of (a) Figure 7 The output data of the forward primary transformation are arranged in the order of (b).
[0194] According to the example, different Figure 7 The arrangement order of (a) and (b) is the arrangement order of (a) and (b), and in order to derive and apply Figure 7If the arrangement order of (a) and (b) is the same as the result (y vector), the column vectors of the matrix G can be rearranged according to the arrangement order. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0195] Since the output y derived by Formula 9 is a one-dimensional vector, when two-dimensional data is required as input data in a process using the result of the forward quadratic transform as input (for example, in the process of performing quantization or residual coding), the output y vector of Formula 9 needs to be properly arranged as 2D data again.
[0196] Figure 8 is a diagram illustrating an order of arranging output data of a forward quadratic transform into two-dimensional blocks according to an example.
[0197] In the case of LFNST, the output values can be arranged in 2D blocks according to a predetermined scanning order. Figure 8 (a) shows that when the output y is a [16×1] vector, the output values are arranged at 16 positions of the 2D block according to the diagonal scanning order. Figure 8 (b) shows that when the output y is an [8×1] vector, the output values are arranged at 8 positions of the 2D block according to the diagonal scanning order, and the remaining 8 positions are filled with zeros. Figure 8 The X in (b) indicates that it is filled with zeros.
[0198] According to another example, since the order of processing the output vector y when performing quantization or residual encoding can be preset, the output vector y may not be arranged in a sequence such as Figure 8 However, in the case of residual coding, data encoding can be performed in 2D block (e.g., 4×4) units (e.g., CG (coefficient group)), and in this case, according to Figure 8 The data is arranged in a specific order in the diagonal scan order of .
[0199] In addition, the decoding apparatus may configure a one-dimensional input vector y by arranging two-dimensional data output through a dequantization process according to a preset scanning order for inverse transform. The input vector y may be output as an output vector x through the following equation.
[0200] [Equation 10]
[0201] x=Gy
[0202] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or a [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.
[0203] The output vector x is based on Figure 7 The order shown in is arranged in a two-dimensional block and is arranged as two-dimensional data, and the two-dimensional data becomes input data (or a part of input data) of an inverse primary transform.
[0204] Therefore, the inverse quadratic transform is overall the reverse of the forward quadratic transform process, and in the case of the inverse transform, unlike in the forward direction, the inverse quadratic transform is applied first and then the inverse primary transform.
[0205] In the inverse LFNST, one of 8 [48×16] matrices and 8 [16×16] matrices can be selected as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.
[0206] In addition, eight matrices can be derived from the four transform sets shown in Table 2 above, and each transform set can be composed of two matrices. Which transform set to use among the four transform sets is determined according to the intra prediction mode, and more specifically, the transform set is determined based on the value of the intra prediction mode extended by considering wide-angle intra prediction (WAIP). Which matrix is selected from the two matrices constituting the selected transform set is derived by index signaling. More specifically, 0, 1, and 2 can be used as the index values sent, 0 can indicate that LFNST is not applied, and 1 and 2 can indicate either of the two transform matrices constituting the transform set selected based on the intra prediction mode value.
[0207] Figure 9 is a diagram illustrating a wide-angle intra prediction mode according to an embodiment of this document.
[0208] The general intra prediction mode value may have values from 0 to 66 and from 81 to 83, and the intra prediction mode value extended due to WAIP may have values from -14 to 83 as shown. The values from 81 to 83 indicate the CCLM (Cross Component Linear Model) mode, and the values from -14 to -1 and the values from 67 to 80 indicate the intra prediction mode extended due to WAIP application.
[0209] When the width of the current prediction block is greater than its height, the upper reference pixel is generally closer to the position inside the block to be predicted. Therefore, prediction in the lower left direction can be more accurate than in the upper right direction. Conversely, when the height of the block is greater than its width, the left reference pixel is generally closer to the position inside the block to be predicted. Therefore, prediction in the upper right direction can be more accurate than in the lower left direction. Therefore, applying remapping (i.e., mode index modification) to the index of the Wide Intra prediction mode can be advantageous.
[0210] When wide-angle intra prediction is applied, information about existing intra predictions may be signaled, and after the information is parsed, the information may be remapped to the index of the wide-angle intra prediction mode. Therefore, the total number of intra prediction modes for a specific block (e.g., a non-square block of a specific size) may not be changed, that is, the total number of intra prediction modes is 67, and the intra prediction mode encoding for the specific block may not be changed.
[0211] Table 3 below shows a process of deriving a modified intra mode by remapping the intra prediction mode to the wide-angle intra prediction mode.
[0212] [Table 3]
[0213]
[0214] In Table 3, the extended intra prediction mode value is finally stored in the predModeIntra variable, and ISP_NO_SPLIT indicates that the CU block is not divided into sub-partitions by the intra sub-partitioning (ISP) technology currently adopted in the VVC standard, and the cIdx variable values 0, 1, and 2 indicate the cases of the luma component, Cb component, and Cr component, respectively. The log2 function shown in Table 3 returns a log value with a base of 2, and the Abs function returns an absolute value.
[0215] The variable predModeIntra indicating the intra prediction mode and the height and width of the transform block are used as input values for the wide-angle intra prediction mode mapping process, and the output value is the modified intra prediction mode predModeIntra. The height and width of the transform block or coding block can be the height and width of the current block used for remapping of the intra prediction mode. At this time, the variable whRatio reflecting the ratio of width to width can be set to Abs(Log2(nW / nH)).
[0216] For non-square blocks, the intra prediction mode can be divided into two cases and modified.
[0217] First, if all of conditions (1) to (3) are satisfied, (1) the width of the current block is greater than the height, (2) the intra prediction mode before modification is equal to or greater than 2, and (3) the intra prediction mode is less than a value derived as (8+2*whRatio) when the variable whRatio is greater than 1 and is less than 8 when the variable whRatio is less than or equal to 1 (predModeIntra is less than (whRatio>1)?(8+2*whRatio):8), the intra prediction mode is set to a value 65 greater than predModeIntra [predModeIntra is set equal to (predModeIntra+65)].
[0218] If different from the above, that is, if conditions (1) to (3) are satisfied, (1) the height of the current block is greater than the width, (2) the intra-frame prediction mode before modification is less than or equal to 66, and (3) the intra-frame prediction mode is greater than the value derived as (60-2*whRatio) when whRatio is greater than 1 and is greater than 60 when whRatio is less than or equal to 1 (predModeIntra is greater than (whRatio>1)?(60-2*whRatio):60), then the intra-frame prediction mode is set to a value that is 67 less than predModeIntra [predModeIntra is set equal to (predModeIntra-67)].
[0219] Table 2 above shows how to select a transform set based on the intra prediction mode value extended by WAIP in LFNST. Figure 9 As shown, modes 14 to 33 and modes 35 to 80 are symmetric about the prediction direction around mode 34. For example, mode 14 and mode 54 are symmetric about the direction corresponding to mode 34. Therefore, the same set of transforms is applied to modes located in mutually symmetric directions, and this symmetry is also reflected in Table 2.
[0220] In addition, it is assumed that the forward LFNST input data of mode 54 is symmetric with the forward LFNST input data of mode 14. For example, for mode 14 and mode 54, according to Figure 7 (a) and Figure 7 The arrangement order shown in (b) rearranges the two-dimensional data into one-dimensional data. In addition, it can be seen that Figure 7 (a) and Figure 7 The pattern of the sequence shown in (b) is symmetrical about the direction indicated by the pattern 34 (diagonal direction).
[0221] Furthermore, as described above, which transform matrix of the [48×16] matrix and the [16×16] matrix is applied to the LFNST is determined by the size and shape of the transform target block.
[0222] Figure 10 is a diagram illustrating a block shape to which LFNST is applied. Figure 10 (a) shows a 4×4 block, Figure 10 (b) shows a 4×8 block and an 8×4 block, Figure 10 (c) shows a 4×N block or an N×4 block, where N is 16 or greater, Figure 10 (d) shows an 8×8 block, Figure 10 (e) shows an M×N block, where M≥8, N≥8, and N>8 or M>8.
[0223] exist Figure 10In , blocks with thick borders indicate the areas where LFNST is applied. Figure 10 For the blocks (a) and (b), LFNST is applied to the top left 4×4 region, and for Figure 10 In the block (c), LFNST is applied separately to the two upper left 4×4 regions that are arranged consecutively. Figure 10 In (a), (b), and (c), since LFNST is applied in units of 4×4 regions, this LFNST will be referred to as “4×4 LFNST” hereinafter. Depending on the matrix dimension of G, a [16×16] or [16×8] matrix can be applied.
[0224] More specifically, the [16×8] matrix is applied to Figure 10 (a) 4×4 block (4×4TU or 4×4CU), and the [16×16] matrix is applied to Figure 10 This is to adjust the worst-case computational complexity to 8 multiplications per sample.
[0225] about Figure 10 In (d) and (e), LFNST is applied to the upper left 8×8 region, and this LFNST is hereinafter referred to as "8×8LFNST". As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of forward LFNST, since a [48×1] vector (the X vector in Equation 9) is input as input data, not all sample values of the upper left 8×8 region are used as input values of the forward LFNST. That is, as can be seen from Figure 7 The left order of (a) or Figure 7 As can be seen from the left order of (b), a [48×1] vector can be constructed based on samples belonging to the remaining three 4×4 blocks while leaving the lower right 4×4 block as it is.
[0226] The [48×8] matrix can be applied to Figure 10 The 8×8 block (8×8TU or 8×8CU) in (d) and the [48×16] matrix can be applied to Figure 10 The 8×8 block in (e) is used to adjust the worst-case computational complexity to 8 multiplications per sample.
[0227] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data are generated (Y vector in Equation 9, [8×1] or [16×1] vector). In the forward LFNST, since the matrix G T The characteristic is that the amount of output data is equal to or less than the amount of input data.
[0228] Figure 11 is a diagram illustrating arrangement of output data of a forward LFNST according to an example, and shows blocks in which the output data of the forward LFNST is arranged according to block shapes.
[0229] exist Figure 11 The shaded area on the upper left of the block shown corresponds to the area where the output data of the forward LFNST is located, the positions marked with 0 indicate samples filled with the value 0, and the remaining area represents the area not changed by the forward LFNST. In the area not changed by the LFNST, the output data of the forward primary transform remains unchanged.
[0230] As described above, since the size of the applied transformation matrix varies according to the shape of the block, the amount of output data also varies. Figure 11 , the output data of the forward LFNST may not completely fill the upper left 4×4 block. Figure 11 In the cases of (a) and (d), the [16×8] matrix and the A[48×8] matrix are applied to the block indicated by the bold line or the partial area inside the block, respectively, and the [8×1] vector is generated as the output of the forward LFNST. That is, according to Figure 8 The scanning order shown in (b) can only fill 8 output data, such as Figure 11 As shown in (a) and (d), the remaining 8 positions can be filled with 0. Figure 10 (d) The case of LFNST application blocks, such as Figure 11 As shown in (d), the two 4×4 blocks on the upper right and lower left adjacent to the upper left 4×4 block are also filled with values of 0.
[0231] As described above, basically, by signaling the LFNST index, it is specified whether to apply LFNST and the transformation matrix to be applied. Figure 11 As shown, when LFNST is applied, since the number of output data of the forward LFNST may be equal to or less than the number of input data, an area filled with zero values occurs as follows.
[0232] 1) If Figure 11 As shown in (a), samples from the eighth position and subsequent positions in the scanning order in the upper left 4×4 block, that is, samples from the ninth to sixteenth positions.
[0233] 2) If Figure 11 As shown in (d) and (e), when the [48×16] matrix or the [48×8] matrix is applied, two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scanning order.
[0234] Therefore, if non-zero data exists by checking areas 1) and 2), it is determined that LFNST is not applied, so that signaling of the corresponding LFNST index can be omitted.
[0235] According to an example, for example, in the case of LFNST adopted in the VVC standard, since LFNST index signaling is performed after residual coding, the encoding device can determine whether non-zero data (significant coefficients) exist at all locations within the TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform LFNST index signaling based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. When non-zero data does not exist in the areas specified in 1) and 2) above, LFNST index signaling is performed.
[0236] Since truncated unary codes are applied as a binarization method for LFNST indexes, the LFNST index consists of up to two bins, and 0, 10, and 11 are assigned as binary codes for possible LFNST index values 0, 1, and 2, respectively. According to an example, context-based CABAC coding can be applied to the first bin (conventional coding), and context-based CABAC coding can also be applied to the second bin. The encoding of the LFNST index is shown in the following table.
[0237] [Table 4]
[0238]
[0239] As shown in Table 4, for the first bin (binIdx=0), context 0 is applied in the case of a single tree, while context 1 can be applied in the case of a non-single tree. In addition, as shown in Table 4, context 2 can be applied to the second bin (binIdx=1). That is, two contexts can be assigned to the first bin, one context can be assigned to the second bin, and each context can be distinguished by the ctxInc value (0, 1, 2).
[0240] Here, a single tree means that the luma component and the chroma component are encoded using the same coding structure. When a coding unit is divided while having the same coding structure, and the size of the coding unit becomes less than or equal to a specific threshold, and the luma component and the chroma component are encoded using separate tree structures, the corresponding coding unit is considered a dual tree, and therefore, the context of the first bin can be determined. That is, as shown in Table 4, context 1 can be assigned.
[0241] Alternatively, when the value of the variable treeType is assigned to SINGLE_TREE of the first bin, context 0 may be used, otherwise context 1 may be used.
[0242] In addition, for the adopted LFNST, the following simplified method can be applied.
[0243] (i) According to an example, the number of output data of the forward LFNST may be limited to a maximum of 16.
[0244] exist Figure 10 In the case of (c), 4×4 LFNST can be applied to two 4×4 regions adjacent to the upper left, respectively, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to a maximum of 16, in the case of a 4×N / N×4 (N≥16) block (TU or CU), 4×4 LFNST is applied only to one 4×4 region in the upper left, and LFNST can be applied only to Figure 10 By doing this, the implementation of image coding can be simplified.
[0245] Figure 12 It is shown that the number of output data of the forward LFNST according to the example is limited to a maximum of 16. Figure 12 , when LFNST is applied to the upper leftmost 4×4 region in a 4×N or N×4 block (where N is 16 or greater), the output data of the forward LFNST becomes 16.
[0246] (ii) According to an example, zeroing can be additionally applied to areas to which LFNST is not applied. In this document, zeroing can mean filling all positions belonging to a specific area with a value of 0. That is, zeroing can be applied to areas that are unchanged due to LFNST, and the result of the forward primary transform can be maintained. As described above, since LFNST is divided into 4×4 LFNST and 8×8 LFNST, zeroing can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.
[0247] (ii)-(A) When 4×4 LFNST is applied, an area to which 4×4 LFNST is not applied may be cleared. Figure 13 is a diagram illustrating clearing of zeros in a block to which 4×4 LFNST is applied according to an example.
[0248] like Figure 13 As shown, for the block to which 4×4 LFNST is applied, that is, for Figure 11 For all blocks in (a), (b), and (c), the entire area where LFNST is not applied can be filled with zeros.
[0249] on the other hand, Figure 13 (d) shows that when the maximum value of the number of output data of the forward LFNST is limited to 16 (as Figure 12), zeroing is performed on the remaining blocks to which 4×4 LFNST is not applied.
[0250] (ii)-(B) When 8×8 LFNST is applied, an area to which 8×8 LFNST is not applied may be cleared. Figure 14 is a diagram illustrating clearing of zeros in a block to which 8×8 LFNST is applied according to an example.
[0251] like Figure 14 As shown, for the block where 8×8 LFNST is applied, i.e., for Figure 11 For all blocks in (d) and (e), the entire area where LFNST is not applied can be filled with zeros.
[0252] (iii) Due to the zeroing presented in (ii) above, the area filled with zeros may not be the same as when LFNST is applied. Figure 11 In the case of LFNST, the wider region performs the zeroing proposed in (ii) to check whether there is non-zero data.
[0253] For example, when (ii)-(B) is applied, when checking Figure 11 After the zero-filled areas in (d) and (e) have non-zero data, additionally check Figure 14 Whether there is non-zero data in the area filled with 0, the signaling of the LFNST index can be performed only when there is no non-zero data.
[0254] Of course, even if the clearing proposed in (ii) is applied, the presence of non-zero data can be checked in the same way as the existing LFNST index signaling. Figure 11 After checking whether there is non-zero data in the block filled with zeros in , LFNST index signaling can be applied. In this case, the encoding device only performs zero clearing and the decoding device does not assume zero clearing, that is, it only checks whether non-zero data exists only in Figure 11 In the regions explicitly marked as 0 in , LFNST index parsing can be performed.
[0255] Various embodiments of applying the combination of the simplified methods ((i), (ii)-(A), (ii)-(B), (iii)) of LFNST can be derived. Of course, the combination of the simplified methods described above is not limited to the following embodiments, and any combination can be applied to LFNST.
[0256] Implementation Method
[0257] -Limit the number of output data of the forward LFNST to a maximum of 16 → (i)
[0258] - When 4×4 LFNST is applied, all regions to which 4×4 LFNST is not applied are cleared → (II)-(A)
[0259] - When 8×8 LFNST is applied, all regions where 8×8 LFNST is not applied are cleared → (II)-(B)
[0260] - After checking whether non-zero data also exists in the existing areas filled with zero values and the areas filled with zeros due to additional clearing ((ii)-(A), (ii)-(B)), signal the LFNST index only if no non-zero data exists → (iii).
[0261] In the case of an embodiment, when LFNST is applied, the area where non-zero output data can exist is limited to the interior of the upper left 4×4 area. Figure 13 (a) and Figure 14 In the case of (a), the eighth position in the scanning order is the last position where non-zero data can exist. Figure 13 (b) and (c) and Figure 14 In the case of (b), the sixteenth position in the scanning order (ie, the position of the lower right edge of the upper left 4×4 block) is the last position in which data other than 0 may exist.
[0262] Therefore, after applying LFNST, after checking whether non-zero data exists at a position not allowed by the residual encoding process (at a position beyond the last position), it may be determined whether to signal the LFNST index.
[0263] In the case of the zeroing method proposed in (ii), the amount of data ultimately generated when both the primary transform and LFNST are applied can be reduced, thereby reducing the amount of computation required to perform the entire transform process. Specifically, when LFNST is applied, since the output data from the forward primary transform is present in areas where LFNST is not applied, there is no need to generate data for areas that were zeroed during the forward primary transform. Consequently, the amount of computation required to generate the corresponding data can be reduced. Additional benefits of the zeroing method proposed in (ii) are summarized below.
[0264] First, as mentioned above, the amount of computation required to perform the entire transformation process is reduced.
[0265] In particular, when (ii)-(B) is applied, the worst-case computational effort is reduced, making the transform process lightweight. In other words, generally speaking, a large amount of computation is required to perform a large-scale transform. By applying (ii)-(B), the amount of data derived as a result of performing forward LFNST can be reduced to 16 or less. In addition, as the size of the entire block (TU or CU) increases, the effect of reducing the number of transform operations further increases.
[0266] Second, the amount of computation required for the entire transformation process can be reduced, thereby reducing the power consumption required to perform the transformation.
[0267] Third, the delay involved in the transformation process is reduced.
[0268] Secondary transforms such as LFNST add computational complexity to the existing primary transform, thus increasing the overall latency involved in performing the transform. In particular, in the case of intra prediction, since reconstructed data from neighboring blocks is used in the prediction process, the increased latency due to the secondary transform during encoding results in an increased latency until reconstruction. This can lead to an increase in the overall latency of intra prediction encoding.
[0269] However, if the clearing proposed in (ii) is applied, the delay time for performing one transform can be greatly reduced when LFNST is applied, maintaining or reducing the delay time of the entire transform, so that the encoding device can be implemented more simply.
[0270] In conventional intra prediction, the block currently to be encoded is regarded as one coding unit, and encoding is performed without segmentation. However, intra subpartitioning (ISP) encoding means performing intra prediction encoding by dividing the block currently to be encoded in the horizontal direction or the vertical direction. In this case, a reconstructed block can be generated by performing encoding / decoding in units of divided blocks, and the reconstructed block can be used as a reference block for the next divided block. According to an embodiment, in ISP encoding, one coding block can be divided into two or four sub-blocks and encoded, and in ISP, in one sub-block, intra prediction is performed with reference to the reconstructed pixel value of the sub-block located on the adjacent left side or the adjacent upper side. Hereinafter, "encoding" may be used as a concept including both encoding performed by an encoding device and decoding performed by a decoding device.
[0271] In addition, the signaling order of the LFNST index and the MTS index will be described below.
[0272] According to an example, the LFNST index signaled in the residual coding may be encoded after the encoding position of the last non-zero coefficient position, and the MTS index may be encoded immediately after the LFNST index. In this configuration, the LFNST index may be signaled for each transform unit. Alternatively, even if not signaled in the residual coding, the LFNST index may be encoded after the encoding of the last significant coefficient position, and the MTS index may be encoded after the LFNST index.
[0273] The syntax of residual coding according to the example is as follows.
[0274] [Table 5]
[0275]
[0276]
[0277] The meanings of the main variables shown in Table 5 are as follows.
[0278] 1.cbWidth, cbHeight: width and height of the current encoding block
[0279] 2. log2TbWidth, log2TbHeight: The base-2 logarithmic values of the width and height of the current transform block, which can be reduced to the upper left area where non-zero coefficients may exist by reflecting zeros.
[0280] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is enabled, if the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in a sequence parameter set (SPS).
[0281] 4. CuPredMode[chType][x0][y0]: A prediction mode corresponding to the variable chType and the (x0, y0) position. chType can have values of 0 and 1, where 0 indicates a luma component and 1 indicates a chroma component. The (x0, y0) position indicates a position on a picture, and MODE_INTRA (intra-frame prediction) and MODE_INTER (inter-frame prediction) can be used as the value of CuPredMode[chType][x0][y0].
[0282] 5. IntraSubPartitionsSplit[x0][y0]: The content of the (x0, y0) position is the same as in item 4. It indicates which ISP partition is applied at the (x0, y0) position, and ISP_NO_SPLIT indicates that the coding unit corresponding to the (x0, y0) position is not divided into partition blocks.
[0283] 6. intra_mip_flag[x0][y0]: The content of the position (x0, y0) is the same as in the above item 4. intra_mip_flag is a flag indicating whether the matrix-based intra prediction (MIP) prediction mode is applied. If the flag value is 0, it indicates that MIP is not enabled, and if the flag value is 1, it indicates that MIP is enabled.
[0284] 7. cIdx: A value of 0 indicates luma, and values of 1 and 2 indicate Cb and Cr of the chroma components, respectively.
[0285] 8.treeType: indicates single tree and dual tree, etc. (SINGLE_TREE: single tree, DUAL_TREE_LUMA: dual tree for luminance component, DUAL_TREE_CHROMA: dual tree for chrominance component)
[0286] 9.tu_cbf_cb[x0][y0]: The content of the (x0, y0) position is the same as in item 4. It indicates the coded block flag (CBF) of the Cb component. If its value is 0, it means that there is no non-zero coefficient in the corresponding transform unit of the Cb component, and if its value is 1, it indicates that there is a non-zero coefficient in the corresponding transform unit of the Cb component.
[0287] 10. lastSubBlock: This indicates the position of the subblock (coefficient group (CG)) where the last non-zero coefficient is located in the scanning order. 0 indicates a subblock containing a DC component, and in the case of greater than 0, it is not a subblock containing a DC component.
[0288] 11.lastScanPos: It indicates the position of the last significant coefficient in a sub-block according to the scanning order. If a sub-block includes 16 positions, it can have a value from 0 to 15.
[0289] 12.lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If not parsed, it is inferred to have a value of 0. That is, the default value is set to 0, indicating that LFNST is not applied.
[0290] 13. LastSignificantCoeffX, LastSignificantCoeffY: These indicate the x and y coordinates of the last significant coefficient in the transform block. The x coordinate starts at 0 and increases from left to right, and the y coordinate starts at 0 and increases from top to bottom. If the value of both variables is 0, it means that the last significant coefficient is located at DC.
[0291] 14. cu_sbt_flag: A flag indicating whether sub-block transform (SBT) included in the current VVC standard is enabled. If the flag value is 0, it indicates that SBT is not enabled, and if the flag value is 1, it indicates that SBT is enabled.
[0292] 15.sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: Flags indicating whether explicit MTS is applied to inter CUs and intra CUs, respectively. If the corresponding flag value is 0, it indicates that MTS is not enabled for inter CUs or intra CUs, and if the corresponding flag value is 1, it indicates that MTS is enabled.
[0293] 16.tu_mts_idx[x0][y0]: MTS index syntax element to be parsed. If not parsed, it is inferred to have a value of 0. That is, the default value is set to 0, indicating that DCT-2 is enabled in both horizontal and vertical directions.
[0294] As shown in Table 5, in the case of a single tree, the last significant coefficient position condition for luma can be used alone to determine whether to signal an LFNST index. That is, if the position of the last significant coefficient is not DC and the last significant coefficient exists in the top left sub-block (CG) (e.g., a 4×4 block), the LFNST index is signaled. In this case, in the case of a 4×4 transform block and an 8×8 transform block, the LFNST index is signaled only when the last significant coefficient exists at positions 0 to 7 in the top left sub-block.
[0295] In the case of a dual tree, the LFNST index is signaled independently for each of luma and chroma, and in the case of chroma, the LFNST index can be signaled by applying the last significant coefficient position condition only to the Cb component. For the Cr component, the corresponding condition may not be checked, and if the CBF value of Cb is 0, the LFNST index may be signaled by applying the last significant coefficient position condition to the Cr component.
[0296] “Min(log2TbWidth,log2TbHeight)>=2” in Table 5 can be expressed as “Min(tbWidth,tbHeight)>=4”, and “Min(log2TbWidth,log2TbHeight)>=4” can be expressed as “Min(tbWidth,tbHeight)>=16”.
[0297] In Table 5, log2ZoTbWidth and log2ZoTbHeight respectively mean the logarithmic values of the width and height with base 2 (base-2) of the upper left area where the last significant coefficient can exist by clearing to zero.
[0298] As shown in Table 5, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places: the first is before parsing the MTS index or LFNST index value, and the second is after parsing the MTS index.
[0299] The first update is before parsing the MTS index (tu_mts_idx[x0][y0]) value, so log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.
[0300] After parsing the MTS index, log2ZoTbWidth and log2ZoTbHeight are set for MTS indices greater than 0 (DST-7 / DCT-8 combination). When DST-7 / DCT-8 is applied independently to each of the horizontal and vertical directions in one transform, there can be up to 16 significant coefficients per row or column in each direction. That is, after applying DST-7 / DCT-8 of length 32 or greater, up to 16 transform coefficients can be derived for each row or column starting from the left or top. Therefore, in a 2D block, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions, significant coefficients can exist in only the upper left region of up to 16×16.
[0301] In addition, when DCT-2 is applied independently to each of the horizontal and vertical directions in the current transform, up to 32 significant coefficients can be present per row or column in each direction. That is, when a DCT-2 length of 64 or greater is applied, up to 32 transform coefficients can be derived for each row or column starting from the left or top. Therefore, in a 2D block, when DCT-2 is applied to both the horizontal and vertical directions, significant coefficients can exist only in the upper left region of up to 32×32.
[0302] In addition, when DST-7 / DCT-8 is applied on one side and DCT-2 is applied on the other side for the horizontal and vertical directions, 16 significant coefficients can be present in the forward direction and 32 significant coefficients can be present in the backward direction. For example, in the case of a 64×8 transform block, if DCT-2 is applied in the horizontal direction and DST-7 is applied in the vertical direction (which may occur when implicit MTS is applied), significant coefficients can be present in up to the upper left 32×8 area.
[0303] If log2ZoTbWidth and log2ZoTbHeight are updated in two places as shown in Table 5, that is, before parsing the MTS index, the range of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight, as shown in the following table.
[0304] [Table 6]
[0305]
[0306] Additionally, in this case, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix may be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values in the binarization process of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.
[0307] [Table 7]
[0308]
[0309] According to an example, in the case where the ISP mode and LFNST are applied, when the signaling of Table 5 is applied, the specification text may be configured as shown in Table 8. Compared with Table 5, the condition for signaling the LFNST index only when the ISP mode is not included is deleted (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT in Table 5).
[0310] In a single tree, when the LFNST index transmitted for the luma component (cIdx=0) is reused for the chroma component, the LFNST index transmitted for the first ISP partition block where a significant coefficient exists can be applied to the chroma transform block. Alternatively, even in a single tree, the LFNST index can be signaled for the chroma component separately from the LFNST index signaled for the luma component. The description of the variables in Table 8 is the same as that in Table 5.
[0311] [Table 8]
[0312]
[0313] According to an example, the LFNST index and / or MTS index can be signaled at the coding unit level. As described above, the LFNST index can have three values: 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate that the selected LFNST set includes the first and second candidates of the two LFNST core candidates, respectively. The LFNST index is encoded by truncated unary binarization, and the values 0, 1, and 2 can be encoded as bin strings of 0, 10, and 11, respectively.
[0314] According to an example, LFNST may be applied only when DCT-2 is applied to both the horizontal and vertical directions in one transform. Therefore, if the MTS index is signaled after the LFNST index is signaled, the MTS index may be signaled only when the LFNST index is 0, and when the LFNST index is not 0, DCT-2 may be applied to both the horizontal and vertical directions to perform one transform without signaling the MTS index.
[0315] The MTS index can have values 0, 1, 2, 3, and 4, where 0, 1, 2, 3, and 4 can indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, and DCT-8 / DCT-8 are applied to the horizontal and vertical directions, respectively. In addition, the MTS index can be encoded by truncated unary binarization, and the values 0, 1, 2, 3, and 4 can be encoded as bin strings of 0, 10, 110, 1110, and 1111, respectively.
[0316] The LFNST index and the MTS index may be signaled at the coding unit level, and the MTS index may be encoded sequentially after the LFNST index at the coding unit level. The coding unit syntax table used for this is as follows.
[0317] [Table 9]
[0318]
[0319] According to Table 9, the existing condition for checking whether the value of tu_mts_idx[x0][y0] is 0 in the condition for signaling lfnst_idx[x0][y0] (i.e., checking whether DCT-2 is applied in both the horizontal and vertical directions) is changed to a condition for checking whether the value of transform_skip_flag[x0][y0] is 0 (! transform_skip_flag[x0][y0]). transform_skip_flag[x0][y0] indicates whether the coding unit is encoded in transform skip mode in which transform is skipped, and this flag is signaled before the MTS index and LFNST index. That is, since lfnst_idx[x0][y0] is signaled before the value of tu_mtx_idx[x0][y0] is signaled, only the condition regarding the value of transform_skip_flag[x0][y0] can be checked.
[0320] As shown in Table 9, multiple conditions are checked when encoding tu_mts_idx[x0][y0], and as described above, tu_mts_idx[x0][y0] is signaled only when the value of lfnst_idx[x0][y0] is 0.
[0321] tu_cbf_luma[x0][y0] is a flag indicating whether there is a significant coefficient for the luma component, and cbWidth and cbHeight indicate the width and height of the coding unit of the luma component, respectively.
[0322] In Table 9, (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT) indicates that the ISP mode is not applied, and (!cu_sbt_flag) indicates that SBT is not applied.
[0323] According to Table 9, when both the width and height of the coding unit of the luma component are 32 or less, tu_mts_idx[x0][y0] is signaled, that is, whether MTS is applied is determined by the width and height of the coding unit of the luma component.
[0324] According to another example, when transform block (TU) tiling occurs (for example, when the maximum transform size is set to 32, a 64×64 coding unit is divided into four 32×32 transform blocks and encoded), the MTS index may be signaled based on the size of each transform block. For example, when both the width and height of the transform block are 32 or less, the same MTS index value may be applied to all transform blocks in the coding unit, thereby applying the same primary transform. In addition, when transform block tiling occurs, the value of tu_cbf_luma[x0][y0] in Table 9 may be the CBF value of the upper left transform block, or may be set to 1 when the CBF value of even one transform block among all transform blocks is 1.
[0325] According to an example, when the ISP mode is applied to the current block, LFNST may be applied, in which case Table 9 may be changed as shown in Table 10.
[0326] [Table 10]
[0327]
[0328] As shown in Table 10, even in ISP mode (IntraSubPartitionsSplitType!=ISP_NO_SPLIT), lfnst_idx[x0][y0] may be configured to be signaled, and the same LFNST index value may be applied to all ISP partition blocks.
[0329] In addition, as shown in Table 10, since tu_mts_idx[x0][y0] can be signaled only in modes other than the ISP mode (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT), the MTS index coding part is the same as that in Table 9.
[0330] As shown in Tables 9 and 10, when the MTS index is signaled immediately after the LFNST index, information about the primary transform cannot be obtained when performing residual coding. In other words, the MTS index is signaled after residual coding. Therefore, in the residual coding portion, the portion where zeroing is performed while retaining only 16 coefficients for DST-7 or DCT-8 with a length of 32 can be changed to the following Table 11.
[0331] [Table 11]
[0332]
[0333]
[0334] As shown in Table 11, in determining log2ZoTbWidth and log2ZoTbHeight (where log2ZoTbWidth and log2ZoTbHeight represent the base-2 logarithmic values of the width and height of the upper left area remaining after clearing is performed), checking the value of tu_mts_idx[x0][y0] may be omitted.
[0335] Binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 9 may be determined based on log2ZoTbWidth and log2ZoTbHeight as shown in Table 11.
[0336] Furthermore, as shown in Table 11, when log2ZoTbWidth and log2ZoTbHeight are determined in residual encoding, a condition for checking sps_mts_enable_flag may be added.
[0337] TR in Table 7 indicates a truncated Rice binarization method, and the last significant coefficient information may be binarized based on cMax and cRiceParam defined in Table 7.
[0338] According to an example, when the position information about the last significant coefficient of the luma transform block is recorded in the residual encoding process, the MTS index may be signaled as shown in Table 12.
[0339] [Table 12]
[0340]
[0341] In Table 12, LumaLastSignificantCoeffX and LumaLastSignificantCoeffY indicate the X coordinate and Y coordinate of the last significant coefficient position of the luma transform block, respectively. A condition that both LumaLastSignificantCoeffX and LumaLastSignificantCoeffY must be less than 16 is added to Table 12. When either is 16 or greater, DCT-2 is applied in both the horizontal and vertical directions, so the signaling of tu_mts_idx[x0][y0] is omitted and it can be inferred that DCT-2 is applied in both the horizontal and vertical directions.
[0342] When both LumaLastSignificantCoeffX and LumaLastSignificantCoeffY are less than 16, this means that the last significant coefficient exists in the upper left 16×16 area. In addition, this indicates that when DST-7 or DCT-8 with a length of 32 is applied in the current VVC standard, there is a possibility that zeroing is applied to retain only 16 transform coefficients from the left or top. Therefore, tu_mts_idx[x0][y0] can be signaled to indicate the transform kernel used for the primary transform.
[0343] Furthermore, according to another example, the coding unit syntax table, transform unit syntax table, and residual coding syntax table are as follows. According to Table 13, the MTS index is moved from the transform unit level to the coding unit level syntax, and the MTS index is signaled after the LFNST index signaling. Furthermore, the following constraint has been removed: LFNST is not allowed when the ISP is applied to a coding unit. The constraint that disallows LFNST when the ISP is applied to a coding unit is removed, allowing LFNST to be applied to all intra-predicted blocks. Furthermore, both the MTS index and the LFNST index are conditionally signaled at the last part of the coding unit level.
[0344] [Table 13]
[0345]
[0346] [Table 14]
[0347]
[0348] [Table 15]
[0349]
[0350] In Table 13, MtsZeroOutSigCoeffFlag is initially set to 1, and the value can be changed in residual coding in Table 15. When there is a significant coefficient in the area to be filled with 0 by clearing (LastSignificantCoeffX>15||LastSignificantCoeffY>15), the value of the variable MtsZeroOutSigCoeffFlag changes from 1 to 0, and in this case, the MTS index is not signaled, as shown in Table 15.
[0351] In addition, as shown in Table 13, when tu_cbf_luma[x0][y0] is 0, mts_idx[x0][y0] encoding can be omitted. That is, when the CBF value of the luma component is 0, since no transform is applied, there is no need to signal the MTS index, so the MTS index encoding can be omitted.
[0352] According to an example, the above technical features can be implemented using another conditional syntax. For example, after performing MTS, a variable indicating whether a significant coefficient exists in an area other than the DC area of the current block can be derived, and when the variable indicates that a significant coefficient exists in an area other than the DC area, an MTS index can be signaled. That is, the presence of a significant coefficient in an area other than the DC area of the current block indicates that the value of tu_cbf_luma[x0][y0] is 1, and in this case, an MTS index can be signaled.
[0353] This variable may be expressed as MtsDcOnly, and after the variable MtsDcOnly is initially set to 1 at the coding unit level, when it is determined that a significant coefficient exists in an area other than the DC area of the current block at the residual coding level, the value is changed to 0. When the variable MtsDcOnly is 0, the image information may be configured so that the MTS index is signaled.
[0354] When tu_cbf_luma[x0][y0] is 0, since the residual coding syntax is not called at the transform unit level of Table 14, the initial value of the variable MtsDcOnly is 1. In this case, since the variable MtsDcOnly is not changed to 0, the image information can be configured not to signal the MTS index. In other words, the MTS index is not parsed and signaled.
[0355] In addition, the decoding device may determine the color index cIdx of the transformation coefficient to derive the variable MtsZeroOutSigCoeffFlag of Table 15. The color index cIdx being 0 indicates a luma component.
[0356] According to an example, since MTS may be applied only to the luma component of the current block, the decoding apparatus may determine whether the color index is luma when deriving the variable MtsZeroOutSigCoeffFlag for determining whether to parse the MTS index.
[0357] The variable MtsZeroOutSigCoeffFlag is a variable indicating whether clearing is performed when MTS is applied. It indicates whether there is a transform coefficient in an area other than the upper left area where the last significant coefficient may exist due to clearing after MTS is performed (that is, an area other than the upper left 16×16 area). The variable MtsZeroOutSigCoeffFlag is initially set to 1 at the coding unit level, as shown in Table 13 (MtsZeroOutSigCoeffFlag=1), and when there is a transform coefficient in an area other than the 16×16 area, its value can be changed from 1 to 0 at the residual coding level, as shown in Table 15 (MtsZeroOutSigCoeffFlag=0). When the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.
[0358] As shown in Table 15, at the residual coding level, a non-cleared area in which non-zero transform coefficients may exist may be set depending on whether clearing of the accompanying MTS is performed, and even in this case, the color index (cIdx) is 0, the non-cleared area may be set to the upper left 16×16 area of the current block.
[0359] Thus, when deriving the variable for determining whether to parse the MTS index, whether the color component is luma or chroma is determined. However, since LFNST can be applied to both luma and chroma components of the current block, the color component is not determined when deriving the variable for determining whether to parse the LFNST index.
[0360] For example, Table 13 shows the variable LfnstZeroOutSigCoeffFlag, which can indicate whether zeroing is performed when LFNST is applied. The variable LfnstZeroOutSigCoeffFlag indicates whether a significant coefficient exists in the second region other than the first region in the upper left of the current block. This value is initially set to 1, and when a significant coefficient exists in the second region, this value can be changed to 0. Only when the value of the initially set variable LfnstZeroOutSigCoeffFlag remains 1 can the LFNST index be parsed. When determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, the color index of the current block is not determined because LFNST can be applied to both the luma component and the chroma component of the current block.
[0361] Figure 15 A CCLM applicable when deriving an intra prediction mode of a chroma block according to an embodiment is illustrated.
[0362] In this specification, a "reference sample template" may refer to a set of adjacent reference samples of a current chroma block used to predict the current chroma block. The reference sample template may be predefined, and information about the reference sample template may be signaled from the encoding device 200 to the decoding device 300.
[0363] Reference Figure 15 , the set of shaded samples in a single row adjacent to the 4×4 block as the current chroma block refers to the reference sample template. The reference sample template is configured as a single row of reference samples, while the reference sample area in the luma area corresponding to the reference sample template is configured as two rows, as shown in Figure 15 shown.
[0364] In an embodiment, when performing intra-frame coding of a chroma image in the Joint Exploration Test Model (JEM) used in the Joint Video Exploration Team (JVET), a cross-component linear model (CCLM) may be used. The CCLM is a method for predicting pixel values of a chroma image based on pixel values of a reconstructed luminance image, and is based on the high correlation between the luminance image and the chroma image.
[0365] CCLM prediction of Cb and Cr chrominance images can be performed based on the following formula.
[0366] [Equation 11]
[0367] Pred C (i, j) = α·Rec′ L (i, j)+β
[0368] Here, Pred c (i, j) represents the Cb or Cr chrominance image to be predicted, Rec L '(i,j) represents the reconstructed luminance image adjusted to the chrominance block size, and (i,j) represents the coordinates of the pixel. In the 4:2:0 color format, since the size of the luminance image is twice that of the chrominance image, it is necessary to generate a Rec image with the chrominance block size by downsampling. L ', so it can be considered that Rec L (2i,2j) and adjacent pixels to be used for the chrominance image Pred c The pixel of the brightness image of (i,j). Rec L '(i,j) can be called the downsampled brightness sample.
[0369] For example, as shown in the following equation, six adjacent pixels can be used to derive Rec L '(i,j).
[0370] [Equation 12]
[0371] Rec′ L(x,y)=(2×Rec L (2x,2y)+2×Rec L (2x,2y+1)+Rec L (2x-1, 2y)+Rec L (2x+1,2y)+Rec L (2x-1,2y+1)+Rec L (2x+1,2y+1)+4)>>3
[0372] α and β represent the adjacent templates of Cb or Cr chrominance blocks and Figure 15 The cross-correlation and average difference between adjacent templates of the luminance block in the middle shadow area. For example, α and β are expressed by Equation 13.
[0373] [Equation 13]
[0374]
[0375] L(n) represents the adjacent reference samples and / or left adjacent samples of the luminance block corresponding to the current chrominance image, C(n) represents the adjacent reference samples and / or left adjacent samples of the current chrominance block to which encoding is currently applied, and (i, j) represents the pixel position. In addition, L(n) can represent the downsampled upper adjacent samples and / or left adjacent samples of the current luminance block. N can represent the total number of pixel pair (luminance and chrominance) values used to calculate the CCLM parameters, and can indicate a value that is twice the smaller value of the width and height of the current chrominance block.
[0376] A picture may be divided into a sequence of coding tree units (CTUs). A CTU may correspond to a coding tree block (CTB). Alternatively, a CTU may include a coding tree block for luma samples and a coding tree block for corresponding chroma samples. Depending on whether the luma block and the corresponding chroma block have separate partition structures, the tree type may be classified as a single tree (SINGLE_TREE) or a dual tree (DUAL_TREE). A single tree may indicate that the chroma block has the same partition structure as the luma block, while a dual tree may indicate that the chroma component block has a partition structure different from that of the luma block.
[0377] When LFNST is applied to a chroma transform block according to an example, it is necessary to refer to information about a collocated luma transform block.
[0378] The existing normative texts for the relevant parts are shown in the table below.
[0379] [Table 16]
[0380]
[0381] As shown in Table 16, when the current intra prediction mode is CCLM mode, the value of the variable predModeIntra for the chroma transform block (in italics) is determined by taking the intra prediction mode value of the co-located chroma transform block. The intra prediction mode value (predModeIntra value) of the luma transform block can then be used to determine the LFNST set.
[0382] However, the variables nTbW and nTbH input as input values for the transform process represent the width and height of the current transform block. Therefore, when the current block is a luma transform block, the variables nTbW and nTbH may represent the width and height of the luma transform block, whereas when the current block is a chroma transform block, the variables nTbW and nTbH represent the width and height of the chroma transform block.
[0383] Here, the variables nTbW and nTbH in the italic portion of Table 16 represent the width and height of the chroma transform block, which do not reflect the color format and therefore do not accurately indicate the reference position of the luma transform block corresponding to the chroma transform block. Therefore, the italic portion of Table 16 can be modified as shown in the following table.
[0384] [Table 17]
[0385]
[0386] As shown in Table 17, nTbW and nTbH are changed to (nTbW*SubWidthC) / 2 and (nTbH*SubHeightC) / 2, respectively. xTbY and yTbY may represent the luma position in the current picture (the top left sample of the current luma transform block relative to the top left luma sample of the current picture), while nTbW and nTbH may represent the width and height of the transform block currently being encoded (the variable nTbW specifies the width of the current transform block, and the variable nTbH specifies the height of the current transform block).
[0387] When the currently encoded transform block is a chroma (Cb or Cr) transform block, nTbW and nTbH are the width and height of the chroma transform block, respectively. Therefore, when the currently encoded transform block is a chroma transform block (cIdx>0), the width and height of the luma transform block are used to obtain the reference position of the collocated luma transform block. In Table 17, SubWidthC and SubHeightC are values set according to the color format (e.g., 4:2:0, 4:2:2, or 4:4:4), specifically, the width ratio and height ratio between the luma component and the chroma component, respectively (see Table 18 below). Therefore, in the case of a chroma transform block, (nTbW*SubWidthC) and (nTbH*SubHeightC) can be the width and height relative to the collocated luma transform block, respectively.
[0388] Therefore, xTbY+(nTbW*SubWidthC) / 2 and yTbY+(nTbH*SubHeightC) / 2 represent values of a central position in a collocated luma transform block based on the upper left position of the current picture, and thus accurately indicate the collocated luma transform block.
[0389] [Table 18]
[0390] Chroma format SubWidthC SubHeightC monochrome 1 1 4:2:0 2 2 4:2:2 2 1 4:4:4 1 1 4:4:4 1 1
[0391] In Table 17, the variable predModeIntra represents an intra-frame prediction mode value. If the value of the variable predModeIntra is equal to INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM, it indicates that the current transform block is a chroma transform block. According to an example, in the current VVC standard, INTRA_LT_CCLM, INTRA_L_CCLM, and INTRA_T_CCLM correspond to mode values 81, 82, and 83, respectively, among the intra-frame prediction mode values. Therefore, as shown in Table 17, it is necessary to use the value of xTbY+(nTbW*SubWidthC) / 2 and the value of yTbY+(nTbH*SubHeightC) / 2 to obtain the reference position of the collocated luma transform block.
[0392] As shown in Table 17, the value of predModeIntra is updated in view of both the variable intra_mip_flag[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] and the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2].
[0393] intra_mip_flag is a variable indicating whether the current transform block (or coding unit) is encoded by a matrix-based intra prediction (MIP) method, and intra_mip_flag[x][y] is a flag value indicating whether MIP is applied to a position corresponding to the coordinates (x, y) based on the luma component when the upper left position of the current picture is defined as (0, 0). The x and y coordinates increase from left to right and from top to bottom, respectively, and when the flag indicating whether MIP is applied is 1, the flag indicates that MIP is applied. When the flag indicating whether MIP is applied is 0, the flag indicates that MIP is not applied. MIP can be applied only to luma blocks.
[0394] According to the modified portion of Table 17, when the value of intra_mip_flag[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] in the collocated luma transform block is 1, the value of predModeIntra is set to the planar mode (INTRA_PLANAR).
[0395] The value of the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents a prediction mode value corresponding to the coordinates (xTbY+(nTbW*SubWidthC) / 2, yTbY+(nTbH*SubHeightC) / 2) when the upper left position of the current picture of the luma component is defined as (0, 0). The prediction mode value may have MODE_INTRA, MODE_IBC, MODE_PLT, and MODE_INTER values representing an intra prediction mode, an intra block copy (IBC) prediction mode, a palette (PLT) coding mode, and an inter prediction mode, respectively. According to Table 17, when the value of the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] is MODE_IBC or MODE_PLT, the value of the variable predModeIntra is set to the DC mode. In other cases, the value of the variable predModeIntra is set to IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] (the intra-frame prediction mode value corresponding to the center position in the juxtaposed luma transform block).
[0396] According to an example, as shown in the following table, considering whether wide-angle intra prediction is performed, the value of the variable predModeIntra may be updated once more based on the predModeIntra value updated in Table 17.
[0397] [Table 19]
[0398]
[0399] The input values of predModeIntra, nTbW, and nTbH in the mapping process shown in Table 19 are the same as the values of the variable predModeIntra updated in Table 17 and nTbW and nTbH referenced in Table 17, respectively.
[0400] In Table 19, nCbW and nCbH represent the width and height of the coding block corresponding to the transform block, respectively, and the variable IntraSubPartitionsSplitType indicates whether the ISP mode is applied, where IntraSubPartitionsSplitType is equal to ISP_NO_SPLIT, indicating that the coding unit is not partitioned by ISP (i.e., the ISP mode is not applied). The variable IntraSubPartitionsSplitType not equal to ISP_NO_SPLIT indicates that the ISP mode is applied, so the coding unit is partitioned into two or four partition blocks. In Table 19, cIdx is an index indicating a color component. A cIdx value equal to 0 indicates a luma block, while a cIdx value not equal to 0 indicates a chroma block. The predModeIntra value output by the mapping process of Table 19 is a value updated taking into account whether the wide-angle intra prediction (WAIP) mode is applied.
[0401] For the predModeIntra value updated by Table 19, the LFNST set can be determined by the mapping relationship shown in the following table.
[0402] [Table 20]
[0403] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0 2<=predModeIntra<=12 1 13<=predModeIntra<=23 2 24<=predModeIntra<=44 3 45<=predModeIntra<=55 2 56<=predModeIntra<=80 1
[0404] In the above table, lfnstTrSetIdx represents an index indicating an LFNST set and has a value from 0 to 3, which indicates that a total of four LFNST sets are configured. Each LFNST set may include two transform cores, i.e., LFNST cores (based on the forward direction of the area where LFNST is applied, the transform core may be a 16×16 matrix or a 16×48 matrix), and the transform core to be applied among the two transform cores may be specified by signaling of the LFNST index. In addition, whether LFNST is applied may also be specified by the LFNST index. In the current VVC standard, the LFNST index may have values 0, 1, and 2, 0 indicating that LFNST is not applied, and 1 and 2 indicating two transform cores, respectively.
[0405] The following figures are provided to describe specific examples of the present disclosure. Since the specific names of the devices or the names of the specific signals / messages / fields illustrated in the figures are provided for example, the technical features of the present disclosure are not limited to the specific names used in the following figures.
[0406] Figure 16 is a flowchart illustrating the operation of a video decoding apparatus according to an embodiment of the present disclosure.
[0407] Figure 16 Each process disclosed in the reference Figures 4 to 15Therefore, some details will be omitted or will be schematically described with reference to Figures 3 to 15 The details described overlap the specific details of the description.
[0408] The decoding apparatus 300 according to an embodiment may obtain intra prediction mode information and an LFNST index from a bitstream ( S1610 ).
[0409] The intra-frame prediction mode information may include the intra-frame prediction mode of the neighboring blocks (e.g., the left neighboring block and / or the upper neighboring block) of the current block, and the MPM index indicating one of the most probable modes (MPM) candidates in the MPM list derived based on the additional candidate mode or the remaining intra-frame prediction mode information indicating one of the remaining intra-frame prediction modes not included in the MPM candidates.
[0410] In addition, the intra mode information may include flag information sps_cclm_enabled_flag indicating whether CCLM is applied to the current block and information intra_chroma_pred_mode about the intra prediction mode of the chroma component.
[0411] The LFNST index information is received as syntax information, and the syntax information is received as a binarized bin string containing 0 and 1.
[0412] The syntax element of the LFNST index according to this embodiment may indicate whether inverse LFNST or inverse inseparable transform is applied and any one of the transform kernel matrices included in the transform set, and when the transform set includes two transform kernel matrices, the syntax element of the transform index may have three values.
[0413] That is, according to an embodiment, the value of the syntax element of the LFNST index may include: 0, which indicates that inverse LFNST is not applied to the target block; 1, which indicates the first transform kernel matrix among the transform kernel matrices; and 2, which indicates the second transform kernel matrix among the transform kernel matrices.
[0414] The decoding device 300 can decode information about the quantized transform coefficients of the current block from the bitstream, and can derive the quantized transform coefficients of the target block based on the information about the quantized transform coefficients of the current block. The information about the quantized transform coefficients of the target block may be included in a sequence parameter set (SPS) or a slice header, and may include information about whether to apply RST, information about a reduction factor, information about a minimum transform size for applying RST, information about a maximum transform size for applying RST, an inverse RST size, and at least one of information about a transform index indicating any one of transform kernel matrices included in a transform set.
[0415] The decoding apparatus 300 may induce transform coefficients by dequantizing residual information (ie, quantized transform coefficients) about the current block, and may arrange the derived transform coefficients in a predetermined scanning order.
[0416] Specifically, the derived transform coefficients can be arranged in units of 4×4 blocks according to the inverse diagonal scanning order, and the transform coefficients in the 4×4 blocks can also be arranged according to the inverse diagonal scanning order. That is, the dequantized transform coefficients can be arranged according to the inverse scanning order applied in the video codec (such as in VVC or HEVC).
[0417] The transform coefficient derived based on the residual information may be a dequantized transform coefficient as described above, or may be a quantized transform coefficient. That is, the transform coefficient may be any data used to check whether there is non-zero data in the current block, regardless of quantization.
[0418] The decoding device can update the intra-frame prediction mode of the chroma block according to the intra-frame prediction mode of the luminance block corresponding to the chroma block based on the intra-frame prediction mode of the chroma block being the CCLM mode. In particular, based on the intra-frame prediction mode of the luminance block being the intra-frame block copy (IBC) mode, the intra-frame prediction mode of the chroma block can be updated to the intra-frame DC mode (S1620).
[0419] The decoding device may derive the intra-frame prediction mode of the chroma block as the CCLM mode based on the intra-frame prediction mode information. For example, the decoding device may receive information about the intra-frame prediction mode of the current chroma block through a bitstream, and may derive the intra-frame prediction mode of the current chroma block as the CCLM mode based on the intra-frame prediction mode information.
[0420] The CCLM mode may include an upper left CCLM mode, an upper CCLM mode, or a left CCLM mode.
[0421] As described above, the decoding apparatus may induce residual samples by applying LFNST as an inseparable transform or MTS as a separable transform, and may perform these transforms based on an LFNST index indicating an LFNST kernel (ie, an LFNST matrix) and an MTS index indicating an MTS kernel, respectively.
[0422] For LFNST, it is necessary to determine an LFNST set, and the LFNST set has a mapping relationship with the intra prediction mode of the current block.
[0423] The decoding apparatus may update the intra prediction mode of the chroma block based on the intra prediction mode of the luma block corresponding to the chroma block for inverse LFNST of the chroma block.
[0424] According to an example, the updated intra prediction mode may be derived as an intra prediction mode corresponding to a specific position in the luma block, and the specific position may be set based on a color format of the chroma block.
[0425] The specific position may be a central position of the luminance block and may be expressed by ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)).
[0426] In the center position, xTbY and yTbY represent the upper left coordinates of the luma block, that is, the upper left position in the luma sample reference of the current transform block, nTbW and nTbH represent the width and height of the chroma block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)) represents the center position of the luma transform block, and IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra prediction mode of the luma block at that position.
[0427] SubWidthC and SubHeightC can be derived as shown in Table 18. That is, when the color format is 4:2:0, SubWidthC and SubHeightC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0428] As shown in Table 17, in order to specify a specific position of a luminance block corresponding to a chrominance block regardless of a color format, the color format is reflected in a variable indicating the specific position.
[0429] As described above, when the intra prediction mode of the luminance block corresponding to a specific position is the IBC mode, the decoding apparatus may update the updated intra prediction mode to the intra DC mode.
[0430] IBC essentially performs prediction within the current picture, but can be performed similarly to inter-frame prediction, except that the reference block is derived within the current picture. That is, IBC can use at least one of the inter-frame prediction techniques described in this document.
[0431] Alternatively, according to an example, when the intra prediction mode corresponding to the specific position is the palette mode, the decoding apparatus may set the updated intra prediction mode to the intra DC mode.
[0432] IBC prediction mode or palette mode can be used to encode content images / videos including games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter-frame prediction, except that the reference block is derived within the current picture. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When palette mode is applied, the value of the sample in the picture can be signaled based on information about the palette table and palette index.
[0433] According to another example, when the intra prediction mode of the luminance block corresponding to a specific position is a matrix-based intra prediction (hereinafter, "MIP") mode, the decoding apparatus may set the updated intra prediction mode to an intra planar mode.
[0434] The MIP mode may be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). When MIP is applied to the current block, the prediction samples of the current block may be derived by i) using adjacent reference samples that have undergone averaging, ii) performing matrix-vector multiplication, and iii) further performing horizontal / vertical interpolation.
[0435] In summary, when the intra prediction mode of the central position is the MIP mode, the IBC mode, or the palette mode, the intra prediction mode of the chroma block may be updated to a specific mode, such as the intra planar mode or the intra DC mode.
[0436] When the intra prediction mode of the central position is not MIP mode, IBC mode, and palette mode, the intra prediction mode of the chroma block can be updated to the intra prediction mode for the luminance block of the central position to reflect the association between the chroma block and the luminance block.
[0437] The decoding apparatus may determine an LFNST set including an LFNST matrix based on the updated intra prediction mode ( S1630 ), and may induce a transform coefficient by performing LFNST on the chroma block based on the LFNST matrix derived from the LFNST set ( S1640 ).
[0438] Any one of a plurality of LFNST matrices may be selected based on the LFNST set and the LFNST index.
[0439] As shown in Table 20, the LFNST transform set is derived according to the intra prediction mode, and 81 to 83 indicating the CCLM mode in the intra prediction mode are omitted because the LFNST transform set is derived using the intra mode value of the corresponding luma block in the CCLM mode.
[0440] According to an example, as shown in Table 20, any one of four LFNST sets may be determined according to the intra prediction mode of the current block, and the LFNST set to be applied to the current chroma block may also be determined.
[0441] The decoding apparatus may perform inverse RST (eg, inverse LFNST) by applying the LFNST matrix to the dequantized transform coefficients, thereby deriving modified transform coefficients of the current chroma block.
[0442] The decoding apparatus may derive residual samples from the transform coefficients through one inverse transform (S1650), and when the current block is a chroma block, may derive residual samples of the chroma block based on the transform coefficients. MTS may be used for one inverse transform.
[0443] In addition, the decoding device may generate reconstructed samples based on the residual samples of the current block and the predicted samples of the current block.
[0444] The following figures are provided to describe specific examples of the present disclosure. Since the specific names of the devices or the names of the specific signals / messages / fields illustrated in the figures are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following figures.
[0445] Figure 17 is a flowchart illustrating the operation of a video encoding apparatus according to an embodiment of the present disclosure.
[0446] Figure 17 Each process disclosed in the reference Figures 4 to 15 Therefore, some details will be omitted or will be schematically described with reference to Figure 2 and Figures 4 to 15 The details described overlap the specific details of the description.
[0447] According to an embodiment, the encoding apparatus 200 may derive a prediction sample of the chroma block based on that the intra prediction mode of the chroma block is the CCLM mode ( S1710 ).
[0448] The encoding device may first derive the intra prediction mode of the chroma block as the CCLM mode.
[0449] For example, the encoding device may determine the intra prediction mode of the current chroma block based on the rate-distortion (RD) cost (or RDO). Here, the RD cost may be derived based on the sum of absolute differences (SAD). The encoding device may determine the CCLM mode as the intra prediction mode of the current chroma block based on the RD cost.
[0450] The CCLM mode may include an upper left CCLM mode, an upper CCLM mode, or a left CCLM mode.
[0451] The encoding device may encode information about the intra prediction mode of the current chroma block and may signal the information about the intra prediction mode through a bitstream. The prediction-related information about the current chroma block may include information about the intra prediction mode.
[0452] According to an embodiment, the encoding apparatus may induce residual samples of the chroma block based on the prediction samples ( S1720 ).
[0453] According to an embodiment, the encoding apparatus may induce transform coefficients of a chroma block based on one transform of residual samples.
[0454] One transform may be performed by multiple transform kernels, in which case the transform kernel may be selected based on the intra prediction mode.
[0455] The encoding device may update the intra prediction mode of the chroma block for LFNST of the chroma block based on the intra prediction mode of the luminance block corresponding to the chroma block, and may update the intra prediction mode of the chroma block to an intra DC mode based on that the intra prediction mode of the luminance block is an intra block copy (IBC) mode (S1730).
[0456] As shown in Table 17, the encoding device can update the CCLM mode of the chroma block based on the intra prediction mode of the luma block corresponding to the chroma block (when predModeIntra is equal to INTRA_LT_CCLM, INTRA_L_CCLM or INTRA_T_CCLM, predModeIntra is derived as follows:).
[0457] According to an example, the updated intra prediction mode may be derived as an intra prediction mode corresponding to a specific position in the luma block, and the specific position may be set based on a color format of the chroma block.
[0458] The specific position may be a central position of the luminance block and may be expressed by ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)).
[0459] In the center position, xTbY and yTbY represent the upper left coordinates of the luma block, that is, the upper left position in the luma sample reference of the current transform block, nTbW and nTbH represent the width and height of the chroma block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)) represents the center position of the luma transform block, and IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra prediction mode of the luma block at that position.
[0460] SubWidthC and SubHeightC can be derived as shown in Table 18. That is, when the color format is 4:2:0, SubWidthC and SubHeightC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0461] As shown in Table 17, in order to specify a specific position of a luminance block corresponding to a chrominance block regardless of a color format, the color format is reflected in a variable indicating the specific position.
[0462] As described above, when the intra prediction mode of the luminance block corresponding to a specific position is the IBC mode, the encoding apparatus may update the updated intra prediction mode to the intra DC mode.
[0463] IBC essentially performs prediction within the current picture, but can be performed similarly to inter-frame prediction, except that the reference block is derived within the current picture. That is, IBC can use at least one of the inter-frame prediction techniques described in this document.
[0464] Alternatively, according to an example, when the intra prediction mode corresponding to the specific position is the palette mode, the decoding apparatus may set the updated intra prediction mode to the intra DC mode.
[0465] IBC prediction mode or palette mode can be used to encode content images / videos including games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter-frame prediction, except that the reference block is derived within the current picture. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When palette mode is applied, the value of the sample in the picture can be signaled based on information about the palette table and palette index.
[0466] According to another example, when the intra prediction mode of the luminance block corresponding to a specific position is a matrix-based intra prediction (hereinafter, "MIP") mode, the encoding apparatus may set the updated intra prediction mode to an intra planar mode.
[0467] The MIP mode may be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). When MIP is applied to the current block, the prediction samples of the current block may be derived by i) using adjacent reference samples that have undergone averaging, ii) performing matrix-vector multiplication, and iii) further performing horizontal / vertical interpolation.
[0468] In summary, when the intra prediction mode of the central position is the MIP mode, the IBC mode, or the palette mode, the intra prediction mode of the chroma block may be updated to a specific mode, such as the intra planar mode or the intra DC mode.
[0469] When the intra prediction mode of the central position is not MIP mode, IBC mode, and palette mode, the intra prediction mode of the chroma block can be updated to the intra prediction mode for the luma block at the central position to reflect the association between the chroma block and the luma block.
[0470] The encoding apparatus may determine an LFNST set including an LFNST matrix based on the updated intra prediction mode ( S1740 ), and may derive a modified transform coefficient by performing LFNST on the chroma block based on the residual sample and the LFNST matrix ( S1750 ).
[0471] The encoding device may determine a transform set based on a mapping relationship according to an intra prediction mode applied to a current block, and may perform LFNST, ie, an inseparable transform, based on any one of two LFNST matrices included in the transform set.
[0472] As described above, a plurality of transform sets may be determined according to the intra prediction mode of the transform block to be transformed.The matrix applied to LFNST is the transpose of the matrix used in inverse LFNST.
[0473] In one example, the LFNST matrix may be a non-square matrix with a smaller number of rows than columns.
[0474] The encoding apparatus may derive a quantized transform coefficient by performing quantization based on the modified transform coefficient of the current chroma block, and may encode and output image information including information about the quantized transform coefficient, information about the intra prediction mode, and an LFNST index indicating an LFNST matrix ( S1760 ).
[0475] Specifically, the encoding apparatus 200 may generate information about the quantized transform coefficient and may encode the generated information about the quantized transform coefficient.
[0476] In one example, the information about the quantized transform coefficient may include at least one of information about whether LFNST is applied, information about a reduction factor, information about a minimum transform size for applying LFNST, and information about a maximum transform size for applying LFNST.
[0477] The encoding apparatus may encode flag information indicating whether CCLM is applied to the current block as sps_cclm_enabled_flag and information about the intra prediction mode of the chroma component as intra_chroma_pred_mode as the information about the intra mode.
[0478] The information on the CCLM mode as intra_chroma_pred_mode may indicate the upper-left CCLM mode, the upper CCLM mode, or the left CCLM mode.
[0479] In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency.
[0480] In addition, in the present disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled through residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.
[0481] In the above embodiments, the method is explained based on a flowchart with the aid of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may be performed in an order or step different from the above order or step, or a certain step may be performed concurrently with other steps. In addition, it will be understood by those skilled in the art that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0482] The above-mentioned method according to the present disclosure may be implemented in a software form, and the encoding device and / or decoding device according to the present disclosure may be included in a device for image processing such as a television, a computer, a smart phone, a set-top box, and a display device.
[0483] When the embodiments in the present disclosure are implemented by software, the above methods can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium and / or other storage devices. In other words, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units shown in each of the accompanying drawings can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip.
[0484] In addition, the decoding device and encoding device to which the present disclosure is applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device (such as video communication), a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0485] In addition, the processing method applied to the present invention can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network. In addition, the embodiments of the present invention can be implemented as a computer program product through program code, and the program code can be executed on a computer according to the embodiments of the present invention. The program code can be stored on a computer-readable carrier.
[0486] Figure 18 The structure of the content streaming system to which the present disclosure is applied is illustrated.
[0487] Furthermore, a content streaming system to which the present disclosure is applied may generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0488] The encoding server is used to compress content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream, and then transmit the bitstream to the streaming server. As another example, if the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. The streaming server can also temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0489] The streaming server transmits multimedia data to user devices via a web server based on user requests. The web server serves as a tool for notifying users of available services. When a user requests a desired service, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case controls commands and responses between the various devices in the content streaming system.
[0490] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to smoothly provide a streaming service, the streaming server can store the bitstream for a predetermined time.
[0491] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system may operate as a distributed server, and in this case, data received by each server may be processed in a distributed manner.
[0492] The claims disclosed herein may be combined in various ways. For example, the technical features of the method claims of the present disclosure may be combined to be implemented or performed in a device, and the technical features of the device claims may be combined to be implemented or performed in a method. Furthermore, the technical features of method claims and device claims may be combined to be implemented or performed in a device, and the technical features of method claims and device claims may be combined to be implemented or performed in a method.
Claims
1. An image decoding device, comprising: Memory; as well as at least one processor coupled to the memory, wherein the at least one processor is configured to: Obtain intra prediction mode information and low-frequency inseparable transform LFNST index from the bitstream; The intra prediction mode based on the chroma block is a cross-component linear model (CCLM) mode, and the intra prediction mode of the chroma block is updated according to the intra prediction mode of the luminance block corresponding to the chroma block; determining an LFNST set including an LFNST matrix based on the updated intra prediction mode; and performing LFNST on the chroma block based on the LFNST matrix derived from the LFNST set, wherein the updated intra-frame prediction mode is derived as an intra-frame prediction mode corresponding to a specific position in the luminance block, Wherein, based on the prediction mode corresponding to the specific position being the intra block copy (IBC) mode, the intra prediction mode of the chroma block is updated to the intra DC mode, and The specific position is set based on the color format of the chroma block.
2. The image decoding device according to claim 1, wherein The specific position is a central position of the luminance block.
3. The image decoding device according to claim 2, wherein The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)), Wherein, xTbY and yTbY represent the upper left coordinates of the luminance block, Where nTbW and nTbH represent the width and height of the chroma block, and Wherein, SubWidthC and SubHeightC represent variables corresponding to the color format.
4. The image decoding device according to claim 3, wherein When the color format is 4:2:0, SubWidthC and SubHeightC are 2, and When the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1. The image decoding device according to claim 1 , wherein: When the prediction mode corresponding to the specific position is a matrix-based intra prediction MIP mode, the intra prediction mode of the chroma block is updated to an intra planar mode. The image decoding device according to claim 1 , wherein: When the prediction mode corresponding to the specific position is the palette mode, the intra prediction mode of the chroma block is updated to the intra DC mode.
7. An image encoding device, comprising: Memory; as well as at least one processor coupled to the memory, wherein the at least one processor is configured to: deriving prediction samples for the chroma block based on that the intra prediction mode for the chroma block is a cross-component linear model CCLM; deriving residual samples for the chroma block based on the prediction samples; updating the intra prediction mode of the chroma block based on the intra prediction mode of the luminance block corresponding to the chroma block; determining a LFNST set including a low frequency inseparable transform LFNST matrix based on the updated intra prediction mode; and performing LFNST on the chroma block based on the residual samples and a LFNST matrix, wherein the updated intra-frame prediction mode is derived as an intra-frame prediction mode corresponding to a specific position in the luminance block, Wherein, based on the prediction mode corresponding to the specific position being the intra block copy (IBC) mode, the intra prediction mode of the chroma block is updated to the intra DC mode, and The specific position is set based on the color format of the chroma block.
8. The image encoding device according to claim 7, wherein The specific position is a central position of the luminance block.
9. The image encoding device according to claim 8, wherein The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)), Wherein, xTbY and yTbY represent the upper left coordinates of the luminance block, Where nTbW and nTbH represent the width and height of the chroma block, and Wherein, SubWidthC and SubHeightC represent variables corresponding to the color format.
10. The image encoding device according to claim 9, wherein When the color format is 4:2:0, SubWidthC and SubHeightC are 2, and When the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
11. The image encoding device according to claim 7, wherein When the prediction mode corresponding to the specific position is a matrix-based intra prediction MIP mode, the intra prediction mode of the chroma block is updated to an intra planar mode.
12. The image encoding device according to claim 7, wherein When the prediction mode corresponding to the specific position is the palette mode, the intra prediction mode of the chroma block is updated to the intra DC mode.
13. A device for transmitting data for image information, the device comprising: At least one processor configured to generate a bitstream for the image information, wherein the bitstream is generated based on the following operations: deriving prediction samples for a chroma block based on an intra prediction mode for the chroma block being a cross-component linear model (CCLM); deriving residual samples for the chroma block based on the prediction samples; updating an intra prediction mode of the chroma block based on an intra prediction mode of a luma block corresponding to the chroma block; determining an LFNST set including a low-frequency inseparable transform (LFNST) matrix based on the updated intra prediction mode; performing LFNST on the chroma block based on the residual samples and the LFNST matrix; and encoding residual information to generate the bitstream; a transmitter configured to transmit the data comprising the bit stream, wherein the updated intra-frame prediction mode is derived as an intra-frame prediction mode corresponding to a specific position in the luminance block, Wherein, based on the prediction mode corresponding to the specific position being the intra block copy (IBC) mode, the intra prediction mode of the chroma block is updated to the intra DC mode, and The specific position is set based on the color format of the chroma block.