Image decoding method, image encoding method, and image data transmission method
By adopting cross-component linear model (CCLM) mode and LFNST index coding in image encoding, the problem of high resolution image/video transmission and storage costs is solved, and image compression efficiency and encoding efficiency are improved.
Patent Information
- Application Number
- CN202510719903.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-29
- Filing Date
- 2020-10-29
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art, when transmitting and storing high resolution, high quality images/videos, increases in the amount of information leads to high transmission and storage costs, and the broadcasting requirements for immersive media and images/videos with different image characteristics are not effectively resolved.
The cross-component linear model (CCLM) mode is adopted, and the intra prediction mode of the chrominance block is updated based on the intra prediction mode of the luminance block, and encoding it through the LFNST index, deriving the LFNST transform set to improve image encoding efficiency.
The image/video compression efficiency is increased, the efficiency of encoding LFNST index and secondary transformation is improved, and the image encoding method and equipment in the CCLM mode are optimized.
Smart Images

Figure CN120302037A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 202080090680.4 (International Application Number: PCT / KR2020 / 014904, Application Date: October 29, 2020, Invention Title: Transform-based Image Coding Method and Apparatus). Technical Field
[0002] The present disclosure relates to image coding technology, and more particularly, to a method and apparatus for coding an image based on a transform in an image coding system. Background Art
[0003] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K, or higher ultra-high definition (UHD) images / videos has been increasing in various fields. As the image / video data becomes higher in resolution and higher in quality, the amount of information or bits to be transmitted increases compared to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0004] In addition, nowadays, the interest in and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and the broadcasting of images / videos having image characteristics different from those of real images such as game images is increasing.
[0005] Therefore, there is a need for an efficient image / video compression technology that can effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention
[0006] Technical Objectives
[0007] One technical aspect of the present disclosure is to provide a method and apparatus for increasing image coding efficiency.
[0008] Another technical aspect of the present disclosure is to provide a method and apparatus for increasing the efficiency of encoding LFNST indices.
[0009] Still another technical aspect of the present disclosure is to provide a method and apparatus for improving the efficiency of a secondary transform by encoding LFNST indices.
[0010] Yet another technical aspect of the present disclosure is to provide an image coding method and an image coding apparatus for deriving an LFNST transform set using an intra mode of a luminance block in a CCLM mode.
[0011] Technical Solutions
[0012] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method may include the following steps: based on the intra prediction mode of a chrominance block being a cross-component linear model (CCLM) mode, updating the intra prediction mode of the chrominance block based on the intra prediction mode of a luminance block corresponding to the chrominance block; determining a set of LFNSTs including an LFNST matrix based on the updated intra prediction mode; and deriving transform coefficients of the chrominance block based on the LFNST matrix derived from the set of LFNSTs, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a specific position in the luminance block, and wherein, based on the intra prediction mode corresponding to the specific position being a MIP mode, updating the updated intra prediction mode to an intra plane mode.
[0013] The specific position is set based on the color format of the chrominance block.
[0014] The specific position is the central position of the luminance block.
[0015] The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)), where xTbY and yTbY represent the upper left coordinates of the luminance block, nTbW and nTbH represent the width and height of the chrominance block, and SubWidthC and SubHeightC represent variables corresponding to the color format.
[0016] When the color format is 4:2:0, SubWidthC and SubHeightC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0017] When the intra prediction mode corresponding to the specific position is an IBC mode, the updated intra prediction mode is an intra DC mode.
[0018] When the intra prediction mode corresponding to the specific position is a palette mode, the updated intra prediction mode is an intra DC mode.
[0019] According to another embodiment of the present disclosure, an image encoding method performed by an encoding device is provided. The method may include the following steps: deriving prediction samples of a chrominance block based on the intra prediction mode of the chrominance block being a cross-component linear model (CCLM); deriving residual samples of the chrominance block based on the prediction samples, the updated intra prediction mode is derived as an intra prediction mode corresponding to a specific position in the luminance block, and based on the intra prediction mode corresponding to the specific position being a MIP mode, updating the updated intra prediction mode to an intra plane mode.
[0020] According to yet another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including a bitstream and encoded image information generated according to an image encoding method performed by an encoding device.
[0021] According to yet another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including encoded image information and a bitstream to cause a decoding device to perform an image decoding method.
[0022] Technical Effects
[0023] According to the present disclosure, the overall image / video compression efficiency can be increased.
[0024] According to the present disclosure, the efficiency of encoding LFNST indexes can be increased.
[0025] According to the present disclosure, the efficiency of secondary transformation can be increased by encoding LFNST indexes.
[0026] According to the present disclosure, an image encoding method and an image encoding device for deriving an LFNST transform set using an intra mode of a luminance block in a CCLM mode can be provided.
[0027] The effects that can be obtained through specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood by those of ordinary skill in the relevant art or derived from the present disclosure. Therefore, the specific effects of the present disclosure are not limited to those explicitly described in the present disclosure and may include various effects that can be understood or derived based on the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a diagram schematically illustrating the configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied.
[0029] Figure 2 is a diagram schematically illustrating the configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.
[0030] Figure 3 schematically illustrates a multi-transformation technique according to an embodiment of the present disclosure.
[0031] Figure 4 schematically shows an intra prediction mode of 65 prediction directions.
[0032] Figure 5 is a diagram illustrating an RST according to an embodiment of the present disclosure.
[0033] Figure 6 is a diagram illustrating the order of arranging output data of a forward primary transformation as a one-dimensional vector according to an example.
[0034] Figure 7 This is a diagram illustrating the order of arranging the output data of the forward quadratic transform into a two-dimensional block according to the example.
[0035] Figure 8 This is a diagram illustrating the block shape to which the LFNST is applied.
[0036] Figure 9 This is a diagram illustrating the arrangement of the output data of the forward LFNST according to the example and shows the block in which the output data of the forward LFNST is arranged according to the example.
[0037] Figure 10 This is a diagram showing that the number of output data of the forward LFNST according to the example is limited to a maximum value of 16.
[0038] Figure 11 This is a diagram illustrating the zeroing in the block to which the 4×4 LFNST is applied according to the example.
[0039] Figure 12 This is a diagram illustrating the zeroing in the block to which the 8×8 LFNST is applied according to the example.
[0040] Figure 13 This is a diagram illustrating the MIP-based predicted sample generation process according to the example.
[0041] Figure 14 This is a diagram illustrating the CCLM applicable when deriving the intra prediction mode of the chrominance block according to the embodiment.
[0042] Figure 15 This is a diagram for explaining the method of decoding an image according to the example.
[0043] Figure 16 This is a diagram for explaining the method of encoding an image according to the example.
[0044] Figure 17 Schematically illustrates an example of a video / image coding system to which the present disclosure can be applied.
[0045] Figure 18 Illustrates the structure of a content stream system to which the present disclosure is applied. Detailed Description of the Invention
[0046] Although the present disclosure may be susceptible to various modifications and include various embodiments, specific embodiments thereof have been shown by way of example in the drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the technical concept of the present disclosure. Unless the context clearly indicates otherwise, the singular forms may include the plural forms. Terms such as "including" and "having" are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus should not be construed as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0047] In addition, for the convenience of describing different characteristic functions from each other, each component in the drawings described herein is independently illustrated. However, it is not meant that each component is implemented by separate hardware or software. For example, any two or more of these components can be combined to form a single component, and any single component can be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent right of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0048] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the drawings. In addition, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.
[0049] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).
[0050] In this document, various embodiments related to video / image coding may be provided, and unless otherwise specified, these embodiments may be combined with each other and executed.
[0051] In this document, video may refer to a collection of a series of images over a period of time. Generally, a picture refers to a unit representing an image of a specific time region, and a slice / tile is a unit that forms a part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / titles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.
[0052] A pixel or picture element (pel) may refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can mean a pixel value in the spatial domain, or when the pixel value is transformed into the frequency domain, it can mean a transform coefficient in the frequency domain.
[0053] A unit can represent the basic unit of image processing. A unit can include at least one of a specific region and information related to that region. A unit can include one luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the situation, terms such as unit, block, region, etc. can be used interchangeably. Generally, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0054] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Additionally, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A, B, C" can mean "at least one of A, B, and / or C".
[0055] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0056] In this disclosure, "at least one of A and B" can mean "only A", "only B", or "both A and B". Additionally, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0057] Furthermore, in this disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0058] Additionally, the parentheses used in this disclosure may denote "for example". Specifically, when indicated as "prediction (intra prediction)", it may mean that "intra prediction" is presented as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra prediction", and "intra prediction" is presented as an example of "prediction". Additionally, when indicated as "prediction (i.e., intra prediction)", this may also mean that "intra prediction" is presented as an example of "prediction".
[0059] The technical features described separately in one of the drawings in this disclosure can be implemented individually or can be implemented simultaneously.
[0060] Figure 1 FIG. is a diagram schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.
[0061] Referring to Figure 1 , the encoding device 100 may include and be configured with an image splitter 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 may include an inter predictor 121 and an intra predictor 122. The residual processor 130 may include a transformer 132, a quantizer 133, a dequantizer 134, an inverse transformer 135. The residual processor 130 may further include a subtractor 131. The adder 150 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the above-described image splitter 110, predictor 120, residual processor 130, entropy encoder 140, adder 150, and filter 160 may be constituted by one or more hardware components (e.g., an encoder chipset or a processor). In addition, the memory 170 may include a decoded picture buffer (DPB) and may also be constituted by a digital storage medium. The hardware components may further include the memory 170 as an internal / external component.
[0062] The image splitter 110 may split an input image (or picture, frame) input to the encoding device 100 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit may be split into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, and later the binary tree structure and / or the ternary tree structure may be applied. Alternatively, the binary tree structure may also be applied first. The encoding process according to this document may be performed based on the final coding unit that is no longer split. In this case, based on the encoding efficiency according to image characteristics, etc., the LCU may be directly used as the final coding unit, or alternatively, the coding unit may be recursively split into coding units of a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction described later. As another example, a processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, each of the PU and TU may be split or divided from the aforementioned final coding unit. The PU may be a unit for sample prediction, and the TU may be a unit for inducing transform coefficients and / or a unit for inducing a residual signal from the transform coefficients.
[0063] In some cases, a unit may be used interchangeably with terms such as a block or a region. Generally, an M×N block may represent samples or a set of transform coefficients composed of M columns and N rows. Samples generally may represent pixels or pixel values, and may also represent only the pixels / pixel values of the luminance component, and may also represent only the pixels / pixel values of the chrominance component. Samples may be used as items corresponding to the pixels or primitives configuring a picture (or image).
[0064] The encoding device 100 can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 121 or the intra-frame predictor 122 from an input image signal (original block, original sample array), and send the generated residual signal to the transformer 132. In this case, as illustrated, the unit in the encoding device 100 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as the subtractor 131. The predictor can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The predictor can generate various information about the prediction, such as prediction mode information, to transmit the generated information to the entropy encoder 140, as described later in the description of each prediction mode. The information about the prediction can be encoded by the entropy encoder 140 and output in the form of a bitstream.
[0065] The intra-frame predictor 122 can predict the current block by referring to the samples within the current picture. Depending on the prediction mode, the samples referred to may be adjacent to the current block or may also be located far from the current block. The prediction modes in intra-frame prediction can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode or the planar mode. Depending on the fineness of the prediction direction, the directional modes can include (for example) 33 directional prediction modes or 65 directional prediction modes. However, this is illustrative, and more or fewer directional prediction modes than the above numbers can be used according to the settings. The intra-frame predictor 122 can also use the prediction mode applied to an adjacent block to determine the prediction mode applied to the current block.
[0066] The inter-frame predictor 121 may derive a prediction block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, the adjacent blocks may include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporally adjacent blocks may be the same as each other and may also be different from each other. The temporally adjacent blocks may be called by names such as, for example, collocated reference blocks, collocated CUs (col CUs), or the like, and the reference picture including the temporally adjacent blocks may also be called by names such as collocated reference blocks, collocated CUs (col CUs), etc., and the reference picture including the temporally adjacent blocks may be called a collocated picture (col Pic). For example, the inter-frame predictor 121 may configure a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. The inter-frame prediction may be performed based on various prediction modes, and for example, in the case of the skip mode and the merge mode, the inter-frame predictor 121 may use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. The motion vector prediction (MVP) mode may indicate the motion vector of the current block by using the motion vector of an adjacent block as the motion vector prediction and by signaling the motion vector difference.
[0067] The predictor 120 may generate a prediction signal based on various prediction methods to be described later. For example, the predictor may apply not only intra-frame prediction or inter-frame prediction to predict a block, but also apply intra-frame prediction and inter-frame prediction simultaneously. This may be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor may perform prediction on a block based on the intra-block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode may be used for content image / video coding such as games, for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but it may perform similarly for inter-frame prediction because it derives a reference block in the current picture. That is, IBC may use at least one of the inter-frame prediction techniques described in this document. The palette mode may be regarded as an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, the sample values in the picture may be signaled based on the information about the palette index and the palette table.
[0068] The prediction signal generated by a predictor (including the inter-frame predictor 121 and / or the intra-frame predictor 122) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 132 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen–Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, GBT represents a transform obtained from the graph. CNT represents a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. Additionally, the transform process can also be applied to a block of pixels in a square having the same size, and can also be applied to a block having a variable size rather than a square.
[0069] Quantizer 133 may quantize the transform coefficients to send the quantized transform coefficients to entropy encoder 140, and entropy encoder 140 may encode the quantized signal (information about the quantized transform coefficients) to output the encoded quantized signal in the form of a bitstream. The information about the quantized transform coefficients may be referred to as residual information. Quantizer 133 may rearrange the quantized transform coefficients in block form in the form of a one-dimensional vector based on the coefficient scan order, and may also generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of a one-dimensional vector. Entropy encoder 140 may perform various coding methods, for example, such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 140 may also encode information required for video / image reconstruction (e.g., values of syntax elements, etc.) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be sent or stored in the form of a bitstream in units of network abstraction layer (NAL). The video / image information may also include information about various parameter sets such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). In addition, the video / image information may also include general constraint information. In this document, the syntax elements and / or information transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the above encoding process and included in the bitstream. The bitstream may be sent through a network or may be stored in a digital storage medium. In this article, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for sending the signal output from entropy encoder 140 and / or a storage device (not shown) for storing the signal may be configured as internal / external elements of encoding device 100, or the transmitter may also be included in entropy encoder 140.
[0070] The quantized transform coefficients output from the quantizer 133 can be used to generate a prediction signal. For example, the dequantizer 134 and the inverse transformer 135 can apply dequantization and inverse transformation to the quantized transform coefficients in order to recover the residual signal (residual block or residual samples). The adder 150 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 121 or the intra-frame predictor 122 in order to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). As in the case of applying the skip mode, if there is no residual in the block to be processed, the predicted block can be used as the reconstructed block. The adder 150 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current picture and, as described below, can also be used for inter-frame prediction of the next picture through filtering.
[0071] In addition, luminance mapping and chrominance scaling (LMCS) can also be applied during picture encoding and / or reconstruction.
[0072] The filter 160 can apply filtering to the reconstructed signal, thereby improving the subjective / objective image quality. For example, the filter 160 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 160 can generate various filtering-related information to transmit the generated information to the entropy encoder 140, as described later in the description of each filtering method. The filtering-related information can be encoded by the entropy encoder 140 and can be output in the form of a bitstream.
[0073] The modified reconstructed picture sent to the memory 170 can be used as a reference picture in the inter-frame predictor 121. If inter-frame prediction is applied by the inter-frame predictor, the encoding device can avoid prediction mismatch between the encoding device 100 and the decoding device and can also improve the encoding efficiency.
[0074] The DPB of the memory 170 can store the modified reconstructed picture to be used as a reference picture in the inter-frame predictor 121. The memory 170 can store the motion information of the block in which the motion information within the current picture is derived (or encoded) and / or the motion information of the block in the previously reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 121 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 170 can store the reconstructed samples of the reconstructed blocks within the current picture and transmit the reconstructed samples to the intra-frame predictor 122.
[0075] Figure 2It is a diagram schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.
[0076] Referring to Figure 2 , the decoding device 200 may include and be configured with an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The predictor 230 may include an inter-predictor 232 and an intra-predictor 231. The residual processor 220 may include a dequantizer 221 and an inverse transformator 222. According to an embodiment, the entropy decoder 210, the residual processor 220, the predictor 230, the adder 240, and the filter 250 described above may be configured by one or more hardware components (e.g., a decoder chipset or a processor). In addition, the memory 260 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 260 as an internal / external component.
[0077] When receiving a bitstream including video / image information, the decoding device 200 may reconstruct an image in response to the process of processing the video / image information in the Figure 1 encoding device shown. For example, the decoding device 200 may derive units / blocks based on block segmentation-related information obtained from the bitstream. The decoding device 200 may perform decoding using the processing units applied to the encoding device. Therefore, the processing units for decoding may be, for example, encoding units, and the encoding units may be segmented from a CTU or an LCU according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the encoding units. In addition, the reconstructed image signal decoded and output by the decoding device 200 may be reproduced by a reproduction device.
[0078] The decoding device 200 may receive, in the form of a bitstream, from Figure 1The signal output by the encoding device as shown, and the received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction) by parsing the bitstream. The video / image information can also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information can also include general constraint information. The decoding device can also decode the picture based on the information about the parameter set and / or the general constraint information. The information and / or syntax elements to be signaled / received (which will be described later in this document) can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 210 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements required for image reconstruction, as well as the quantization values of the residual-related transform coefficients. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the syntax element information to be decoded and the decoding information of the adjacent blocks and the block to be decoded or the information of the symbols / bins decoded in the previous step to determine the context model, and perform arithmetic decoding on the bins by predicting the bin generation probability according to the determined context model to generate the symbols corresponding to the values of each syntax element. At this time, the CABAC entropy decoding method can determine the context model and then update the context model using the information of the symbols / bins used for decoding the context model of the next symbol / bin. The information about prediction among the information decoded by the entropy decoder 210 can be provided to the predictors (inter-frame predictor 232 and intra-frame predictor 231), and the residual values obtained by the entropy decoder 210 performing entropy decoding, that is, the quantization transform coefficients and the related parameter information, can be input to the residual processor 220. The residual processor 220 can derive the residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded by the entropy decoder 210 can be provided to the filter 250. In addition, the receiver (not shown) for receiving the signal output by the encoding device can also be configured as an internal / external component of the decoding device 200, or the receiver can also be a component of the entropy decoder 210. In addition, the decoding device according to the present disclosure can be referred to as a video / image / picture decoding device, and the decoding device can also be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210, and the sample decoder can include at least one of a dequantizer 221, an inverse transformer 222, an adder 240, a filter 250, a memory 260, an inter-frame predictor 232, and an intra-frame predictor 231.
[0079] The dequantizer 221 may dequantize the quantized transform coefficients to output the transform coefficients. The dequantizer 221 may rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order executed by the encoding device. The dequantizer 221 may perform dequantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step information) and obtain the transform coefficients.
[0080] The inverse transformer 222 inverse-transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0081] The predictor 230 may perform prediction on a current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block and determine a specific intra / inter prediction mode based on the information about prediction output from the entropy decoder 210.
[0082] The predictor may generate a prediction signal based on various prediction methods to be described later. For example, the predictor may apply not only intra prediction or inter prediction to the prediction of a block but also apply intra prediction and inter prediction simultaneously. This may be referred to as combined inter and intra prediction (CIIP). In addition, the predictor may perform prediction on a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode may be used for content image / video coding such as games, e.g., screen content coding (SCC). IBC basically performs prediction in the current picture, but it may be similarly performed for inter prediction since it derives a reference block in the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information about a palette table and a palette index may be signaled through video / image information.
[0083] The intra predictor 231 may predict the current block by referring to samples within the current picture. Depending on the prediction mode, the samples referred to may be adjacent to the current block or may also be located away from the current block. The prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 231 may also use the prediction mode applied to an adjacent block to determine the prediction mode applied to the current block.
[0084] The inter - frame predictor 232 can induce a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, adjacent blocks can include spatially adjacent blocks existing within the current picture and temporally adjacent blocks existing in the reference picture. For example, the inter - frame predictor 232 can configure a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter - frame prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of the inter - frame prediction of the current block.
[0085] The adder 240 can add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter - frame predictor 232 and / or the intra - frame predictor 231) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). As in the case of applying the skip mode, if there is no residual for the block to be processed, the prediction block can be used as the reconstructed block.
[0086] The adder 240 can be referred to as a reconstructor or a reconstructed - block generator. The generated reconstructed signal can be used for the intra - frame prediction of the next block to be processed within the current picture, and as described below, can also be output through filtering or can be used for the inter - frame prediction of the next picture.
[0087] In addition, luminance mapping and chroma scaling (LMCS) can also be applied during the picture decoding process.
[0088] The filter 250 can apply filtering to the reconstructed signal, thereby improving the subjective / objective image quality. For example, the filter 250 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and send the modified reconstructed picture to the memory 260, specifically, the DPB of the memory 260. The various filtering methods can include (for example) de - blocking filtering, sample - adaptive offset, adaptive loop filter, bilateral filter, etc.
[0089] The (modified) reconstructed picture stored in the DPB of the memory 260 can be used as a reference picture in the inter - predictor 232. The memory 260 can store the motion information of the blocks in which the motion information within the current picture is derived (decoded) and / or the motion information of the blocks within the previous reconstructed pictures. The stored motion information can be transmitted to the inter - predictor 260 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 260 can store the reconstructed samples of the reconstructed blocks within the current picture and transmit the stored reconstructed samples to the intra - predictor 231.
[0090] In this document, the exemplary embodiments described in the filter 160, the inter - predictor 121, and the intra - predictor 122 of the encoding device 100 can be equally applied to the filter 250, the inter - predictor 232, and the intra - predictor 231 of the decoding device 200, respectively.
[0091] As described above, prediction is performed to improve the compression efficiency when performing video coding. Accordingly, a prediction block including prediction samples for the current block that is the target block to be encoded can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve the image coding efficiency by signaling to the decoding device not the original sample values of the original block itself but the information about the residual between the original block and the prediction block (residual information). The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.
[0092] Residual information can be generated through a transformation process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients so that it can signal the associated residual information to the decoding device (through the bitstream). Here, the residual information can include value information, position information, transformation technique, transformation kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / de - quantization process based on the residual information and derive residual samples (or a residual sample block). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can de - quantize / inverse - transform the quantized transform coefficients to derive a residual block for use as a reference for inter - prediction of the next picture and can generate a reconstructed picture based on this.
[0093] Figure 3 A multi - transformation technique according to an embodiment of the present disclosure is schematically illustrated.
[0094] Refer toFigure 3 , the transformer may correspond to the transformer in the encoding device described above Figure 1 , and the inverse transformer may correspond to the inverse transformer in the encoding device described above Figure 1 , or the inverse transformer in the decoding device described above Figure 2 .
[0095] The transformer may derive (primary) transform coefficients (S310) by performing a single transformation based on residual samples (residual sample arrays) in a residual block. This single transformation may be referred to as the core transformation. In this document, the single transformation may be based on multiple transform selection (MTS), and when multiple transform is used as the single transformation, it may be referred to as multi-core transformation.
[0096] The multi-core transformation may represent a method of additionally using Discrete Cosine Transform (DCT) type 2 and Discrete Sine Transform (DST) type 7, DCT type 8, and / or DST type 1 for transformation. That is, the multi-core transformation may represent a transformation method of transforming a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this document, from the perspective of the transformer, the primary transform coefficients may be referred to as temporary transform coefficients.
[0097] In other words, when applying a conventional transformation method, transform coefficients may be generated by applying a transformation from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, when applying the multi-core transformation, transform coefficients (or primary transform coefficients) may be generated by applying a transformation from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this document, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transformation types, transform kernels, or transform cores. These DCT / DST transformation types may be defined based on basis functions.
[0098] When performing the multi-core transformation, a vertical transform kernel and a horizontal transform kernel for a target block may be selected from among the transform kernels, a vertical transform may be performed on the target block based on the vertical transform kernel, and a horizontal transform may be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform may indicate a transformation of the horizontal component of the target block, and the vertical transform may indicate a transformation of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel may be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block) including the residual block.
[0099] In addition, according to the example, if a transformation is performed by applying MTS, the mapping relationship of the transformation kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transformation or the horizontal transformation. For example, when the horizontal transformation kernel is represented as trTypeHor and the vertical transformation kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT8.
[0100] In this case, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transformation kernel sets. For example, MTS index 0 can indicate that both the trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both the trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that the trTypeHor value is 2 and the trTypeVer value is 1, MTS index 3 can indicate that the trTypeHor value is 1 and the trTypeVer value is 2, and MTS index 4 can indicate that both the trTypeHor and trTypeVer values are 2.
[0101] In one example, the transformation kernel sets according to the MTS index information are shown in the following table.
[0102] [Table 1]
[0103] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0104] The transformer may perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients (S320). The primary transform is a transform from the spatial domain to the frequency domain, and the secondary transform refers to using the correlation existing between the (primary) transform coefficients to transform into a more compact representation. The secondary transform may include an inseparable transform. In this case, the secondary transform may be referred to as an inseparable secondary transform (NSST) or a mode-dependent inseparable secondary transform (MDNSST). NSST may represent a transform that performs a secondary transform on the (primary) transform coefficients derived by the primary transform based on an inseparable transform matrix to generate modified transform coefficients (or secondary transform coefficients) for the residual signal. Here, based on the inseparable transform matrix, the transform can be applied to the (primary) transform coefficients once without separating the vertical transform and the horizontal transform (or applying the horizontal / vertical transform independently). In other words, NSST is not applied separately to the (primary) transform coefficients in the vertical and horizontal directions, and may represent, for example, a transform method of rearranging a two-dimensional signal (transform coefficients) into a one-dimensional signal in a specific predetermined direction (e.g., row-major order or column-major order) and then generating modified transform coefficients (or secondary transform coefficients) based on the inseparable transform matrix. For example, the row-major order is arranged in rows in the order of the first row, the second row, …, and the Nth row for M×N blocks, and the column-major order is arranged in columns in the order of the first column, the second column, …, and the Mth column for M×N blocks. NSST may be applied to the upper left region of a block configured with (primary) transform coefficients (hereinafter referred to as a transform coefficient block). For example, when both the width W and the height H of the transform coefficient block are 8 or greater, 8×8 NSST may be applied to the upper left 8×8 region of the transform coefficient block. In addition, while both the width (W) and the height (H) of the transform coefficient block are 4 or greater, when the width (W) or the height (H) of the transform coefficient block is less than 8, 4×4 NSST may be applied to the upper left min(8, W)×min(8, H) region of the transform coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that the width W or the height H of the transform coefficient block is 4 or greater is satisfied, 4×4 NSST may be applied to the upper left end min(8, W)×min(8, H) region of the transform coefficient block.
[0105] Specifically, for example, if a 4×4 input block is used, the inseparable secondary transform may be performed as follows.
[0106] The 4×4 input block X may be represented as follows.
[0107] [Equation 1]
[0108]
[0109] If X is represented in the form of a vector, then the vector It can be expressed as follows.
[0110] [Equation 2]
[0111]
[0112] In Equation 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to row-major order.
[0113] In this case, the inseparable quadratic transformation can be calculated as follows.
[0114] [Equation 3]
[0115]
[0116] In this equation, represents the transform coefficient vector, and T represents a 16×16 (inseparable) transform matrix.
[0117] Through the aforementioned Equation 3, a 16×1 transform coefficient vector can be derived, and the vector can be reorganized into 4×4 blocks in a scan order (such as horizontal, vertical, and diagonal, etc.). However, the above calculation is an example, and the hypercube-Givens transform (HyGT), etc. can also be used for the calculation of the inseparable quadratic transformation in order to reduce the computational complexity of the inseparable quadratic transformation.
[0118] In addition, in the inseparable quadratic transformation, the transform kernel (or transform core, transform type) can be selected to be mode-dependent. In this case, the mode can include an intra prediction mode and / or an inter prediction mode.
[0119] As described above, the inseparable quadratic transformation can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 region included in the transform coefficient block when both W and H are equal to or greater than 8, and the 8×8 region can be the upper left 8×8 region in the transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 region included in the transform coefficient block when both W and H are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0120] Here, in order to select the transform kernels related to the mode, two non-separable quadratic transform kernels for each transform set for non-separable quadratic transforms can be configured for both the 8×8 transform and the 4×4 transform, and there can be four transform sets. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.
[0121] However, as the size of the transform (i.e., the size of the region to which the transform is applied) can be a size other than 8×8 or 4×4, for example, the number of sets can be n, and the number of transform kernels in each set can be k.
[0122] The transform sets can be referred to as NSST sets or LFNST sets. A specific set among the transform sets can be selected, for example, based on the intra prediction mode of the current block (CU or sub-block). The low-frequency non-separable transform (LFNST) can be an example of a reduced non-separable transform, which will be described later, and represents a non-separable transform for low-frequency components.
[0123] As a reference, for example, the intra prediction mode can include two non-directional (or non-angle) intra prediction modes and 65 directional (or angle) intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include 65 intra prediction modes numbered from 2 to 66. However, this is an example, and this document can be applied even if the number of intra prediction modes is different. In addition, in some cases, the intra prediction mode numbered 67 can also be used, and the intra prediction mode numbered 67 can represent a linear model (LM) mode.
[0124] Figure 4 The intra-directional mode with 65 prediction directions is schematically shown.
[0125] Referring to Figure 4 , based on the intra prediction mode 34 with the upper left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode with horizontal directivity and an intra prediction mode with vertical directivity. In Figure 4In this case, H and V respectively indicate horizontal directionality and vertical directionality, and the numbers -32 to 32 indicate displacements of 1 / 32 units at the sample grid positions. These numbers can represent offsets for the mode index values. Intra prediction modes 2 to 33 have horizontal directionality, and intra prediction modes 34 to 66 have vertical directionality. Strictly speaking, intra prediction mode 34 can be regarded as neither horizontal nor vertical, but can be classified as belonging to horizontal directionality when determining the transform set of the secondary transform. This is because the input data is transposed for the vertically oriented mode symmetric to intra prediction mode 34, and the input data alignment method for the horizontal mode is used for intra prediction mode 34. Transposing the input data means switching the rows and columns of the two-dimensional M×N block data into N×M data. Intra prediction mode 18 and intra prediction mode 50 can respectively represent a horizontal intra prediction mode and a vertical intra prediction mode, and intra prediction mode 2 can be called a top-right diagonal intra prediction mode because intra prediction mode 2 has a left reference pixel and performs prediction in the top-right direction. Similarly, intra prediction mode 34 can be called a bottom-right diagonal intra prediction mode, and intra prediction mode 66 can be called a bottom-left diagonal intra prediction mode.
[0126] According to an example, four transform sets can be mapped according to the intra prediction mode, for example, as shown in the following table.
[0127] [Table 2]
[0128] lfnstPredModeIntra lfnstTrSetIdx lfnstPredModeIntra<0 1 0 <= lfnstPredModeIntra <= 1 0 2 <= lfnstPredModeIntra <= 12 1 13 <= lfnstPredModeIntra <= 23 2 24 <= lfnstPredModeIntra <= 44 3 45 <= lfnstPredModeIntra <= 55 2 56 <= lfnstPredModeIntra <= 80 1 81 <= lfnstPredModeIntra <= 83 0
[0129] As shown in Table 2, any one of the four transform sets, that is, lfnstTrSetIdx, can be mapped to any one of the four indices (i.e., 0 to 3) according to the intra prediction mode.
[0130] When determining a specific set for the non-separable transform, one of the k transform kernels in the specific set can be selected by the non-separable secondary transform index. The encoding device can derive the non-separable secondary transform index indicating the specific transform kernel based on rate distortion (RD) checking, and can signal the non-separable secondary transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable secondary transform index. For example, the lfnst index value 0 can refer to the first non-separable secondary transform kernel, the lfnst index value 1 can refer to the second non-separable secondary transform kernel, and the lfnst index value 2 can refer to the third non-separable secondary transform kernel. Alternatively, the lfnst index value 0 can indicate that the first non-separable secondary transform is not applied to the target block, and the lfnst index values 1 to 3 can indicate three transform kernels.
[0131] The transformer may perform an inseparable quadratic transform based on the selected transform kernel and may obtain modified (quadratic) transform coefficients. As described above, the modified transform coefficients may be derived as the transform coefficients quantized by the quantizer, may be encoded and signaled to the decoding device, and be transmitted to the dequantizer / inverse transformer in the encoding device.
[0132] In addition, as described above, if the quadratic transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform may be derived as the transform coefficients quantized by the quantizer as described above, may be encoded and signaled to the decoding device, and be transmitted to the dequantizer / inverse transformer in the encoding device.
[0133] The inverse transformer may perform a series of processes in an order opposite to the order that has been performed in the above transformer. The inverse transformer may receive the (dequantized) transform coefficients, and derive the (primary) transform coefficients by performing a quadratic (inverse) transform (S350), and may obtain a residual block (residual samples) by performing a primary (inverse) transform on the (primary) transform coefficients (S360). In this regard, from the perspective of the inverse transformer, the primary transform coefficients may be referred to as modified transform coefficients. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.
[0134] The decoding device may further include a quadratic inverse transform application determiner (or an element for determining whether to apply the quadratic inverse transform) and a quadratic inverse transform determiner (or an element for determining the quadratic inverse transform). The quadratic inverse transform application determiner may determine whether to apply the quadratic inverse transform. For example, the quadratic inverse transform may be NSST, RST, or LFNST, and the quadratic inverse transform application determiner may determine whether to apply the quadratic inverse transform based on the quadratic transform flag obtained by parsing the bitstream. In another example, the quadratic inverse transform application determiner may determine whether to apply the quadratic inverse transform based on the transform coefficients of the residual block.
[0135] The quadratic inverse transform determiner may determine the quadratic inverse transform. In this case, the quadratic inverse transform determiner may determine the quadratic inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra prediction mode. In an embodiment, the quadratic transform determination method may be determined depending on the primary transform determination method. Various combinations of the primary transform and the quadratic transform may be determined according to the intra prediction mode. In addition, in an example, the quadratic inverse transform determiner may determine the region to which the quadratic inverse transform is applied based on the size of the current block.
[0136] In addition, as described above, if the second (inverse) transformation is omitted, the (dequantized) transform coefficients can be received, a single (separable) inverse transformation can be performed, and a residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0137] In addition, in the present disclosure, a reduced second transformation (RST) in which the size of the transformation matrix (kernel) is reduced can be applied in the concept of NSST, so as to reduce the amount of computation and storage required for the non-separable second transformation.
[0138] In addition, the transform kernel, the transform matrix, and the coefficients constituting the transform kernel matrix described in the present disclosure, that is, the kernel coefficients or the matrix coefficients, can be represented by 8 bits. This can be a condition implemented in the decoding device and the encoding device, and compared with the existing 9 bits or 10 bits, it can reduce the storage amount required for storing the transform kernel, and can reasonably adapt to the performance degradation. In addition, representing the kernel matrix by 8 bits can allow the use of a small multiplier, and can be more suitable for the single instruction multiple data (SIMD) instruction for optimal software implementation.
[0139] In this specification, the term "RST" may refer to a transformation performed on the residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. In the case of performing the reduction transformation, due to the reduction in the size of the transform matrix, the amount of computation required for the transformation can be reduced. That is, RST can be used to solve the computational complexity problem that occurs in the transformation of a large-sized block or a non-separable transformation.
[0140] RST can be referred to by various terms such as reduction transformation, reduced second transformation, shrinking transformation, simplified transformation, and simple transformation, and the names that RST can be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in the low-frequency region including non-zero coefficients in the transform block, it can be referred to as a low-frequency non-separable transformation (LFNST). The transform index can be referred to as the LFNST index.
[0141] In addition, when performing the second inverse transformation based on RST, the inverse transformers 135 of the encoding device 100 and the inverse transformers 222 of the decoding device 200 can include: an inverse reduced second transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients; and an inverse single transformer that derives the residual samples of the target block based on the inverse single transformation of the modified transform coefficients. The inverse single transformation refers to the inverse transformation of the single transformation applied to the residual. In the present disclosure, deriving the transform coefficients based on the transformation may refer to deriving the transform coefficients by applying the transformation.
[0142] Figure 5FIG. is an illustration of an RST according to an embodiment of the present disclosure.
[0143] In the present disclosure, a "target block" may refer to a current block to be encoded, a residual block, or a transform block.
[0144] In the RST according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, so that a reduction transform matrix can be determined, where R is less than N. N may refer to the square of the length of a side of a block to which a transform is applied, or the total number of transform coefficients corresponding to a block to which a transform is applied, and the reduction factor may refer to the R / N value. The reduction factor may be referred to as a reduction factor, a shrinking factor, a simplification factor, a simple factor, or various other terms. In addition, R may be referred to as a reduction coefficient, but depending on the situation, the reduction factor may refer to R. In addition, depending on the situation, the reduction factor may refer to the N / R value.
[0145] In the example, the reduction factor or reduction coefficient may be signaled through a bitstream, but the example is not limited thereto. For example, a predetermined value for the reduction factor or reduction coefficient may be stored in each of the encoding device 100 and the decoding device 200, and in this case, the reduction factor or reduction coefficient may not be signaled separately.
[0146] The size of the reduction transform matrix according to the example may be R×N, which is less than N×N (the size of a conventional transform matrix), and may be defined as in Equation 4 below.
[0147] [Equation 4]
[0148]
[0149] Figure 5 The matrix T in the reduction transform block shown in (a) of may refer to the matrix T of Equation 4 R×N . As Figure 5 shown in (a) of, when the reduction transform matrix T R×N is multiplied by the residual samples of the target block, the transform coefficients of the current block can be derived.
[0150] In the example, if the size of the block to which a transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST according to Figure 5 (a) of can be expressed as the matrix operation shown in Equation 5 below. In this case, the storage and multiplication calculations can be reduced to approximately 1 / 4 by the reduction factor.
[0151] In the present disclosure, matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix set on the left side of the column vector.
[0152] [Equation 5]
[0153]
[0154] In Equation 5, r1 to r 64 can represent the residual samples of the target block, and specifically can be the transform coefficients generated by applying a single transform. As a result of the calculation of Equation 5, the transform coefficients c of the target block can be derived i , and the process of deriving c i can be as shown in Equation 6.
[0155] [Equation 6]
[0156]
[0157] As a result of the calculation of Equation 6, the transform coefficients c1 to c of the target block can be derived R . That is, when R = 16, the transform coefficients c1 to c of the target block can be derived 16 . If a conventional transform is applied instead of RST, and a transform matrix of size 64×64 (N×N) is multiplied by a residual sample of size 64×1 (N×1), then only 16 (R) transform coefficients are derived for the target block because RST is applied, although 64 (N) transform coefficients are derived for the target block. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data sent from the encoding device 100 to the decoding device 200 is reduced, and thus the transmission efficiency between the encoding device 100 and the decoding device 200 can be improved.
[0158] When considered from the perspective of the size of the transform matrix, the size of the conventional transform matrix is 64×64 (N×N), but the size of the reduced transform matrix is reduced to 16×64 (R×N). Therefore, compared with the case of performing a conventional transform, the storage utilization rate in the case of performing RST can be reduced by the ratio of R / N. In addition, when compared with the number of multiplication calculations N×N in the case of using a conventional transform matrix, using the reduced transform matrix can reduce the number of multiplication calculations (R×N) by the ratio of R / N.
[0159] In the example, the transformer 132 of the encoding device 100 can derive the transform coefficients of the target block by performing a single transform on the residual samples of the target block and a second transform based on RST. These transform coefficients can be transmitted to the inverse transformer of the decoding device 200, and the inverse transformer 222 of the decoding device 200 can derive the modified transform coefficients based on the inverse reduced second transform (RST) for the transform coefficients, and can derive the residual samples of the target block based on the inverse single transform for the modified transform coefficients.
[0160] Inverse RST matrix T according to the example N×Ris sized N×R, which is smaller than the size of a conventional inverse transform matrix N×N, and is related to the reduced transform matrix T shown in Equation 4 R×N has a transpose relationship.
[0161] Figure 5 The matrix T in the reduced inverse transform block shown in (b) of t can refer to the inverse RST matrix T N×R T (the superscript T refers to transpose). As Figure 5 shown in (b) of N×R T when the inverse RST matrix T R×N T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. The inverse RST matrix T R×N can be expressed as (T T N×R ).
[0162] More specifically, when the inverse RST is used as a secondary inverse transform, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the residual samples of the target block can be derived.
[0163] In an example, if the size of the block to which the inverse transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST in (b) of Figure 5 can be represented as the matrix operation shown in Equation 7 below.
[0164] [Equation 7]
[0165]
[0166] In Equation 7, c1 to c 16 can represent the transform coefficients of the target block. As a result of the calculation in Equation 7, r j representing the modified transform coefficients of the target block or the residual samples of the target block can be derived, and the process of deriving r j can be as shown in Equation 8.
[0167] [Equation 8]
[0168]
[0169] As a result of the calculation of Equation 8, r1 to r representing the modified transform coefficients of the target block or the residual samples of the target block can be derived. N From the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduction transform matrix is reduced to 64×16 (R×N). Therefore, compared with the case of performing a conventional inverse transform, the storage utilization rate in the case of performing an inverse RST can be reduced by the ratio of R / N. In addition, when compared with the number of multiplication calculations N×N in the case of using a conventional inverse transform matrix, using the inverse reduction transform matrix can reduce the number of multiplication calculations (N×R) by the ratio of R / N.
[0170] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied according to the transform set in Table 2. Since, according to the intra prediction mode, one transform set includes two or three transforms (kernels), it can be configured to select one of up to four transforms including the case where no secondary transform is applied. Among the transforms without applying the secondary transform, applying an identity matrix can be considered. Assuming that indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case of applying the identity matrix, that is, the case where no secondary transform is applied), the transform index or lfnst index as a syntax element can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the upper left 8×8 block, through the transform index, 8×8 NSST in the RST configuration can be specified, or 8×8 lfnst can be specified when LFNST is applied. 8×8 lfnst and 8×8 RST refer to the transforms that can be applied to the 8×8 region included in the transform coefficient block when both the W and H of the target block to be transformed are equal to or greater than 8, and the 8×8 region can be the upper left 8×8 region in the transform coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to the transforms that can be applied to the 4×4 region included in the transform coefficient block when both the W and H of the target block are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transform coefficient block.
[0171] According to an embodiment of the present disclosure, for the transform in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 region. Here, "maximum" means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is, when performing RST by applying an m×48 transform kernel matrix (m≤16) to an 8×8 region, 48 pieces of data are input, and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming the 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on 48 pieces of data constituting a region other than the lower-right 4×4 region among the 8×8 region. Here, when performing matrix operations by applying a maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper-left 4×4 region according to the scanning order, and the upper-right 4×4 region and the lower-left 4×4 region can be filled with zeros.
[0172] For the inverse transform in the decoding process, the transpose matrix of the aforementioned transform kernel matrix can be used. That is, when performing inverse RST or LFNST in the inverse transform process executed by the decoding device, the input coefficient data for applying inverse RST is configured in a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector can be arranged into a two-dimensional block according to a predetermined arrangement order.
[0173] In summary, in the transform process, when RST or LFNST is applied to an 8×8 region, matrix operations are performed on 48 transform coefficients in the upper-left region, upper-right region, and lower-left region of the 8×8 region except for the lower-right region with a 16×48 transform kernel matrix. For the matrix operations, 48 transform coefficients are input in a one-dimensional array. When performing matrix operations, 16 modified transform coefficients are derived, and the modified transform coefficients can be arranged in the upper-left region of the 8×8 region.
[0174] Conversely, in the inverse transformation process, when applying the inverse RST or LFNST to an 8×8 region, the corresponding 16 transform coefficients in the 8×8 region corresponding to the upper left region of the 8×8 region among the transform coefficients in the 8×8 region can be input in a one-dimensional array according to the scanning order, and can undergo matrix operations with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix) * (16×1 transform coefficient vector) = (48×1 modified transform coefficient vector). Here, an n×1 vector can be interpreted as having the same meaning as an n×1 matrix, and thus can be represented as an n×1 column vector. In addition, * represents matrix multiplication. When performing the matrix operation, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left region, upper right region, and lower left region of the 8×8 region except for the lower right region.
[0175] When the second inverse transformation is based on RST, the inverse transformer 135 of the encoding device 100 and the inverse transformer 222 of the decoding device 200 can include an inverse reduced second transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients and an inverse first transformer for deriving the residual samples of the target block based on the inverse first transformation of the modified transform coefficients. The inverse first transformation refers to the inverse transformation of the first transformation applied to the residual. In the present disclosure, deriving transform coefficients based on a transformation can refer to deriving transform coefficients by applying the transformation.
[0176] The non-separable transformation (LFNST) described above will be described in detail below. The LFNST can include a forward transformation performed by the encoding device and an inverse transformation performed by the decoding device.
[0177] The encoding device receives the result (or a part of the result) derived after applying the first (core) transformation as input and applies a forward second transformation (second transformation).
[0178] [Equation 9]
[0179] y = G T x
[0180] In Equation 9, x and y are the input and output of the second transformation respectively, G is a matrix representing the second transformation, and the transform basis vectors are composed of column vectors. In the case of the inverse LFNST, when the dimension of the transformation matrix G is expressed as [number of rows × number of columns], in the case of the forward LFNST, the transpose of the matrix G becomes the dimension of G T of.
[0181] For the inverse LFNST, the dimensions of matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of 8 transformed basis vectors sampled from the left sides of the [48×16] matrix and the [16×16] matrix, respectively.
[0182] On the other hand, for the forward LFNST, matrix G T has dimensions of [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transformed basis vectors from the upper parts of the [16×48] matrix and the [16×16] matrix, respectively.
[0183] Therefore, in the case of the forward LFNST, a [48×1] vector or a [16×1] vector can be used as the input x, and a [16×1] vector or an [8×1] vector can be used as the output y. In video encoding and decoding, the output of the forward single transform is two-dimensional (2D) data. Therefore, in order to construct a [48×1] vector or a [16×1] vector as the input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data that is the output of the forward transform.
[0184] Figure 6 is a diagram illustrating the order of arranging the output data of the forward single transform into a one-dimensional vector according to an example. Figure 6 The left diagrams of (a) and (b) of illustrate the order for constructing a [48×1] vector, and Figure 6 the right diagrams of (a) and (b) of illustrate the order for constructing a [16×1] vector. In the case of LFNST, a one-dimensional vector x can be obtained by arranging the 2D data in the same order as in Figure 6 the (a) and (b) of.
[0185] The arrangement direction of the output data of the forward single transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in the horizontal direction with respect to the diagonal direction, the output data of the forward single transform can be arranged in the order of Figure 6 the (a) of, and when the intra prediction mode of the current block is in the vertical direction with respect to the diagonal direction, the output data of the forward single transform can be arranged in the order of Figure 6 the (b) of.
[0186] According to the example, an arrangement order different from Figure 6 the (a) and (b) of can be applied, and in order to derive an arrangement order different from that applied in Figure 6For the same result (y vector) in the arrangement orders of (a) and (b), the column vectors of matrix G can be rearranged according to the arrangement order. That is, the column vectors of G can be rearranged such that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0187] Since the output y derived by Equation 9 is a one-dimensional vector, when two-dimensional data is required as input data during the process of using the result of the forward quadratic transform as input (for example, during quantization or residual coding), the output y vector of Equation 9 needs to be appropriately rearranged into 2D data again.
[0188] Figure 7 FIG. is a diagram illustrating the order of arranging the output data of the forward quadratic transform into two-dimensional blocks according to an example.
[0189] In the case of LFNST, the output values can be arranged in a 2D block according to a predetermined scan order. Figure 7 FIG. (a) shows that when the output y is a [16×1] vector, the output values are arranged at 16 positions in a 2D block according to the diagonal scan order. Figure 7 FIG. (b) shows that when the output y is an [8×1] vector, the output values are arranged at 8 positions in a 2D block according to the diagonal scan order, and the remaining 8 positions are filled with zeros. Figure 7 X in FIG. (b) indicates that it is filled with zeros.
[0190] According to another example, since the order of processing the output vector y can be preset during quantization or residual coding, the output vector y may not be arranged in a 2D block as shown in Figure 7 . However, in the case of residual coding, data coding can be performed in units of 2D blocks (for example, 4×4) (for example, CG (coefficient group)), and in this case, the data is arranged according to a specific order in the diagonal scan order as shown in Figure 7 .
[0191] In addition, the decoding device can configure the one-dimensional input vector y by arranging the two-dimensional data output through the dequantization process according to a preset scan order for inverse transform. The input vector y can be output as the output vector x through the following formula.
[0192] [Equation 10]
[0193] x = Gy
[0194] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.
[0195] The output vector x is arranged in a two-dimensional block according to Figure 6 the order shown in Figure 6 , and is arranged as two-dimensional data, and this two-dimensional data becomes the input data (or a part of the input data) of the inverse first transformation.
[0196] Therefore, the inverse second transformation is overall the reverse of the forward second transformation process, and in the case of the inverse transformation, different from the forward direction, the inverse second transformation is first applied, and then the inverse first transformation is applied.
[0197] In the inverse LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be selected as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.
[0198] In addition, eight matrices can be derived from the four transformation sets shown in Table 2 above, and each transformation set can be composed of two matrices. Which transformation set to use among the four transformation sets is determined according to the intra-frame prediction mode, and more specifically, based on the value of the intra-frame prediction mode extended by considering wide-angle intra-frame prediction (WAIP). Which matrix to select from the two matrices constituting the selected transformation set is derived by index signaling. More specifically, 0, 1, and 2 can be used as the transmitted index values, 0 can indicate that LFNST is not applied, and 1 and 2 can indicate any one of the two transformation matrices constituting the transformation set selected based on the intra-frame prediction mode value.
[0199] Furthermore, as described above, which transformation matrix among the [48×16] matrix and the [16×16] matrix to apply to the LFNST is determined by the size and shape of the block to be transformed.
[0200] Figure 8 is a diagram illustrating the block shapes to which LFNST is applied. Figure 8 (a) of Figure 8 shows a 4×4 block, Figure 8 (b) of Figure 8 shows 4×8 blocks and 8×4 blocks, Figure 8 (c) of Figure 8 shows 4×N blocks or N×4 blocks, where N is 16 or greater, Figure 8 (d) of Figure 8 shows an 8×8 block, Figure 8 (e) of Figure 8 shows M×N blocks, where M≥8, N≥8 and N>8 or M>8.
[0201] In Figure 8 , the blocks with thick boundaries indicate the regions to which LFNST is applied. For the Figure 8 blocks of (a) and (b) of Figure 8 , LFNST is applied to the upper left 4×4 region, and for the Figure 8 block of (c) of Figure 8 , LFNST is separately applied to two continuously arranged upper left 4×4 regions. InFigure 8 In (a), (b), and (c), since the LFNST is applied in units of 4×4 regions, the LFNST will be referred to hereinafter as the "4×4 LFNST". Based on the matrix dimensions of G, a [16×16] or [16×8] matrix can be applied.
[0202] More specifically, a [16×8] matrix is applied to Figure 8 the 4×4 blocks (4×4 TUs or 4×4 CUs) in (a), and a [16×16] matrix is applied to Figure 8 the blocks in (b) and (c). This is to adjust the worst-case computational complexity to 8 multiplications per sample.
[0203] Regarding Figure 8 (d) and (e), the LFNST is applied to the upper-left 8×8 region, and this LFNST will be referred to hereinafter as the "8×8 LFNST". As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of the forward LFNST, since a [48×1] vector (the X vector in Equation 9) is input as the input data, not all the sample values in the upper-left 8×8 region are used as the input values of the forward LFNST. That is, as can be seen from Figure 6 the left-side order in (a) or Figure 6 the left-side order in (b), a [48×1] vector can be constructed based on the samples belonging to the remaining 3 4×4 blocks while leaving the lower-right 4×4 block intact.
[0204] [48×8] matrix can be applied to Figure 8 the 8×8 blocks (8×8 TUs or 8×8 CUs) in (d), and a [48×16] matrix can be applied to Figure 8 the 8×8 blocks in (e). This is also to adjust the worst-case computational complexity to 8 multiplications per sample.
[0205] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data (the Y vector in Equation 9, [8×1] or [16×1] vector) are generated. In the forward LFNST, due to the characteristics of the matrix G T the number of output data is equal to or less than the number of input data.
[0206] Figure 9 is a diagram illustrating the arrangement of the output data of the forward LFNST according to the example, and shows the blocks in which the output data of the forward LFNST are arranged according to the block shape.
[0207] In Figure 9The shaded area in the upper left of the block shown corresponds to the area where the output data of the forward LFNST is located. The positions marked with 0 indicate samples filled with the value 0, and the remaining area represents the area not changed by the forward LFNST. In the area not changed by the LFNST, the output data of the forward transform remains unchanged.
[0208] As described above, since the size of the applied transformation matrix varies according to the shape of the block, the number of output data also varies. As Figure 9 , the output data of the forward LFNST may not completely fill the upper left 4×4 block. In Figure 9 cases (a) and (d), a [16×8] matrix and an A[48×8] matrix are applied to the block indicated by the thick line or a partial area inside the block, respectively, and an [8×1] vector is generated as the output of the forward LFNST. That is, according to Figure 7 scan order shown in (b), only 8 output data can be filled, as shown in Figure 9 cases (a) and (d), and 0 can be filled in the remaining 8 positions. In Figure 8 case of the block to which the LFNST in (d) is applied, as shown in Figure 9 (d), the two 4×4 blocks in the upper right and lower left adjacent to the upper left 4×4 block are also filled with the value 0.
[0209] As described above, basically, by signaling the LFNST index, it is specified whether the LFNST is applied and which transformation matrix is to be applied. As Figure 9 shown, when the LFNST is applied, since the number of output data of the forward LFNST can be equal to or less than the number of input data, areas filled with zero values appear as follows.
[0210] 1) As shown in Figure 9 (a), samples from the eighth position and subsequent positions in the scan order in the upper left 4×4 block, that is, samples from the ninth to the sixteenth.
[0211] 2) As shown in Figure 9 (d) and (e), when a [48×16] matrix or a [48×8] matrix is applied, the two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scan order.
[0212] Therefore, if non-zero data exists by checking areas 1) and 2), it is determined that the LFNST is not applied, so that signaling of the corresponding LFNST index can be omitted.
[0213] According to an example, for instance, in the case of LFNST adopted in the VVC standard, since signaling of the LFNST index is performed after residual coding, the encoding device can know whether non-zero data (valid coefficients) exist at all positions within a TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform signaling of the LFNST index based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. When non-zero data does not exist in the regions specified in 1) and 2) above, signaling of the LFNST index is performed.
[0214] Since the truncated unary code is applied as the binarization method of the LFNST index, the LFNST index consists of up to two bins, and 0, 10, and 11 are assigned as the binary codes for the possible LFNST index values 0, 1, and 2, respectively. According to an example, context-based CABAC coding can be applied to the first bin (regular coding), and context-based CABAC coding can also be applied to the second bin. The coding of the LFNST index is shown in the following table.
[0215] [[Table 3]
[0216]
[0217] As shown in Table 3, for the first bin (binIdx = 0), context 0 is applied in the case of a single tree, while in the case of a non-single tree, context 1 can be applied. Additionally, as shown in Table 3, context 2 can be applied to the second bin (binIdx = 1). That is, two contexts can be assigned to the first bin, one context can be assigned to the second bin, and each context can be distinguished by the ctxInc value (0, 1, 2).
[0218] Here, a single tree means that the luminance component and the chrominance component are encoded using the same coding structure. When the coding unit is divided while having the same coding structure, and the size of the coding unit becomes less than or equal to a specific threshold, and the luminance component and the chrominance component are encoded in separate tree structures, the corresponding coding unit is regarded as a double tree, and thus, the context of the first bin can be determined. That is, as shown in Table 3, context 1 can be assigned.
[0219] Alternatively, when the value of the variable treeType is assigned to SINGLE_TREE for the first bin, context 0 can be used, otherwise context 1 can be used.
[0220] In addition, for the adopted LFNST, the following simplified method can be applied.
[0221] (i) According to the example, the number of output data of the forward LFNST can be limited to a maximum of 16.
[0222] In Figure 8 case (c), the 4×4 LFNST can be applied to two adjacent 4×4 regions in the upper left, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to the maximum of 16, in the case of a 4×N / N×4 (N≥16) block (TU or CU), the 4×4 LFNST is only applied to one 4×4 region in the upper left, and the LFNST can only be applied to Figure 8 all the blocks once. By this, the implementation of image coding can be simplified.
[0223] Figure 10 shows that the number of output data of the forward LFNST according to the example is limited to the maximum of 16. As Figure 10 , when the LFNST is applied to the uppermost left 4×4 region in a 4×N or N×4 block (where N is 16 or greater), the output data of the forward LFNST becomes 16.
[0224] (ii) According to the example, the regions to which the LFNST is not applied can be additionally cleared. In this document, clearing can mean filling all positions belonging to a specific region with a value of 0. That is, clearing can be applied to the regions that are not changed due to the LFNST, and the result of the forward single transformation can be maintained. As described above, since the LFNST is divided into 4×4 LFNST and 8×8 LFNST, the clearing can be divided into two types as follows ((ii)-(A) and (ii)-(B)).
[0225] (ii)-(A) When the 4×4 LFNST is applied, the regions to which the 4×4 LFNST is not applied can be cleared. Figure 11 is a diagram illustrating the clearing in the block to which the 4×4 LFNST is applied according to the example.
[0226] As Figure 11 shown, regarding the block to which the 4×4 LFNST is applied, that is, for all the blocks in Figure 9 (a), (b), and (c), the entire region to which the LFNST is not applied can be filled with zeros.
[0227] On the other hand, Figure 11 (d) of Figure 10 shows that when the maximum value of the number of output data of the forward LFNST is limited to 16 (as
[0228] (ii)-(B) When applying the 8×8 LFNST, the areas where the 8×8 LFNST is not applied can be cleared to zero. Figure 12 It is a diagram illustrating the clearing in the block for applying the 8×8 LFNST according to the example.
[0229] As Figure 12 shown, regarding the block for applying the 8×8 LFNST, that is, for all the blocks in Figure 9 (d) and (e), the entire area where the LFNST is not applied can be filled with zeros.
[0230] (iii) Due to the clearing presented in (ii) above, the area filled with zeros may not be the same as when the LFNST is applied. Therefore, the clearing proposed in (ii) can be performed for a wider area according to the case of comparing the Figure 9 LFNST to check for the presence of non-zero data.
[0231] For example, when (ii)-(B) is applied, after checking for the presence of non-zero data in the area filled with zeros in Figure 9 (d) and (e), additionally checking for the presence of non-zero data in the area filled with 0 in Figure 12 , the signaling for the LFNST index can be performed only when there is no non-zero data.
[0232] Of course, even when the clearing proposed in (ii) is applied, the presence of non-zero data can be checked in the same way as the existing LFNST index signaling. That is, after checking for the presence of non-zero data in the blocks filled with zeros in Figure 9 , the LFNST index signaling can be applied. In this case, the encoding device only performs the clearing and the decoding device does not assume the clearing, that is, only checks whether non-zero data exists only in the area Figure 9 explicitly marked as 0 in, and the LFNST index parsing can be performed.
[0233] Various embodiments of the combination of the simplified methods ((i), (ii)-(A), (ii)-(B), (iii)) for applying the LFNST can be derived. Of course, the combination of the above simplified methods is not limited to the following embodiments, and any combination can be applied to the LFNST.
[0234] Embodiment
[0235] - Limit the number of output data of the forward LFNST to a maximum of 16 → (i)
[0236] - When applying the 4×4 LFNST, all areas where the 4×4 LFNST is not applied are cleared to zero → (II)-(A)
[0237] - When applying 8×8 LFNST, all regions where 8×8 LFNST is not applied are cleared → (II)-(B)
[0238] - After checking whether non-zero data also exists in the existing regions filled with zero values and the regions filled with zero due to additional clearing ((ii)-(A), (ii)-(B)), signal the LFNST index only when no non-zero data exists → (iii).
[0239] In the case of the embodiment, when applying LFNST, the regions where non-cleared data can exist are limited to the inside of the upper left 4×4 region. More specifically, in Figure 11 of (a) and Figure 12 of (a), the eighth position in the scanning order is the last position where non-zero data can exist. In Figure 11 of (b) and (c) and Figure 12 of (b), the sixteenth position in the scanning order (i.e., the position of the lower right edge of the upper left 4×4 block) is the last position where data other than 0 can exist.
[0240] Therefore, after applying LFNST, after checking whether non-zero data exists at positions where the residual coding process does not allow (at positions beyond the last position), it is possible to determine whether to signal the LFNST index.
[0241] In the case of the clearing method proposed in (ii), due to the amount of data finally generated when both a transform and LFNST are applied once, the computational amount required to perform the entire transform process can be reduced. That is, when LFNST is applied, since clearing is applied to the region where the forward one-time transform output data exists in the region where LFNST is not applied, there is no need to generate data for the region that becomes cleared during the execution of the forward one-time transform. Therefore, the computational amount required to generate the corresponding data can be reduced. The additional effects of the clearing method proposed in (ii) are summarized as follows.
[0242] First, as described above, the computational amount required to perform the entire transform process is reduced.
[0243] In particular, when applying (ii)-(B), the worst-case computational amount is reduced, making the transform process lighter. In other words, generally, a large amount of computation is required to perform a large-size one-time transform. By applying (ii)-(B), the amount of data derived as a result of performing the forward LFNST can be reduced to 16 or less. Additionally, as the size of the entire block (TU or CU) increases, the effect of reducing the amount of transform operations further increases.
[0244] Second, the computational load required for the entire transformation process can be reduced, thereby reducing the power consumption required to perform the transformation.
[0245] Third, the latency involved in the transformation process is reduced.
[0246] Secondary transformations such as LFNST add computational load to the existing primary transformation, thus increasing the overall latency time involved in performing the transformation. In particular, in the case of intra prediction, since the reconstructed data of adjacent blocks is used during the prediction process, during encoding, the increase in latency due to the secondary transformation leads to an increase in the latency until reconstruction. This can result in an increase in the overall latency of intra prediction coding.
[0247] However, if the zeroing proposed in application (ii) is applied, the latency time for performing the primary transformation can be greatly reduced when LFNST is applied, maintaining or reducing the latency time of the entire transformation, such that the encoding device can be implemented more simply.
[0248] Furthermore, in traditional intra prediction, the block to be currently encoded is regarded as one encoding unit, and encoding is performed without division. However, intra sub - partition (ISP) coding means performing intra prediction coding by dividing the block to be currently encoded in the horizontal or vertical direction. In this case, reconstructed blocks can be generated by performing encoding / decoding in units of the divided blocks, and the reconstructed blocks can be used as reference blocks for the next divided block. According to an embodiment, in ISP coding, one coding block can be divided into two or four sub - blocks and encoded, and in ISP, within one sub - block, intra prediction is performed by referring to the reconstructed pixel values of the sub - blocks located adjacent on the left or adjacent on the upper side. Hereinafter, "encoding" can be used as a concept including both the encoding performed by the encoding device and the decoding performed by the decoding device.
[0249] Furthermore, the signaling order of the LFNST index and the MTS index will be described below.
[0250] According to an example, the LFNST index signaled in the residual coding can be encoded after the coding position of the last non - zero coefficient position, and the MTS index can be encoded immediately after the LFNST index. In the case of this configuration, the LFNST index can be signaled for each transform unit. Alternatively, even if not signaled in the residual coding, the LFNST index can be encoded after the coding of the last valid coefficient position, and the MTS index can be encoded after the LFNST index.
[0251] The syntax of the residual coding according to an example is as follows.
[0252] [Table 4]
[0253]
[0254]
[0255] The meanings of the main variables shown in Table 4 are as follows.
[0256] 1. cbWidth, cbHeight: The width and height of the current coding block
[0257] 2. log2TbWidth, log2TbHeight: The base-2 logarithms of the width and height of the current transform block, which can be reduced by reflection zeroing to the upper left region where non-zero coefficients can exist.
[0258] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is enabled. If the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in the Sequence Parameter Set (SPS).
[0259] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variables chType and the position (x0, y0). chType can have values of 0 and 1, where 0 indicates the luma component and 1 indicates the chroma component. The (x0, y0) position indicates a position on the picture, and MODE_INTRA (intra prediction) and MODE_INTER (inter prediction) can be used as the values of CuPredMode[chType][x0][y0].
[0260] 5. IntraSubPartitionsSplit[x0][y0]: The content at the position (x0, y0) is the same as in item 4 above. It indicates which ISP partition is applied at the position (x0, y0). ISP_NO_SPLIT indicates that the coding unit corresponding to the position (x0, y0) is not divided into sub-blocks.
[0261] 6. intra_mip_flag[x0][y0]: The content at the position (x0, y0) is the same as in item 4 above. intra_mip_flag is a flag indicating whether the matrix-based intra prediction (MIP) prediction mode is applied. If the flag value is 0, it indicates that MIP is not enabled, and if the flag value is 1, it indicates that MIP is enabled.
[0262] 7. cIdx: The value 0 indicates luma, and the values 1 and 2 indicate Cb and Cr of the chroma components respectively.
[0263] 8. treeType: Indicates single tree, dual tree, etc. (SINGLE_TREE: single tree, DUAL_TREE_LUMA: dual tree for luminance component, DUAL_TREE_CHROMA: dual tree for chrominance component)
[0264] 9. tu_cbf_cb[x0][y0]: The content at the position (x0, y0) is the same as in item 4. It indicates the Coding Block Flag (CBF) of the Cb component. If its value is 0, it means that there are no non-zero coefficients in the corresponding transform unit of the Cb component, and if its value is 1, it indicates that there are non-zero coefficients in the corresponding transform unit of the Cb component.
[0265] 10. lastSubBlock: It indicates the position of the sub-block (Coefficient Group (CG)) where the last non-zero coefficient is located in the scan order. 0 indicates the sub-block containing the DC component, and in the case of being greater than 0, it is not the sub-block containing the DC component.
[0266] 11. lastScanPos: It indicates the position where the last valid coefficient is located in a sub-block in the scan order. If a sub-block includes 16 positions, values from 0 to 15 can be there.
[0267] 12. lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If not parsed, it is inferred to be the value 0. That is, the default value is set to 0, indicating that LFNST is not applied.
[0268] 13. LastSignificantCoeffX, LastSignificantCoeffY: They indicate the x and y coordinates where the last valid coefficient is located in the transform block. The x coordinate starts from 0 and increases from left to right, and the y coordinate starts from 0 and increases from top to bottom. If the values of both variables are 0, it means that the last valid coefficient is located at DC.
[0269] 14. cu_sbt_flag: A flag indicating whether Sub-Block Transform (SBT) included in the current VVC standard is enabled. If the flag value is 0, it indicates that SBT is not enabled, and if the flag value is 1, it indicates that SBT is enabled.
[0270] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: Flags indicating whether explicit MTS is applied to inter CUs and intra CUs respectively. If the corresponding flag value is 0, it indicates that MTS is not enabled for inter CUs or intra CUs, and if the corresponding flag value is 1, it indicates that MTS is enabled.
[0271] 16. tu_mts_idx[x0][y0]: The MTS index syntax element to be parsed. If not parsed, it is inferred to be the value 0. That is, the default value is set to 0, indicating that DCT-2 is enabled in both the horizontal and vertical directions.
[0272] As shown in Table 4, in the case of a single tree, it is possible to determine whether to signal the LFNST index using only the last significant coefficient position condition for luminance. That is, if the position of the last significant coefficient is not DC and the last significant coefficient exists in the top-left sub-block (CG) (e.g., a 4×4 block), the LFNST index is signaled. In this case, in the case of 4×4 transform blocks and 8×8 transform blocks, the LFNST index is signaled only when the last significant coefficient exists at positions 0 to 7 in the top-left sub-block.
[0273] In the case of a dual tree, independent of each of luminance and chrominance, the LFNST index is signaled, and in the case of chrominance, the LFNST index can be signaled by applying only the last significant coefficient position condition to the Cb component. For the Cr component, the corresponding condition may not be checked, and if the CBF value of Cb is 0, the LFNST index can be signaled by applying the last significant coefficient position condition to the Cr component.
[0274] "Min(log2TbWidth, log2TbHeight) >= 2" in Table 4 can be expressed as "Min(tbWidth, tbHeight) >= 4", and "Min(log2TbWidth, log2TbHeight) >= 4" can be expressed as "Min(tbWidth, tbHeight) >= 16".
[0275] In Table 4, log2ZoTbWidth and log2ZoTbHeight respectively mean the base-2 logarithms of the width and height of the top-left region where the last significant coefficient can exist by zeroing.
[0276] As shown in Table 4, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places. The first is before parsing the MTS index or the LFNST index value, and the second is after parsing the MTS index.
[0277] The first update is before parsing the value of the MTS index (tu_mts_idx[x0][y0]), so the log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.
[0278] After parsing the MTS index, set log2ZoTbWidth and log2ZoTbHeigh for MTS indices (DST-7 / DCT-8 combination) greater than 0. When DST-7 / DCT-8 is independently applied to each of the horizontal and vertical directions in a single transformation, there can be up to 16 valid coefficients per row or column in each direction. That is, after applying DST-7 / DCT-8 with a length of 32 or greater, up to 16 transform coefficients can be derived for each row or column starting from the left or top. Therefore, in a 2D block, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions, valid coefficients can exist only in the upper-left region of up to 16×16.
[0279] In addition, when DCT-2 is independently applied to each of the horizontal and vertical directions in the current single transformation, there can be up to 32 valid coefficients per row or column in each direction. That is, when applying DCT-2 with a length of 64 or greater, up to 32 transform coefficients can be derived for each row or column starting from the left or top. Therefore, in a 2D block, when DCT-2 is applied to both the horizontal and vertical directions, valid coefficients can exist only in the upper-left region of up to 32×32.
[0280] In addition, when DST-7 / DCT-8 is applied to one side and DCT-2 is applied to the other side for the horizontal and vertical directions, there can be 16 valid coefficients in the former direction and 32 valid coefficients in the latter direction. For example, in the case of a 64×8 transform block, if DCT-2 is applied in the horizontal direction and DST-7 is applied in the vertical direction (which may occur when applying implicit MTS), valid coefficients can exist in the upper-left 32×8 region up to.
[0281] If, as shown in Table 4, log2ZoTbWidth and log2ZoTbHeight are updated in two places, that is, before parsing the MTS index, the ranges of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight as shown in the following table.
[0282] [Table 5]
[0283]
[0284] Additionally, in this case, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values during the binarization process of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.
[0285] [Table 6]
[0286]
[0287] According to the example, in the case of applying the ISP mode and LFNST, when applying the signaling of Table 4, the specification text can be configured as shown in Table 7. Compared with Table 4, the condition for signaling the LFNST index only in the case of not including the ISP mode (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT in Table 4) is removed.
[0288] In a single tree, when the LFNST index sent for the luminance component (cIdx = 0) is reused for the chrominance component, the LFNST index sent for the first ISP partition block with valid coefficients can be applied to the chrominance transform block. Alternatively, even in a single tree, the LFNST index for the chrominance component can be signaled separately from the LFNST index signaled for the luminance component. The descriptions of the variables in Table 7 are the same as those in Table 4.
[0289] [Table 7]
[0290]
[0291] According to the example, the LFNST index and / or the MTS index can be signaled at the coding unit level. As described above, the LFNST index can have three values 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate the first candidate and the second candidate among the two LFNST kernel candidates included in the selected LFNST set, respectively. The LFNST index is encoded by truncated unary binarization, and the values 0, 1, 2 can be encoded as the bin strings 0, 10, 11, respectively.
[0292] According to the example, the LFNST can be applied only when the DCT-2 is applied to both the horizontal and vertical directions in one transformation. Therefore, if the MTS index is signaled after the LFNST index is signaled, the MTS index can be signaled only when the LFNST index is 0, and when the LFNST index is not 0, a single transformation can be performed by applying the DCT-2 to both the horizontal and vertical directions without signaling the MTS index.
[0293] The MTS index can have values 0, 1, 2, 3, and 4, where 0, 1, 2, 3, and 4 can indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, DCT-8 / DCT-8 are applied to the horizontal and vertical directions respectively. Additionally, the MTS index can be encoded by truncated unary binarization, and the values 0, 1, 2, 3, 4 can be encoded as bin strings 0, 10, 110, 1110, 1111 respectively.
[0294] Signaling the LFNST index at the coding unit level can be indicated as shown in the following table. The LFNST index can be signaled in the second half of the coding unit syntax table.
[0295] [Table 8]
[0296]
[0297] The variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag in Table 8 can be set as shown in Table 11 below.
[0298] The variable LfnstDcOnly is equal to 1 when for a transform block with a coding block flag (CBF) of 1 (equal to 0 if there is at least one valid coefficient in the corresponding block, otherwise equal to 0), all the last valid coefficients are located at the DC position (top-left position), and equal to 0 otherwise. Specifically, in the case of dual-tree luminance, the position of the last valid coefficient is checked for a luminance transform block, and in the case of dual-tree chrominance, the position of the last valid coefficient is checked for both the transform blocks of Cb and Cr. In the case of single-tree, the position of the last valid coefficient can be checked for the transform blocks of luminance, Cb, and Cr.
[0299] If there is a valid coefficient at the zeroing position when applying the LFNST, the variable LfnstZeroOutSigCoeffFlag is equal to 0, otherwise equal to 1.
[0300] The lfnst_idx[x0][y0] contained in Table 8 and subsequent tables indicates the LFNST index of the corresponding coding unit, while tu_mts_idx[x0][y0] indicates the MTS index of the corresponding coding unit.
[0301] According to the example, in order to code the MTS index continuously after the LFNST index at the coding unit level, the coding unit syntax table can be configured as shown in Table 9.
[0302] [Table 9]
[0303]
[0304] Comparing Table 9 with Table 8, the condition for checking whether the value of tu_mts_idx[x0][y0] is 0 in the condition for signaling lfnst_idx[x0][y0] (i.e., checking whether DCT-2 is applied to both the horizontal and vertical directions) is changed to the condition for checking whether the value of transform_skip_flag[x0][y0] is 0 (!transform_skip_flag[x0][y0]). transform_skip_flag[x0][y0] indicates whether the coding unit is coded in a transform skip mode in which the transform is skipped, and this flag is signaled before the MTS index and the LFNST index. That is, since lfnst_idx[x0][y0] is signaled before signaling the value of tu_mtx_idx[x0][y0], the condition regarding the value of transform_skip_flag[x0][y0] can be checked only.
[0305] As shown in Table 9, multiple conditions are checked when coding tu_mts_idx[x0][y0], and as described above, tu_mts_idx[x0][y0] is signaled only when the value of lfnst_idx[x0][y0] is 0.
[0306] tu_cbf_luma[x0][y0] is a flag indicating whether there are valid coefficients for the luminance component, and cbWidth and cbHeight indicate the width and height of the coding unit of the luminance component, respectively.
[0307] In Table 9, (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT) indicates that the ISP mode is not applied, and (!cu_sbt_flag) indicates that SBT is not applied.
[0308] According to Table 9, when both the width and height of the coding unit of the luma component are 32 or smaller, tu_mts_idx[x0][y0] is signaled, that is, whether to apply MTS is determined by the width and height of the coding unit of the luma component.
[0309] According to another example, when transform unit (TU) tiling occurs (e.g., when the maximum transform size is set to 32, a 64×64 coding unit is divided into 4 32×32 transform blocks and encoded), the MTS index can be signaled based on the size of each transform block. For example, when both the width and height of the transform block are 32 or smaller, the same MTS index value can be applied to all transform blocks in the coding unit, thereby applying the same single transform. Additionally, when transform block tiling occurs, the value of tu_cbf_luma[x0][y0] in Table 9 can be the CBF value of the top-left transform block, or can be set to 1 when the CBF value of even one transform block among all transform blocks is 1.
[0310] According to the example, when the ISP mode is applied to the current block, LFNST can be applied, and in this case Table 9 can be changed as shown in Table 10.
[0311] [Table 10]
[0312]
[0313] As shown in Table 10, even in the ISP mode (IntraSubPartitionsSplitType!= ISP_NO_SPLIT), lfnst_idx[x0][y0] can be configured to be signaled, and the same LFNST index value can be applied to all ISP sub-blocks.
[0314] Furthermore, as shown in Table 10, since tu_mts_idx[x0][y0] is signaled only in modes other than the ISP mode, the MTS index coding part is the same as in Table 9.
[0315] As shown in Tables 9 and 10, when the MTS index is signaled immediately after the LFNST index, information about the single transform cannot be known during residual coding. That is, the MTS index is signaled after residual coding. Therefore, in the residual coding part, the part where zeroing is performed while only retaining 16 coefficients for DST-7 or DCT-8 with a length of 32 can be changed as shown in Table 11 below.
[0316] [Table 11]
[0317]
[0318]
[0319] As shown in Table 11, in the process of determining log2ZoTbWidth and log2ZoTbHeight (where log2ZoTbWidth and log2ZoTbHeight respectively represent the base-2 logarithms of the width and height of the upper-left region remaining after performing zero clearing), the check of the value of tu_mts_idx[x0][y0] can be omitted.
[0320] The binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 11 can be determined based on log2ZoTbWidth and log2ZoTbHeight as shown in Table 6.
[0321] In addition, as shown in Table 11, when determining log2ZoTbWidth and log2ZoTbHeight in residual coding, a condition for checking the sps_mts_enable_flag can be added.
[0322] TR in Table 6 indicates the truncated Rice binarization method, and the last significant coefficient information can be binarized based on cMax and cRiceParam defined in Table 6 according to the method described in the following table.
[0323] [Table 12]
[0324]
[0325] In addition, according to another example, the coding unit syntax table, transform unit syntax table, and residual coding syntax table are as follows. According to Table 13, the MTS index moves from the transform unit level to the coding unit level syntax and is signaled after the LFNST index signaling. Additionally, the constraint that does not allow LFNST when ISP is applied to the coding unit has been removed. When ISP is applied to the coding unit, the constraint that does not allow LFNST is removed so that LFNST can be applied to all intra prediction blocks. Additionally, both the MTS index and the LFNST index are conditionally signaled at the end part of the coding unit level.
[0326] [Table 13]
[0327]
[0328] [Table 14]
[0329]
[0330] [Table 15]
[0331]
[0332] In Table 13, MtsZeroOutSigCoeffFlag is initially set to 1 and this value can be changed in the residual coding of Table 15. When there are valid coefficients in the area filled with 0 by zeroing (LastSignificantCoeffX>15||LastSignificantCoeffY>15), the value of the variable MtsZeroOutSigCoeffFlag changes from 1 to 0. In this case, the MTS index is not signaled, as shown in Table 15.
[0333] In addition, as shown in Table 13, when tu_cbf_luma[x0][y0] is 0, the coding of mts_idx[x0][y0] can be omitted. That is, when the CBF value of the luminance component is 0, since no transform is applied, there is no need to signal the MTS index, so the MTS index coding can be omitted.
[0334] According to the example, the above technical features can be implemented with another conditional syntax. For example, after performing MTS, a variable indicating whether there are valid coefficients in the area of the current block except for the DC area can be derived, and when this variable indicates that there are valid coefficients in the area except for the DC area, the MTS index can be signaled. That is, the presence of valid coefficients in the area of the current block except for the DC area indicates that the value of tu_cbf_luma[x0][y0] is 1, and in this case, the MTS index can be signaled.
[0335] This variable can be denoted as MtsDcOnly, and after the variable MtsDcOnly is initially set to 1 at the coding unit level, this value is changed to 0 when it is determined at the residual coding level that there are valid coefficients in the area of the current block except for the DC area. When the variable MtsDcOnly is 0, the image information can be configured to signal the MTS index.
[0336] When tu_cbf_luma[x0][y0] is 0, since the residual coding syntax is not called at the transform unit level of Table 14, the initial value of the variable MtsDcOnly is maintained as 1. In this case, since the variable MtsDcOnly is not changed to 0, the image information can be configured so that the MTS index is not signaled. That is, the MTS index is not parsed and signaled.
[0337] In addition, the decoding device may determine the color index cIdx of the transform coefficient to derive the variable MtsZeroOutSigCoeffFlag in Table 15. A color index cIdx of 0 represents the luminance component.
[0338] According to the example, since MTS can be applied only to the luminance component of the current block, the decoding device may determine whether the color index is luminance when deriving the variable MtsZeroOutSigCoeffFlag used to determine whether to parse the MTS index.
[0339] The variable MtsZeroOutSigCoeffFlag is a variable indicating whether zeroing is performed when applying MTS. It indicates whether there are transform coefficients in the area outside the upper-left area where the last valid coefficient can exist due to zeroing after performing MTS (i.e., in the area outside the upper-left 16×16 area). The variable MtsZeroOutSigCoeffFlag is initially set to 1 at the coding unit level, as shown in Table 13 (MtsZeroOutSigCoeffFlag = 1), and its value may change from 1 to 0 at the residual coding level when there are transform coefficients in the area outside the 16×16 area, as shown in Table 15 (MtsZeroOutSigCoeffFlag = 0). When the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.
[0340] As shown in Table 15, at the residual coding level, the non-zeroing area where non-zero transform coefficients can exist may be set according to whether zeroing accompanying MTS is performed. And even in this case, with the color index (cIdx) being 0, the non-zeroing area may be set to the upper-left 16×16 area of the current block.
[0341] Thus, when deriving the variable for determining whether the MTS index is parsed, the color component is determined to be luminance or chrominance. However, since LFNST can be applied to both the luminance component and the chrominance component of the current block, the color component is not determined when deriving the variable for determining whether to parse the LFNST index.
[0342] For example, Table 13 shows the variable LfnstZeroOutSigCoeffFlag, which can indicate zeroing when applying LFNST. The variable LfnstZeroOutSigCoeffFlag indicates whether there are valid coefficients in the second region of the current block except for the first region in the upper left. The value is initially set to 1, and when there are valid coefficients in the second region, the value can be changed to 0. Only when the value of the initially set variable LfnstZeroOutSigCoeffFlag remains 1 can the LFNST index be parsed. When determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, since LFNST can be applied to both the luminance component and the chrominance component of the current block, the color index of the current block is not determined.
[0343] Figure 13 An example of MIP-based predicted sample generation processing according to an example is illustrated. Referring to Figure 13 , the MIP processing is described as follows.
[0344] 1. Averaging processing
[0345] Among the boundary samples, four samples for the case of W = H = 4 and eight samples for any other case are extracted through averaging processing.
[0346] 2. Matrix-vector multiplication processing
[0347] Matrix-vector multiplication is performed using the averaged samples as input, followed by adding an offset. Through this operation, a reduced set of predicted samples of the subsampled samples in the original block can be derived.
[0348] 3. (Linear) interpolation processing
[0349] The predicted samples at the remaining positions are generated based on the predicted samples of the subsampled sample set through linear interpolation, which is single-step linear interpolation in each direction.
[0350] For the matrix, the matrices and offset vectors required to generate the predicted block or predicted samples can be selected from three sets S0, S1, and S2.
[0351] Set S0 can include 16 matrices A0 i , i ∈ {0,..., 15} and 16 offset vectors b0 i , i ∈ {0,..., 15}, and each matrix can include 16 rows and 4 columns. The matrices and offset vectors in set S0 can be used for 4×4 blocks. In another example, set S0 can include 18 matrices.
[0352] Set S1 can include 8 matrices A1i where \(i\in\{0,\ldots,7\}\) and there are 8 offset vectors \(\mathbf{b}_1\) i where \(i\in\{0,\ldots,7\}\), and each matrix can include 16 rows and 8 columns. In another example, the set \(S_1\) can include 6 matrices. The matrices and offset vectors in the set \(S_1\) can be used for \(4\times8\) blocks, \(8\times4\) blocks, and \(8\times8\) blocks. Alternatively, the matrices and offset vectors in the set \(S_1\) can be used for \(4\times H\) blocks or \(W\times4\) blocks.
[0353] Finally, the set \(S_2\) can include 6 matrices \(\mathbf{A}_2\) i where \(i\in\{0,\ldots,5\}\) and there are 6 offset vectors \(\mathbf{b}_2\) i where \(i\in\{0,\ldots,5\}\), and each matrix can include 64 rows and 8 columns. The matrices and offset vectors in the set \(S_2\) or some of them can be used for any blocks with different sizes to which the sets \(S_0\) and \(S_1\) are not applied. For example, the matrices and offset vectors in the set \(S_2\) can be used for operations on blocks with a height and width of 8 or greater.
[0354] The total number of multiplications required to compute the matrix - vector product is always less than or equal to \(4\times W\times H\). That is, in the MIP mode, up to four multiplications per sample are required.
[0355] Figure 14 Illustrates the CCLM applicable when deriving the intra - prediction mode of a chrominance block according to an embodiment.
[0356] In this specification, a "reference sample template" may refer to a set of adjacent reference samples of the current chrominance block used to predict the current chrominance block. The reference sample template can be predefined, and information about the reference sample template can be signaled from the encoding device 100 to the decoding device 200.
[0357] Referring to Figure 14 , the set of shaded samples in a single row adjacent to the \(4\times4\) block that is the current chrominance block refers to the reference sample template. The reference sample template is configured as the reference samples of a single row, while the reference sample region in the luminance region corresponding to the reference sample template is configured as two rows, as Figure 14 shown.
[0358] In an embodiment, when performing intra - coding of a chrominance image in the Joint Exploration Test Model (JEM) used in the Joint Video Exploration Team (JVET), a Cross - Component Linear Model (CCLM) can be used. CCLM is a method of predicting the pixel values of a chrominance image based on the pixel values of a reconstructed luminance image and is based on the high correlation between the luminance image and the chrominance image.
[0359] The CCLM prediction of the \(C_b\) and \(C_r\) chrominance images can be performed based on the following formula.
[0360] [Formula 11]
[0361] Pred C (i,j) = α·Rec' L (i,) + β
[0362] Here, Pred c (i,j) represents the Cb or Cr chrominance image to be predicted, Rec L '(i,j) represents the reconstructed luminance image adjusted to the chrominance block size, and (i,j) represents the coordinates of the pixel. In the 4:2:0 color format, since the size of the luminance image is twice that of the chrominance image, it is necessary to generate Rec L ' with the chrominance block size by downsampling. Therefore, the pixels of the luminance image to be used for the chrominance image Pred L (i,j) can be adopted considering Rec c (2i,2j) and the adjacent pixels. Rec L '(i,j) can be called the downsampled luminance sample.
[0363] For example, as shown in the following formula, Rec L '(i,j) can be derived using six adjacent pixels.
[0364] [Formula 12]
[0365] Rec' L (x,y) = (2×Rec L (2x,2y) + 2×Rec L (2x,2y + 1) + Rgc L (2x - 1,2y) + Rec L (2x + 1,2y) + Rec L (2x - 1,2y + 1) + Rec L (2x + 1,2y + 1) + 4) >> 3
[0366] α and β represent the cross - correlation and the average difference between the adjacent templates of the Cb or Cr chrominance block and Figure 14 the adjacent templates of the luminance block in the shaded area in. For example, α and β are represented by Formula 13.
[0367] [Formula 13]
[0368]
[0369]
[0370] L(n) represents the neighboring reference samples and / or the left neighboring samples of the luminance block corresponding to the current chrominance image, C(n) represents the neighboring reference samples and / or the left neighboring samples of the current chrominance block to which coding is currently applied, and (i,j) represents the pixel position. Additionally, L(n) can represent the upsampled upper neighboring samples and / or the left neighboring samples of the current luminance block. N can represent the total number of pixel pairs (luminance and chrominance) values used to calculate the CCLM parameter, and can indicate a value that is twice the smaller value of the width and height of the current chrominance block.
[0371] A picture can be partitioned into a sequence of coding tree units (CTUs). A CTU can correspond to a coding tree block (CTB). Alternatively, a CTU can include a coding tree block of luminance samples and a coding tree block of corresponding chrominance samples. Depending on whether the luminance block and the corresponding chrominance block have a separate partitioning structure, the tree type can be classified as a single tree (SINGLE_TREE) or a dual tree (DUAL_TREE). A single tree can indicate that the chrominance block has the same partitioning structure as the luminance block, while a dual tree can indicate that the chrominance component block has a partitioning structure different from that of the luminance block.
[0372] When applying LFNST to a chrominance transform block according to an example, information about the collocated luminance transform block needs to be referenced.
[0373] The existing specification text regarding relevant parts is shown in the following table.
[0374] [Table 16]
[0375]
[0376] As shown in Table 16, when the intra prediction mode in the current frame is the CCLM mode, the value of the variable predModeIntra for the chrominance transform block is determined by taking the intra prediction mode value of the co-located chrominance transform block (the italicized part). The intra prediction mode value (predModeIntra value) of the luminance transform block can subsequently be used to determine the LFNST set.
[0377] However, the variables nTbW and nTbH input as input values for this transform process represent the width and height of the current transform block. Therefore, when the current block is a luminance transform block, the variables nTbW and nTbH can represent the width and height of the luminance transform block, while when the current block is a chrominance transform block, the variables nTbW and nTbH represent the width and height of the chrominance transform block.
[0378] Here, the variables nTbW and nTbH in the italicized part of Table 16 represent the width and height of the chrominance transform block that do not reflect the color format and thus do not accurately indicate the reference position of the luminance transform block corresponding to the chrominance transform block. Therefore, the italicized part of Table 16 can be modified as shown in the following table.
[0379] [Table 17]
[0380]
[0381] As shown in Table 17, nTbW and nTbH are respectively changed to (nTbW * SubWidthC) / 2 and (nTbH * SubHeightC) / 2. xTbY and yTbY can represent the luminance position in the current picture (the upper-left sample of the current luminance transform block relative to the upper-left luminance sample of the current picture), while nTbW and nTbH can represent the width and height of the currently encoded transform block (the variable nTbW specifies the width of the current transform block, and the variable nTbH specifies the height of the current transform block).
[0382] When the currently encoded transform block is a chrominance (Cb or Cr) transform block, nTbW and nTbH are respectively the width and height of the chrominance transform block. Therefore, when the currently encoded transform block is a chrominance transform block (cIdx > 0), the width and height of the luminance transform block need to be used to obtain the reference position of the juxtaposed luminance transform block when obtaining the reference position. In Table 17, SubWidthC and SubHeightC are values set according to the color format (such as 4:2:0, 4:2:2, or 4:4:4). Specifically, they are respectively the width ratio and height ratio between the luminance component and the chrominance component (see Table 18 below). Therefore, in the case of a chrominance transform block, (nTbW * SubWidthC) and (nTbH * SubHeightC) can be respectively the width and height relative to the juxtaposed luminance transform block.
[0383] Therefore, xTbY + (nTbW * SubWidthC) / 2 and yTbY + (nTbH * SubHeightC) / 2 represent the values of the central position in the juxtaposed luminance transform block based on the upper-left position of the current picture, and thus precisely indicate the juxtaposed luminance transform block.
[0384] [Table 18]
[0385] Chrominance format SubWidthC SubHeightC Monochrome 1 1 4:2:0 2 2 4:2:2 2 1 4:4:4 1 1 4:4:4 1 1
[0386] In Table 17, the variable predModeIntra represents the intra prediction mode value. When the value of the variable predModeIntra is equal to INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM, it indicates that the current transform block is a chrominance transform block. According to the example, in the current VVC standard, INTRA_LT_CCLM, INTRA_L_CCLM, and INTRA_T_CCLM correspond to the mode values 81, 82, and 83 among the intra prediction mode values respectively. Therefore, as shown in Table 17, the values of xTbY+(nTbW*SubWidthC) / 2 and yTbY+(nTbH*SubHeightC) / 2 are required to obtain the reference positions of the adjacent luminance transform blocks.
[0387] As shown in Table 17, given that both the variable intra_mip_flag[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] and the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] update the value of predModeIntra.
[0388] intra_mip_flag is a variable indicating whether the current transform block (or coding unit) is encoded by the matrix-based intra prediction (MIP) method, and intra_mip_flag[x][y] is a flag value indicating whether MIP is applied to the position corresponding to the coordinates (x, y) based on the luminance component when the upper left position of the current picture is defined as (0,0). The x and y coordinates increase from left to right and from top to bottom respectively, and when the flag indicating whether to apply MIP is 1, the flag indicates that MIP is applied. When the flag indicating whether to apply MIP is 0, the flag indicates that MIP is not applied. MIP can be applied only to luminance blocks.
[0389] According to the modified part of Table 17, when the value of intra_mip_flag[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] in the adjacent luminance transform block is 1, the value of predModeIntra is set to the planar mode (INTRA_PLANAR).
[0390] The value of the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the prediction mode value corresponding to the coordinates (xTbY+(nTbW*SubWidthC) / 2, yTbY+(nTbH*SubHeightC) / 2) when the upper left position of the current picture in the luma component is defined as (0,0). The prediction mode value can have the values MODE_INTRA, MODE_IBC, MODE_PLT, and MODE_INTER which represent the intra prediction mode, intra block copy (IBC) prediction mode, palette (PLT) coding mode, and inter prediction mode respectively. According to Table 17, when the value of the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] is MODE_IBC or MODE_PLT, the value of the variable predModeIntra is set to the DC mode. In cases other than these two cases, the value of the variable predModeIntra is set to IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] (the intra prediction mode value corresponding to the central position in the collocated luma transform block).
[0391] According to the example, as shown in the following table, considering whether to perform wide-angle intra prediction, the value of the variable predModeIntra can be updated one more time based on the predModeIntra value updated in Table 17.
[0392] [Table 19]
[0393]
[0394] In the mapping process shown in Table 19, the input values of predModeIntra, nTbW, and nTbH are the same as the updated variable predModeIntra in Table 17 and the values of nTbW and nTbH referenced in Table 17 respectively.
[0395] In Table 19, nCbW and nCbH respectively represent the width and height of the coding block corresponding to the transform block, and the variable IntraSubPartitionsSplitType indicates whether the ISP mode is applied. Where IntraSubPartitionsSplitType being equal to ISP_NO_SPLIT indicates that the coding unit is not partitioned by ISP (i.e., the ISP mode is not applied). The variable IntraSubPartitionsSplitType not equal to ISP_NO_SPLIT indicates that the ISP mode is applied, and thus the coding unit is partitioned into two or four sub-blocks. In Table 19, cIdx is an index indicating the color component. A cIdx value equal to 0 represents a luminance block, while a cIdx value not equal to 0 indicates a chrominance block. The predModeIntra value output through the mapping process in Table 19 is a value updated considering whether the wide-angle intra prediction (WAIP) mode is applied.
[0396] For the predModeIntra value updated through Table 19, the LFNST set can be determined through the mapping relationship shown in the following table.
[0397] [Table 20]
[0398] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 44 3 45 <= predModeIntra <= 55 2 56 <= predModeIntra <= 80 1
[0399] In the above table, lfnstTrSetIdx represents an index indicating the LFNST set and has values from 0 to 3, which indicates that a total of four LFNST sets are configured. Each LFNST set can include two transform kernels, i.e., LFNST kernels (depending on the region where LFNST is applied and based on the forward direction, the transform kernel can be a 16×16 matrix or a 16×48 matrix), and the transform kernel to be applied among the two transform kernels can be specified by signaling of the LFNST index. Additionally, whether to apply LFNST can also be specified by the LFNST index. In the current VVC standard, the LFNST index can have values 0, 1, and 2, where 0 indicates not applying LFNST, and 1 and 2 respectively indicate the two transform kernels.
[0400] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices illustrated in the drawings or the names of specific signals / messages / fields are provided for illustration purposes, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0401] Figure 15 is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.
[0402] Figure 15 Each process disclosed in Figures 3 to 14Some details of the description. Therefore, the description of specific details that overlap with the details described with reference to Figures 3 to 14 will be omitted or will be described schematically.
[0403] According to an embodiment, the decoding device 200 may obtain intra prediction mode information and LFNST index from a bitstream (S1510).
[0404] The intra prediction mode information may include the intra prediction modes of adjacent blocks (e.g., left adjacent block and / or upper adjacent block) of the current block, and an MPM index indicating one of the MPM candidates in the list of most probable modes (MPM) derived based on additional candidate modes or remaining intra prediction mode information indicating one of the remaining intra prediction modes not included in the MPM candidates.
[0405] In addition, the intra mode information may include flag information sps_cclm_enabled_flag indicating whether CCLM is applied to the current block and information intra_chroma_pred_mode about the intra prediction mode of the chrominance component.
[0406] The LFNST index information is received as syntax information, and the syntax information is received as a binary bin string including 0 and 1.
[0407] The syntax element of the LFNST index according to the present embodiment may indicate whether to apply inverse LFNST or inverse non-separable transform and any one of the transform kernel matrices included in the transform set, and when the transform set includes two transform kernel matrices, the syntax element of the transform index may have three values.
[0408] That is, according to an embodiment, the value of the syntax element of the LFNST index may include: 0, which indicates that inverse LFNST is not applied to the target block; 1, which indicates the first transform kernel matrix among the transform kernel matrices; and 2, which indicates the second transform kernel matrix among the transform kernel matrices.
[0409] The decoding device 200 may decode information about the quantized transform coefficients of the current block from the bitstream, and may derive the quantized transform coefficients of the target block based on the information about the quantized transform coefficients of the current block. The information about the quantized transform coefficients of the target block may be included in the sequence parameter set (SPS) or the slice header, and may include at least one of information about whether to apply RST, information about the reduction factor, information about the minimum transform size for applying RST, information about the maximum transform size for applying RST, inverse RST size, and information about the transform index indicating any one of the transform kernel matrices included in the transform set.
[0410] The decoding device 200 can derive transform coefficients by dequantizing the residual information (i.e., quantized transform coefficients) regarding the current block, and can arrange the derived transform coefficients in a predetermined scanning order.
[0411] Specifically, the derived transform coefficients can be arranged in units of 4×4 blocks according to the inverse diagonal scanning order, and the transform coefficients in the 4×4 blocks can also be arranged according to the inverse diagonal scanning order. That is to say, the dequantized transform coefficients can be arranged according to the inverse scanning order applied in a video codec (such as in VVC or HEVC).
[0412] The transform coefficients derived based on the residual information can be the dequantized transform coefficients as described above, or can be quantized transform coefficients. That is to say, the transform coefficients can be any data used to check for non-zero data in the current block, regardless of quantization.
[0413] The decoding device can update the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block when the intra prediction mode of the chrominance block is the CCLM mode. In particular, when the intra prediction mode of the luminance block is the MIP mode, the intra prediction mode of the chrominance block can be updated to the intra-plane mode (S1520).
[0414] The decoding device can derive the intra prediction mode of the chrominance block as the CCLM mode based on the intra prediction mode information. For example, the decoding device can receive information about the intra prediction mode of the current chrominance block through a bitstream, and can derive the intra prediction mode of the current chrominance block as the CCLM mode based on the intra prediction mode information.
[0415] The CCLM mode can include the upper-left CCLM mode, the upper CCLM mode, or the left CCLM mode.
[0416] As described above, the decoding device can derive residual samples by applying the LFNST as an inseparable transform or the MTS as a separable transform, and can perform these transforms based on the LFNST index indicating the LFNST kernel (i.e., the LFNST matrix) and the MTS index indicating the MTS kernel, respectively.
[0417] For the LFNST, it is necessary to determine the LFNST set, and this LFNST set has a mapping relationship with the intra prediction mode of the current block.
[0418] The decoding device can update the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block for the inverse LFNST of the chrominance block.
[0419] According to the example, the updated intra prediction mode can be derived as the intra prediction mode corresponding to a specific position in the luminance block, and this specific position can be set based on the color format of the chrominance block.
[0420] The specific position can be the central position of the luminance block and can be represented by ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)).
[0421] In the central position, xTbY and yTbY represent the upper-left coordinates of the luminance block, that is, the upper-left position in the luminance sample reference of the current transform block, nTbW and nTbH represent the width and height of the chrominance block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)) represents the central position of the luminance transform block, and IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra prediction mode of the luminance block at this position.
[0422] SubWidthC and SubHeightC can be derived as shown in Table 18. That is, when the color format is 4:2:0, SubWidthC and SubHeighC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0423] As shown in Table 17, in order to specify the specific position of the luminance block corresponding to the chrominance block regardless of the color format, the color format is reflected in the variable indicating the specific position.
[0424] As described above, when the intra prediction mode corresponding to the specific position of the luminance block is the matrix-based intra prediction (hereinafter, "MIP") mode, the decoding device can set the updated intra prediction mode to the intra planar mode.
[0425] The MIP mode can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). When MIP is applied to the current block, the predicted samples of the current block can be derived by i) using adjacent reference samples that have undergone averaging processing, ii) performing matrix-vector multiplication processing, and iii) further performing horizontal / vertical interpolation processing.
[0426] Alternatively, according to the example, when the intra prediction mode corresponding to a specific position is the intra block copy (IBC) mode or the palette mode, the decoding device may set the updated intra prediction mode to the intra DC mode.
[0427] The IBC prediction mode or the palette mode may be used to encode content images / videos including games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but may be performed similarly to inter prediction, except that the reference block is derived within the current picture. That is, IBC may use at least one of the inter prediction techniques described in the present disclosure. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, the values of the samples in the picture may be signaled based on the information about the palette table and the palette index.
[0428] In summary, when the intra prediction mode at the central position is the MIP mode, the IBC mode, or the palette mode, the intra prediction mode of the chrominance block may be updated to a specific mode, such as the intra planar mode or the intra DC mode.
[0429] When the intra prediction mode at the central position is not the MIP mode, the IBC mode, and the palette mode, the intra prediction mode of the chrominance block may be updated to the intra prediction mode of the luma block at the central position in order to reflect the correlation between the chrominance block and the luma block.
[0430] The decoding device may determine an LFNST set including the LFNST matrix based on the updated intra prediction mode (S1530), and may derive the transform coefficients of the chrominance block based on the LFNST matrix derived from the LFNST set (S1540).
[0431] Any one of the multiple LFNST matrices may be selected based on the LFNST set and the LFNST index.
[0432] As shown in Table 20, the LFNST transform set is derived according to the intra prediction mode, and 81 to 83 indicating the CCLM mode in the intra prediction mode are omitted because the LFNST transform set is derived using the intra mode value of the corresponding luma block in the CCLM mode.
[0433] According to the example, as shown in Table 20, any one of the four LFNST sets may be determined according to the intra prediction mode of the current block, and the LFNST set to be applied to the current chrominance block may also be determined.
[0434] The decoding device may perform an inverse RST (e.g., inverse LFNST) by applying the LFNST matrix to the dequantized transform coefficients, thereby deriving the modified transform coefficients of the current chrominance block.
[0435] The decoding device can derive the residual samples from the transform coefficients through a single inverse transform (S1550), and when the current block is a chrominance block, it can derive the residual samples of the chrominance block based on the transform coefficients. MTS can be used for the single inverse transform.
[0436] In addition, the decoding device can generate reconstructed samples based on the residual samples of the current block and the prediction samples of the current block.
[0437] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices illustrated in the drawings or the names of specific signals / messages / fields are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0438] Figure 16 is a flowchart illustrating the operations of a video encoding device according to an embodiment of the present disclosure.
[0439] Figure 16 Each process disclosed herein is based on some details described with reference to Figures 3 to 14 Therefore, the specific details overlapping with those described with reference to Figure 1 and Figures 3 to 14 will be omitted or described schematically.
[0440] According to an embodiment, the encoding device 100 can derive the prediction samples of the chrominance block based on the fact that the intra prediction mode of the chrominance block is the CCLM mode (S1610).
[0441] The encoding device can first derive the intra prediction mode of the chrominance block as the CCLM mode.
[0442] For example, the encoding device can determine the intra prediction mode of the current chrominance block based on the rate-distortion (RD) cost (or RDO). Here, the RD cost can be derived based on the sum of absolute differences (SAD). The encoding device can determine the CCLM mode as the intra prediction mode of the current chrominance block based on the RD cost.
[0443] The CCLM mode can include the upper-left CCLM mode, the upper CCLM mode, or the left CCLM mode.
[0444] The encoding device can encode the information about the intra prediction mode of the current chrominance block and can signal the information about the intra prediction mode through the bitstream. The information related to the prediction of the current chrominance block can include the information about the intra prediction mode.
[0445] According to an embodiment, the encoding device can derive the residual samples of the chrominance block based on the prediction samples (S1620).
[0446] According to an embodiment, an encoding device may derive transform coefficients of a chrominance block based on a single transform of residual samples.
[0447] The single transform may be performed by a plurality of transform kernels, and in this case, a transform kernel may be selected based on an intra prediction mode.
[0448] The encoding device may update the intra prediction mode of the chrominance block for LFNST of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block, and may update the intra prediction mode of the chrominance block to an intra plane mode based on the intra prediction mode of the luminance block being a MIP mode (S1630).
[0449] As shown in Table 17, the encoding device may update the CCLM mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block (when predModeIntra is equal to INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM, derive predModeIntra as follows).
[0450] According to an example, the updated intra prediction mode may be derived as an intra prediction mode corresponding to a specific position in the luminance block, and the specific position may be set based on the color format of the chrominance block.
[0451] The specific position may be a central position of the luminance block, and may be represented by ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)).
[0452] In the central position, xTbY and yTbY represent the upper left coordinates of the luminance block, that is, the upper left position in the luminance sample reference of the current transform block, nTbW and nTbH represent the width and height of the chrominance block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)) represents the central position of the luminance transform block, and IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra prediction mode of the luminance block at that position.
[0453] SubWidthC and SubHeightC may be derived as shown in Table 18. That is, when the color format is 4:2:0, SubWidthC and SubHeighC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0454] As shown in Table 17, in order to specify the specific position of the luminance block corresponding to the chrominance block regardless of the color format, the color format is reflected in the variable indicating the specific position.
[0455] As described above, when the intra prediction mode of the luminance block corresponding to the specific position is the matrix-based intra prediction (hereinafter, "MIP") mode, the encoding device may set the updated intra prediction mode to the intra planar mode.
[0456] The MIP mode may be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). When MIP is applied to the current block, the predicted samples of the current block can be derived by i) using adjacent reference samples that have undergone averaging processing, ii) performing matrix-vector multiplication processing, and iii) further performing horizontal / vertical interpolation processing.
[0457] Alternatively, according to an example, when the intra prediction mode corresponding to the specific position is the intra block copy (IBC) mode or the palette mode, the decoding device may set the updated intra prediction mode to the intra DC mode.
[0458] The IBC prediction mode or the palette mode can be used to encode content images / videos including games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction, except that the reference block is derived within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, the values of the samples in the picture can be signaled based on the information about the palette table and the palette index.
[0459] In summary, when the intra prediction mode at the central position is the MIP mode, the IBC mode, or the palette mode, the intra prediction mode of the chrominance block can be updated to a specific mode, such as the intra planar mode or the intra DC mode.
[0460] When the intra prediction mode at the central position is not the MIP mode, the IBC mode, and the palette mode, the intra prediction mode of the chrominance block can be updated to the intra prediction mode of the luminance block at the central position to reflect the association between the chrominance block and the luminance block.
[0461] The encoding device may determine the LFNST set including the LFNST matrix based on the updated intra prediction mode (S1640), and may derive the transform coefficients of the chrominance block based on the residual samples and the LFNST matrix (S1650).
[0462] The encoding device may determine a transform set based on a mapping relationship according to the intra prediction mode applied to the current block, and may perform LFNST (i.e., non-separable transform) based on any one of the two LFNST matrices included in the transform set.
[0463] As described above, multiple transform sets may be determined according to the intra prediction mode of the transform block to be transformed. The matrix applied to LFNST is the transpose of the matrix used in inverse LFNST.
[0464] In one example, the LFNST matrix may be a non-square matrix with the number of rows less than the number of columns.
[0465] The encoding device may derive quantized transform coefficients by performing quantization based on the modified transform coefficients of the current chroma block, and may encode and output image information including information about the quantized transform coefficients, information about the intra prediction mode, and an LFNST index indicating the LFNST matrix (S1660).
[0466] Specifically, the encoding device 100 may generate information about the quantized transform coefficients and may encode the generated information about the quantized transform coefficients.
[0467] In one example, the information about the quantized transform coefficients may include at least one of information about whether LFNST is applied, information about a reduction factor, information about the minimum transform size for applying LFNST, and information about the maximum transform size for applying LFNST.
[0468] The encoding device may encode flag information indicating whether CCLM is applied to the current block as sps_cclm_enabled_flag and information about the intra prediction mode of the chroma component as intra_chroma_pred_mode as information about the intra mode.
[0469] The information about the CCLM mode as intra_chroma_pred_mode may indicate the top-left CCLM mode, the top CCLM mode, or the left CCLM mode.
[0470] In the present disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transformation / inverse transformation is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of consistent expression.
[0471] In addition, in the present disclosure, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled by the residual coding syntax. The transform coefficients may be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients may be derived by the inverse transform (scaling) of the transform coefficients. The residual samples may be derived based on the inverse transform (transformation) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.
[0472] In the above embodiments, the method is explained based on the flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may be executed in an order or steps different from the above order or steps, or a certain step may be executed concurrently with other steps. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and one or more steps in the flowchart may be incorporated or deleted without affecting the scope of the present disclosure.
[0473] The above method according to the present disclosure may be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in a device for image processing such as a television, a computer, a smart phone, a set-top box, and a display device.
[0474] When the embodiments in the present disclosure are implemented by software, the above method may be implemented as modules (steps, functions, etc.) for performing the above functions. These modules may be stored in a memory and may be executed by a processor. The memory may be inside or outside the processor and may be connected to the processor in various well-known ways. The processor may include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure may be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing may be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.
[0475] In addition, the decoding device and encoding device applying the present disclosure may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device (such as video communication), a mobile streaming device, a storage medium, a camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0476] In addition, the processing method applying the present disclosure may be produced in the form of a program executable by a computer and may be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure may also be stored in the computer-readable recording medium. The computer-readable recording medium includes various storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, the bitstream generated by the encoding method may be stored in the computer-readable recording medium or transmitted through a wired or wireless communication network. In addition, embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer according to the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0477] Figure 17 Examples of a video / image encoding system to which the present disclosure may be applied are illustrated.
[0478] Referring to Figure 17 , the video / image encoding system may include a source device and a receiving device. The source device may transfer the encoded video / image information or data to the receiving device in the form of a file or a stream through a digital storage medium or a network.
[0479] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0480] The video source can obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet computer, and a smart phone, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer or the like. In this case, the video / image capture process can be replaced by a process of generating relevant data.
[0481] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0482] The transmitter can send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include an element for generating a media file in a predetermined file format and can include an element for sending through a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to the decoding device.
[0483] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc., corresponding to the operations of the encoding device.
[0484] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0485] Figure 18 The structure of a content stream system to which the present disclosure is applied is illustrated.
[0486] In addition, the content stream system to which the present disclosure is applied can generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0487] The encoding server is used to compress the content input from multimedia input devices such as smart phones, cameras, video cameras, etc. into digital data to generate a bitstream, and send it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, camera, video camera, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method of the present disclosure. And the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0488] Based on the user's request, the streaming server sends multimedia data to the user device through the web server, and the web server serves as a device to notify the user of what services exist. When the user requests the service they want, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content streaming system may include a separate control server, and in this case, the control server is used to control the commands / responses between the corresponding devices in the content streaming system.
[0489] The streaming server can receive content from the media storage device and / or the encoding server. For example, in the case of receiving content from the encoding server, the content can be received in real time. In this case, in order to smoothly provide the streaming service, the streaming server can store the bitstream for a predetermined time.
[0490] For example, the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigator, a slate PC, a tablet PC, a Ultrabook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital sign, etc. Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0491] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined to be implemented or executed in a device, and the technical features of the device claims can be combined to be implemented or executed in a method. In addition, the technical features of the method claims and the device claims can be combined to be implemented or executed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or executed in a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtain intra prediction mode information from a bitstream; Based on the intra prediction mode of a chrominance block being a cross-component linear model (CCLM) mode, update the intra prediction mode of the chrominance block based on the intra prediction mode of a luminance block corresponding to the chrominance block; Determine an LFNST set including an LFNST matrix based on the updated intra prediction mode; Derive transform coefficients for the chrominance block based on the LFNST matrix derived from the LFNST set; And Derive residual samples for the chrominance block based on the transform coefficients, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a specific position in the luminance block, wherein based on the intra prediction type corresponding to the specific position being a MIP mode, the intra prediction mode of the chrominance block is updated to an intra plane mode, wherein the specific position is set based on the color format of the chrominance block, wherein the specific position is set to ((xTbY + (nTbW * SubWidthC) / 2), (yTbY + (nTbH * SubHeightC) / 2)), wherein xTbY and yTbY represent the upper left coordinates of the luminance block, wherein nTbW and nTbH represent the width and height of the chrominance block, wherein SubWidthC and SubHeightC represent variables corresponding to the color format, wherein when the color format is 4:2:0, SubWidthC and SubHeightC are 2, and wherein when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
2. The image decoding method according to claim 1, wherein, The specific position is the central position of the luminance block.
3. The image decoding method according to claim 1, wherein, When the prediction mode corresponding to the specific position is an IBC mode, the intra prediction mode of the chrominance block is updated to an intra DC mode.
4. The image decoding method according to claim 1, wherein, When the prediction mode corresponding to the specific position is a palette mode, the intra prediction mode of the chrominance block is updated to an intra DC mode.
5. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Derive prediction samples for the chrominance block based on the intra prediction mode of the chrominance block being a cross-component linear model (CCLM); Derive residual samples for the chrominance block based on the prediction samples; Update the intra prediction mode of the chrominance block based on the intra prediction mode of a luminance block corresponding to the chrominance block; Determine an LFNST set including an LFNST matrix based on the updated intra prediction mode; And Derive modified transform coefficients for the chrominance block based on the residual samples and the LFNST matrix, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a specific position in the luminance block, Wherein, based on that the intra prediction type corresponding to the specific position is the MIP mode, the intra prediction mode of the chrominance block is updated to the intra plane mode. Wherein, the specific position is set based on the color format of the chrominance block. Wherein, the specific position is set to ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)). Wherein, xTbY and yTbY represent the upper left coordinates of the luminance block. Wherein, nTbW and nTbH represent the width and height of the chrominance block. Wherein, SubWidthC and SubHeightC represent variables corresponding to the color format. Wherein, when the color format is 4:2:0, SubWidthC and SubHeightC are 2, and Wherein, when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
6. The image encoding method according to claim 5, wherein, The specific position is the central position of the luminance block.
7. The image encoding method according to claim 5, wherein, When the prediction mode corresponding to the specific position is the IBC mode, the intra prediction mode of the chrominance block is updated to the intra DC mode.
8. The image encoding method according to claim 5, wherein, When the prediction mode corresponding to the specific position is the palette mode, the intra prediction mode of the chrominance block is updated to the intra DC mode.
9. A method for transmitting data for image information, the method comprising the following steps: Obtaining a bitstream for the image information, wherein the bitstream is generated based on the following steps: deriving prediction samples for the chrominance block based on that the intra prediction mode for the chrominance block is the cross-component linear model CCLM; deriving residual samples for the chrominance block based on the prediction samples; updating the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block; determining an LFNST set including an LFNST matrix based on the updated intra prediction mode; deriving modified transform coefficients for the chrominance block based on the residual samples and the LFNST matrix; and encoding residual information to generate the bitstream. Transmitting the data including the bitstream. Wherein, the updated intra prediction mode is derived as the intra prediction mode corresponding to a specific position in the luminance block. Wherein, based on that the intra prediction type corresponding to the specific position is the MIP mode, the intra prediction mode of the chrominance block is updated to the intra plane mode. Wherein, the specific position is set based on the color format of the chrominance block. Wherein, the specific position is set to ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)). Wherein, xTbY and yTbY represent the upper left coordinates of the luminance block. Wherein, nTbW and nTbH represent the width and height of the chrominance block. Wherein, SubWidthC and SubHeightC represent variables corresponding to the color format. Among them, when the color format is 4:2:0, SubWidthC and SubHeightC are 2, and Among them, when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.