Transform-based Image Coding Method and Apparatus
By using the CCLM mode and the intra-frame mode of the brightness block in the image encoding system, the LFNST transformation set is derived, which solves the problem of low compression efficiency of high resolution and high-quality image/video data, and achieves more efficient image/video compression and transmission.
Patent Information
- Application Number
- CN202080090680.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-29
- Filing Date
- 2020-10-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-10-29
AI Technical Summary
The prior art is difficult to effectively compress and transmit or store high resolution and high quality image/video data, especially when meeting the needs of ultra-high definition images/video, virtual reality and holograms.
By adopting the cross-component linear model (CCLM) mode and the intra-frame mode of the brightness block in the image encoding system, the LFNST transformation set is derived, thereby improving the efficiency of image encoding.
The image/video compression efficiency is improved, the efficiency of encoding LFNST indexes is increased, and the efficiency of secondary transformation is improved by encoding LFNST indexes.
Smart Images

Figure CN114902678B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image coding techniques, and more particularly, to methods and apparatuses for encoding an image based on transform coding in an image coding system. Background Art
[0002] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K, or higher ultra-high definition (UHD) images / videos has been continuously increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases compared to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, nowadays, the interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and the broadcasting of images / videos having image characteristics different from those of real images such as game images is increasing.
[0004] Therefore, there is a need for an efficient image / video compression technique for effectively compressing and transmitting or storing and reproducing information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention
[0005] Technical Objectives
[0006] One technical aspect of the present disclosure is to provide a method and an apparatus for increasing image coding efficiency.
[0007] Another technical aspect of the present disclosure is to provide a method and an apparatus for increasing the efficiency of encoding an LFNST index.
[0008] Still another technical aspect of the present disclosure is to provide a method and an apparatus for improving the efficiency of secondary transform by encoding an LFNST index.
[0009] Yet another technical aspect of the present disclosure is to provide an image coding method and an image coding apparatus for deriving an LFNST transform set using an intra mode of a luminance block in a CCLM mode.
[0010] Technical Solutions
[0011] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method may include the following steps: Based on the fact that the intra prediction mode of a chrominance block is a cross-component linear model (CCLM) mode, update the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block; Determine an LFNST set including an LFNST matrix based on the updated intra prediction mode; and Derive the transform coefficients of the chrominance block based on the LFNST matrix derived from the LFNST set, where the updated intra prediction mode is derived as the intra prediction mode corresponding to a specific position in the luminance block, and where, based on the fact that the intra prediction mode corresponding to the specific position is a MIP mode, update the updated intra prediction mode to an intra planar mode.
[0012] The specific position is set based on the color format of the chrominance block.
[0013] The specific position is the central position of the luminance block.
[0014] The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)), where xTbY and yTbY represent the upper left coordinates of the luminance block, nTbW and nTbH represent the width and height of the chrominance block, and SubWidthC and SubHeightC represent variables corresponding to the color format.
[0015] When the color format is 4:2:0, SubWidthC and SubHeightC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0016] When the intra prediction mode corresponding to the specific position is an IBC mode, the updated intra prediction mode is an intra DC mode.
[0017] When the intra prediction mode corresponding to the specific position is a palette mode, the updated intra prediction mode is an intra DC mode.
[0018] According to another embodiment of the present disclosure, an image encoding method performed by an encoding device is provided. The method may include the following steps: Based on the fact that the intra prediction mode of a chrominance block is a cross-component linear model (CCLM), derive the prediction samples of the chrominance block; Derive the residual samples of the chrominance block based on the prediction samples, the updated intra prediction mode is derived as the intra prediction mode corresponding to a specific position in the luminance block, and based on the fact that the intra prediction mode corresponding to the specific position is a MIP mode, update the updated intra prediction mode to an intra planar mode.
[0019] According to another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including a bitstream and encoded image information generated according to an image encoding method performed by an encoding device.
[0020] According to another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including encoded image information and a bitstream to cause a decoding device to perform an image decoding method.
[0021] Technical Effects
[0022] According to the present disclosure, the overall image / video compression efficiency can be increased.
[0023] According to the present disclosure, the efficiency of encoding LFNST indexes can be increased.
[0024] According to the present disclosure, the efficiency of secondary transformation can be increased by encoding LFNST indexes.
[0025] According to the present disclosure, an image encoding method and an image encoding device for deriving an LFNST transform set using an intra mode of a luminance block in a CCLM mode can be provided.
[0026] The effects obtainable through specific examples of the present disclosure are not limited to those listed above. For example, there may be various technical effects that can be understood by those of ordinary skill in the relevant art or derived from the present disclosure. Therefore, the specific effects of the present disclosure are not limited to those explicitly described in the present disclosure and may include various effects that can be understood or derived based on the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a diagram schematically illustrating the configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied.
[0028] Figure 2 is a diagram schematically illustrating the configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.
[0029] Figure 3 Schematically illustrates a multi-transformation technique according to an embodiment of the present disclosure.
[0030] Figure 4 Schematically shows an intra-directional mode with 65 prediction directions.
[0031] Figure 5 is a diagram illustrating an RST according to an embodiment of the present disclosure.
[0032] Figure 6 is a diagram illustrating the order of arranging output data of a forward primary transformation into a one-dimensional vector according to an example.
[0033] Figure 7 It is a diagram illustrating the order of arranging the output data of the forward quadratic transform into a two-dimensional block according to the example.
[0034] Figure 8 It is a diagram illustrating the block shape to which the LFNST is applied.
[0035] Figure 9 It is a diagram illustrating the arrangement of the output data of the forward LFNST according to the example and shows the block in which the output data of the forward LFNST is arranged according to the example.
[0036] Figure 10 It shows a diagram in which the number of output data of the forward LFNST according to the example is limited to a maximum value of 16.
[0037] Figure 11 It is a diagram illustrating the zeroing in the block to which the 4×4 LFNST is applied according to the example.
[0038] Figure 12 It is a diagram illustrating the zeroing in the block to which the 8×8 LFNST is applied according to the example.
[0039] Figure 13 It is a diagram illustrating the MIP-based predicted sample generation process according to the example.
[0040] Figure 14 It is a diagram illustrating the CCLM applicable when deriving the intra prediction mode of the chrominance block according to the embodiment.
[0041] Figure 15 It is a diagram for explaining the method of decoding an image according to the example.
[0042] Figure 16 It is a diagram for explaining the method of encoding an image according to the example.
[0043] Figure 17 It schematically illustrates an example of a video / image coding system to which the present disclosure can be applied.
[0044] Figure 18 It illustrates the structure of a content stream system to which the present disclosure is applied. Detailed Description of the Invention
[0045] Although the present disclosure may be susceptible to various modifications and include various embodiments, specific embodiments thereof have been shown by way of example in the drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the technical concept of the present disclosure. Unless the context clearly indicates otherwise, the singular forms may include the plural forms. Terms such as "including" and "having" are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus should not be construed as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0046] In addition, for the convenience of describing different characteristic functions, each component in the drawings described herein is illustrated independently. However, it is not meant that each component is implemented by separate hardware or software. For example, any two or more of these components can be combined to form a single component, and any single component can be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent right of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0047] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the drawings. In addition, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.
[0048] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).
[0049] In this document, various embodiments related to video / image coding may be provided, and unless otherwise specified, these embodiments may be combined with each other and executed.
[0050] In this document, video may refer to a collection of a series of images over a period of time. Generally, a picture refers to a unit representing an image of a specific time region, and a slice / tile is a unit that forms part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / titles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.
[0051] A pixel or pel may refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can mean a pixel value in the spatial domain, or when the pixel value is transformed into the frequency domain, it can mean a transform coefficient in the frequency domain.
[0052] A unit can represent the basic unit of image processing. A unit can include at least one of a specific region and information related to that region. A unit can include one luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the situation, terms such as unit and terms like block, region, etc. can be used interchangeably. Generally, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients composed of M columns and N rows.
[0053] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Additionally, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A, B, C" can mean "at least one of A, B, and / or C".
[0054] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0055] In this disclosure, "at least one of A and B" can mean "only A", "only B", or "both A and B". Additionally, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0056] Furthermore, in this disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0057] In addition, the parentheses used in the present disclosure may indicate "for example". Specifically, when it is indicated as "prediction (intra prediction)", it may mean that "intra prediction" is presented as an example of "prediction". In other words, "prediction" in the present disclosure is not limited to "intra prediction", and "intra prediction" is presented as an example of "prediction". In addition, when it is indicated as "prediction (i.e., intra prediction)", this may also mean that "intra prediction" is presented as an example of "prediction".
[0058] The technical features separately described in one drawing in the present disclosure may be implemented separately or may be implemented simultaneously.
[0059] Figure 1 FIG. is a diagram schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.
[0060] Referring to Figure 1 , the encoding device 100 may include and be configured with an image splitter 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 may include an inter predictor 121 and an intra predictor 122. The residual processor 130 may include a transformer 132, a quantizer 133, a dequantizer 134, and an inverse transformer 135. The residual processor 130 may further include a subtractor 131. The adder 150 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the above-described image splitter 110, predictor 120, residual processor 130, entropy encoder 140, adder 150, and filter 160 may be constituted by one or more hardware components (e.g., an encoder chipset or a processor). In addition, the memory 170 may include a decoded picture buffer (DPB) and may also be constituted by a digital storage medium. The hardware components may further include the memory 170 as an internal / external component.
[0061] The image splitter 110 may split an input image (or picture, frame) input to the encoding device 100 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, one coding unit may be split into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, and later the binary tree structure and / or the ternary tree structure may be applied. Alternatively, the binary tree structure may also be applied first. The encoding process according to this document may be performed based on the final coding unit that is no longer split. In this case, based on the encoding efficiency according to the image characteristics, etc., the LCU may be directly used as the final coding unit, or alternatively, the coding unit may be recursively split into coding units of a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction described later. As another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, each of the PU and TU may be split or divided from the aforementioned final coding unit. The PU may be a unit for sample prediction, and the TU may be a unit for inducing transform coefficients and / or a unit for inducing a residual signal from the transform coefficients.
[0062] In some cases, the unit may be used interchangeably with terms such as block or region. Generally, an M×N block may represent samples or a set of transform coefficients composed of M columns and N rows. Samples generally may represent pixels or pixel values, and may also represent only the pixels / pixel values of the luminance component, and may also represent only the pixels / pixel values of the chrominance component. Samples may be used as items corresponding to the pixels or primitives configuring a picture (or image).
[0063] The encoding device 100 can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 121 or the intra-frame predictor 122 from an input image signal (original block, original sample array), and send the generated residual signal to the transformer 132. In this case, as illustrated, the unit in the encoding device 100 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as the subtractor 131. The predictor can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The predictor can generate various information about the prediction, such as prediction mode information, to transmit the generated information to the entropy encoder 140, as described later in the description of each prediction mode. The information about the prediction can be encoded by the entropy encoder 140 and output in the form of a bitstream.
[0064] The intra-frame predictor 122 can predict the current block by referring to the samples within the current picture. Depending on the prediction mode, the samples referred to may be adjacent to the current block or may also be located far from the current block. The prediction modes in intra-frame prediction can include multiple non-directional modes and multiple directional modes. The non-directional modes can include, for example, the DC mode or the planar mode. Depending on the fineness of the prediction direction, the directional modes can include (for example) 33 directional prediction modes or 65 directional prediction modes. However, this is illustrative, and more or fewer directional prediction modes than the above numbers can be used according to the settings. The intra-frame predictor 122 can also use the prediction mode applied to an adjacent block to determine the prediction mode applied to the current block.
[0065] The inter - frame predictor 121 can induce a predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, adjacent blocks can include spatially adjacent blocks within the current picture and temporally adjacent blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally adjacent block can be the same as each other and can also be different from each other. The temporally adjacent block can be called by names such as, for example, a collocated reference block, a collocated CU (col CU), or the like, and the reference picture including the temporally adjacent block can also be called by names such as a collocated reference block, a collocated CU (col CU), etc., and the reference picture including the temporally neighboring block can be called a collocated picture (colPic). For example, the inter - frame predictor 121 can configure a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. The inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame predictor 121 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. The motion vector prediction (MVP) mode can indicate the motion vector of the current block by using the motion vector of an adjacent block as a motion vector prediction and by signaling the motion vector difference.
[0066] The predictor 120 can generate a prediction signal based on various prediction methods to be described later. For example, the predictor can not only apply intra - frame prediction or inter - frame prediction to predict a block, but also apply intra - frame prediction and inter - frame prediction simultaneously. This can be called combined intra - and inter - frame prediction (CIIP). In addition, the predictor can perform prediction on a block based on the intra - block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but it can be similarly performed for inter - frame prediction because it derives a reference block in the current picture. That is, IBC can use at least one of the inter - frame prediction techniques described in this document. The palette mode can be regarded as an example of intra - frame coding or intra - frame prediction. When the palette mode is applied, the sample values in the picture can be signaled based on information about the palette index and the palette table.
[0067] The prediction signal generated by a predictor (including the inter-frame predictor 121 and / or the intra-frame predictor 122) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 132 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen–Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, the GBT represents a transform obtained from the graph. The CNT represents a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. Additionally, the transform process can also be applied to a pixel block of a square with the same size, and can also be applied to a block with a variable size rather than a square.
[0068] Quantizer 133 may quantize the transform coefficients to send the quantized transform coefficients to entropy encoder 140, and entropy encoder 140 may encode the quantized signal (information about the quantized transform coefficients) to output the encoded quantized signal in the form of a bitstream. The information about the quantized transform coefficients may be referred to as residual information. Quantizer 133 may rearrange the quantized transform coefficients in block form into a one-dimensional vector based on the coefficient scan order, and may also generate information about the quantized transform coefficients based on the quantized transform coefficients in one-dimensional vector form. Entropy encoder 140 may perform various coding methods, for example, such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 140 may also encode the information required for video / image reconstruction (e.g., values of syntax elements, etc.) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be sent or stored in units of network abstraction layer (NAL) in the form of a bitstream. The video / image information may also include information about various parameter sets such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). In addition, the video / image information may also include general constraint information. In this document, the syntax elements and / or information transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the above encoding process and included in the bitstream. The bitstream may be sent through a network or may be stored in a digital storage medium. In this document, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for sending the signal output from entropy encoder 140 and / or a storage device (not shown) for storing the signal may be configured as internal / external elements of encoding device 100, or the transmitter may also be included in entropy encoder 140.
[0069] The quantized transform coefficients output from the quantizer 133 can be used to generate a prediction signal. For example, the dequantizer 134 and the inverse transformer 135 can apply dequantization and inverse transformation to the quantized transform coefficients in order to recover the residual signal (residual block or residual samples). The adder 150 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 121 or the intra-frame predictor 122 in order to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). As in the case of applying the skip mode, if there is no residual in the block to be processed, the predicted block can be used as the reconstructed block. The adder 150 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current picture and, as described below, can also be used for inter-frame prediction of the next picture through filtering.
[0070] In addition, luminance mapping and chrominance scaling (LMCS) can also be applied during the picture encoding and / or reconstruction process.
[0071] The filter 160 can apply filtering to the reconstructed signal, thereby improving the subjective / objective image quality. For example, the filter 160 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 160 can generate various filtering-related information to transmit the generated information to the entropy encoder 140, as described later in the description of each filtering method. The filtering-related information can be encoded by the entropy encoder 140 and can be output in the form of a bitstream.
[0072] The modified reconstructed picture sent to the memory 170 can be used as a reference picture in the inter-frame predictor 121. If inter-frame prediction is applied by the inter-frame predictor, the encoding device can avoid prediction mismatch between the encoding device 100 and the decoding device and can also improve the encoding efficiency.
[0073] The DPB of the memory 170 can store the modified reconstructed picture to be used as a reference picture in the inter-frame predictor 121. The memory 170 can store the motion information of the block in which the motion information within the current picture is derived (or encoded) and / or the motion information of the block within the previously reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 121 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 170 can store the reconstructed samples of the reconstructed blocks within the current picture and transmit the reconstructed samples to the intra-frame predictor 122.
[0074] Figure 2It is a diagram schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.
[0075] Referring to Figure 2 , the decoding device 200 may include and be configured with an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The predictor 230 may include an inter-predictor 232 and an intra-predictor 231. The residual processor 220 may include a dequantizer 221 and an inverse transformer 222. According to an embodiment, the entropy decoder 210, the residual processor 220, the predictor 230, the adder 240, and the filter 250 described above may be configured by one or more hardware components (e.g., a decoder chipset or a processor). In addition, the memory 260 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 260 as an internal / external component.
[0076] When receiving a bitstream including video / image information, the decoding device 200 may reconstruct an image in response to the process of processing the video / image information in the Figure 1 encoding device shown. For example, the decoding device 200 may derive units / blocks based on block segmentation-related information obtained from the bitstream. The decoding device 200 may perform decoding using the processing units applied to the encoding device. Thus, the processing units for decoding may be, for example, encoding units, and the encoding units may be segmented from a CTU or an LCU according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the encoding units. In addition, the reconstructed image signal decoded and output by the decoding device 200 may be reproduced by a reproduction device.
[0077] The decoding device 200 may receive, in the form of a bitstream, from Figure 1The signal output by the encoding device as shown, and the received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction) by parsing the bitstream. The video / image information can also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information can also include general constraint information. The decoding device can also decode pictures based on the information about the parameter sets and / or the general constraint information. The information and / or syntax elements to be signaled / received (which will be described later in this document) can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 210 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements required for image reconstruction, as well as the quantization values of the residual-related transform coefficients. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the syntax element information to be decoded and the decoding information of adjacent blocks and the block to be decoded or the information of the symbols / bins decoded in the previous step to determine the context model, and perform arithmetic decoding on the bins by predicting the bin generation probability according to the determined context model to generate symbols corresponding to the values of each syntax element. At this time, the CABAC entropy decoding method can determine the context model and then update the context model using the information of the symbols / bins decoded for the decoding of the context model of the next symbol / bin. The information about prediction among the information decoded by the entropy decoder 210 can be provided to the predictors (inter-frame predictor 232 and intra-frame predictor 231), and the residual values obtained by the entropy decoder 210 performing entropy decoding, that is, the quantization transform coefficients and the related parameter information, can be input to the residual processor 220. The residual processor 220 can derive the residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded by the entropy decoder 210 can be provided to the filter 250. In addition, the receiver (not shown) for receiving the signal output by the encoding device can also be configured as an internal / external component of the decoding device 200, or the receiver can also be a component of the entropy decoder 210. In addition, the decoding device according to the present disclosure can be referred to as a video / image / picture decoding device, and the decoding device can also be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210, and the sample decoder can include at least one of a dequantizer 221, an inverse transformer 222, an adder 240, a filter 250, a memory 260, an inter-frame predictor 232, and an intra-frame predictor 231.
[0078] The dequantizer 221 may dequantize the quantized transform coefficients to output the transform coefficients. The dequantizer 221 may rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order executed by the encoding device. The dequantizer 221 may perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0079] The inverse transformer 222 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0080] The predictor 230 may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 210 and determine a specific intra / inter prediction mode.
[0081] The predictor may generate a prediction signal based on various prediction methods to be described later. For example, the predictor may apply not only intra prediction or inter prediction to the prediction of a block but also apply intra prediction and inter prediction simultaneously. This may be referred to as combined inter and intra prediction (CIIP). In addition, the predictor may perform prediction on the block based on the intra block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode may be used for content image / video coding such as games, e.g., screen content coding (SCC). IBC basically performs prediction in the current picture, but it may perform similarly for inter prediction because it derives a reference block in the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and the palette index may be signaled by being included in the video / image information.
[0082] The intra predictor 231 may predict the current block by referring to samples within the current picture. Depending on the prediction mode, the samples referred to may be adjacent to the current block or may also be located far from the current block. The prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 231 may also use the prediction mode applied to an adjacent block to determine the prediction mode applied to the current block.
[0083] The inter-frame predictor 232 can induce a predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, adjacent blocks can include spatially adjacent blocks existing within the current picture and temporally adjacent blocks existing in the reference picture. For example, the inter-frame predictor 232 can configure a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter-frame prediction of the current block.
[0084] The adder 240 can add the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a predictor (including the inter-frame predictor 232 and / or the intra-frame predictor 231) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). As in the case of applying the skip mode, if there is no residual for the block to be processed, the predicted block can be used as the reconstructed block.
[0085] The adder 240 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current picture, and as described below, can also be output through filtering or can also be used for inter-frame prediction of the next picture.
[0086] In addition, luminance mapping and chrominance scaling (LMCS) can also be applied during picture decoding.
[0087] The filter 250 can apply filtering to the reconstructed signal, thereby improving the subjective / objective image quality. For example, the filter 250 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and send the modified reconstructed picture to the memory 260, specifically, the DPB of the memory 260. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0088] The (modified) reconstructed picture stored in the DPB of the memory 260 can be used as a reference picture in the inter - predictor 232. The memory 260 can store the motion information of the blocks in which the motion information within the current picture is derived (decoded) and / or the motion information of the blocks within the previous reconstructed pictures. The stored motion information can be transmitted to the inter - predictor 232 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 260 can store the reconstructed samples of the reconstructed blocks within the current picture and transmit the stored reconstructed samples to the intra - predictor 231.
[0089] In this document, the exemplary embodiments described in the filter 160, inter - predictor 121, and intra - predictor 122 of the encoding device 100 can be equally applied to the filter 250, inter - predictor 232, and intra - predictor 231 of the decoding device 200, respectively.
[0090] As described above, prediction is performed to improve the compression efficiency when performing video coding. Accordingly, a prediction block including prediction samples for the current block that is the target block to be encoded can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve the image coding efficiency by signaling to the decoding device not the original sample values of the original block itself but the information about the residual between the original block and the prediction block (residual information). The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.
[0091] The residual information can be generated through a transformation process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients, so that it can signal the associated residual information to the decoding device (through the bitstream). Here, the residual information can include value information, position information, transformation technique, transformation kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / de - quantization process based on the residual information and derive residual samples (or a residual sample block). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can de - quantize / inverse - transform the quantized transform coefficients to derive a residual block for use as a reference for inter - prediction of the next picture and can generate a reconstructed picture based on this.
[0092] Figure 3 An example of a multi - transformation technique according to an embodiment of the present disclosure is schematically illustrated.
[0093] Refer toFigure 3 , the transformer may correspond to the transformer in the encoding device described above Figure 1 , and the inverse transformer may correspond to the inverse transformer in the encoding device described above Figure 1 , or the inverse transformer in the decoding device described above. Figure 2
[0094] The transformer may derive (primary) transform coefficients (S310) by performing a single transformation based on residual samples (residual sample array) in the residual block. This single transformation may be referred to as the core transformation. In this document, the single transformation may be based on multi-transformation selection (MTS), and when multi-transformation is used as the single transformation, it may be referred to as multi-core transformation.
[0095] The multi-core transformation may represent a method of additionally performing transformations using Discrete Cosine Transform (DCT) type 2 and Discrete Sine Transform (DST) type 7, DCT type 8, and / or DST type 1. That is, the multi-core transformation may represent a transformation method of transforming a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this document, from the perspective of the transformer, the primary transform coefficients may be referred to as temporary transform coefficients.
[0096] In other words, when applying a conventional transformation method, transform coefficients can be generated by applying a transformation from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, when applying multi-core transformation, transform coefficients (or primary transform coefficients) can be generated by applying a transformation from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this document, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transformation types, transform kernels, or transform cores. These DCT / DST transformation types can be defined based on basis functions.
[0097] When performing multi-core transformation, a vertical transform kernel and a horizontal transform kernel for the target block can be selected from among the transform kernels, the vertical transform can be performed on the target block based on the vertical transform kernel, and the horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform may indicate the transformation of the horizontal component of the target block, and the vertical transform may indicate the transformation of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block) including the residual block.
[0098] In addition, according to the example, if a transformation is performed by applying MTS, the mapping relationship of the transformation kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transformation or the horizontal transformation. For example, when the horizontal transformation kernel is represented as trTypeHor and the vertical transformation kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT8.
[0099] In this case, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transformation kernel sets. For example, MTS index 0 can indicate that both the trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both the trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that the trTypeHor value is 2 and the trTypeVer value is 1, MTS index 3 can indicate that the trTypeHor value is 1 and the trTypeVer value is 2, and MTS index 4 can indicate that both the trTypeHor and trTypeVer values are 2.
[0100] In one example, the transformation kernel sets according to the MTS index information are shown in the following table.
[0101] [Table 1]
[0102] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0103] The transformer may perform a secondary transformation based on the (primary) transformation coefficients to derive modified (secondary) transformation coefficients (S320). The primary transformation is a transformation from the spatial domain to the frequency domain, and the secondary transformation refers to a transformation into a more compact representation using the correlations existing between the (primary) transformation coefficients. The secondary transformation may include an inseparable transformation. In this case, the secondary transformation may be referred to as an inseparable secondary transformation (NSST) or a mode-dependent inseparable secondary transformation (MDNSST). The NSST may represent a transformation of performing a secondary transformation on the (primary) transformation coefficients derived through the primary transformation based on an inseparable transformation matrix to generate modified transformation coefficients (or secondary transformation coefficients) for the residual signal. Here, based on the inseparable transformation matrix, the transformation may be applied once to the (primary) transformation coefficients without separating the vertical transformation and the horizontal transformation (or applying the horizontal / vertical transformation independently). In other words, the NSST does not separately apply to the (primary) transformation coefficients in the vertical and horizontal directions, and may represent, for example, a transformation method of rearranging a two-dimensional signal (transformation coefficients) into a one-dimensional signal in a specific predetermined direction (e.g., row-major order or column-major order) and then generating modified transformation coefficients (or secondary transformation coefficients) based on the inseparable transformation matrix. For example, the row-major order is set in rows in the order of the first row, the second row, …, and the Nth row for M×N blocks, and the column-major order is set in rows in the order of the first column, the second column, …, and the Mth column for M×N blocks. The NSST may be applied to the upper left region of a block configured with (primary) transformation coefficients (hereinafter referred to as a transformation coefficient block). For example, when both the width W and the height H of the transformation coefficient block are 8 or greater, an 8×8 NSST may be applied to the upper left 8×8 region of the transformation coefficient block. In addition, while both the width (W) and the height (H) of the transformation coefficient block are 4 or greater, when the width (W) or the height (H) of the transformation coefficient block is less than 8, a 4×4 NSST may be applied to the upper left min(8, W)×min(8, H) region of the transformation coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that the width W or the height H of the transformation coefficient block is 4 or greater is satisfied, a 4×4 NSST may be applied to the upper left end min(8, W)×min(8, H) region of the transformation coefficient block.
[0104] Specifically, for example, if a 4×4 input block is used, the inseparable secondary transformation may be performed as follows.
[0105] The 4×4 input block X may be represented as follows.
[0106] [Equation 1]
[0107]
[0108] If X is represented in the form of a vector, the vector It can be expressed as follows.
[0109] [Equation 2]
[0110]
[0111] In Equation 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to row-major order.
[0112] In this case, the inseparable quadratic transform can be calculated as follows.
[0113] [Equation 3]
[0114]
[0115] In this equation, represents the transform coefficient vector, and T represents a 16×16 (inseparable) transform matrix.
[0116] Through the aforementioned Equation 3, the 16×1 transform coefficient vector can be derived, and the vector can be reorganized into 4×4 blocks in scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is an example, and the hypercube-Givens transform (HyGT), etc. can also be used for the calculation of the inseparable quadratic transform to reduce the computational complexity of the inseparable quadratic transform.
[0117] In addition, in the inseparable quadratic transform, the transform kernel (or transform core, transform type) can be selected to be mode-dependent. In this case, the mode can include the intra prediction mode and / or the inter prediction mode.
[0118] As described above, the inseparable quadratic transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 region included in the transform coefficient block when both W and H are equal to or greater than 8, and the 8×8 region can be the upper left 8×8 region in the transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 region included in the transform coefficient block when both W and H are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0119] Here, in order to select a transform kernel related to a mode, two non-separable quadratic transform kernels for each transform set for non-separable quadratic transforms can be configured for both the 8×8 transform and the 4×4 transform, and there can be four transform sets. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.
[0120] However, as the size of the transform (i.e., the size of the region to which the transform is applied) can be a size other than, for example, 8×8 or 4×4, the number of sets can be n, and the number of transform kernels in each set can be k.
[0121] The transform sets can be referred to as NSST sets or LFNST sets. A specific set among the transform sets can be selected, for example, based on the intra prediction mode of the current block (CU or sub-block). The low-frequency non-separable transform (LFNST) can be an example of a reduced non-separable transform, which will be described later, and represents a non-separable transform for low-frequency components.
[0122] As a reference, for example, the intra prediction mode can include two non-directional (or non-angle) intra prediction modes and 65 directional (or angle) intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode of No. 0 and the DC intra prediction mode of No. 1, and the directional intra prediction modes can include 65 intra prediction modes from No. 2 to No. 66. However, this is an example, and this document can be applied even if the number of intra prediction modes is different. In addition, in some cases, the intra prediction mode of No. 67 can also be used, and the intra prediction mode of No. 67 can represent a linear model (LM) mode.
[0123] Figure 4 The intra-frame directional mode with 65 prediction directions is schematically shown.
[0124] Referring to Figure 4 , based on the intra prediction mode 34 with the upper left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode with horizontal directivity and an intra prediction mode with vertical directivity. In Figure 4In this case, H and V respectively indicate horizontal directionality and vertical directionality, and the numbers -32 to 32 indicate displacements of 1 / 32 units at the sample grid positions. These numbers can represent offsets for the mode index values. Intra prediction modes 2 to 33 have horizontal directionality, and intra prediction modes 34 to 66 have vertical directionality. Strictly speaking, intra prediction mode 34 can be regarded as neither horizontal nor vertical, but can be classified as belonging to horizontal directionality when determining the transform set of the secondary transform. This is because the input data is transposed for the vertically oriented mode symmetric to intra prediction mode 34, and the input data alignment method for the horizontal mode is used for intra prediction mode 34. Transposing the input data means switching the rows and columns of the two-dimensional M×N block data to N×M data. Intra prediction mode 18 and intra prediction mode 50 can respectively represent the horizontal intra prediction mode and the vertical intra prediction mode, and intra prediction mode 2 can be called the upper right diagonal intra prediction mode because intra prediction mode 2 has a left reference pixel and performs prediction in the upper right direction. Similarly, intra prediction mode 34 can be called the lower right diagonal intra prediction mode, and intra prediction mode 66 can be called the lower left diagonal intra prediction mode.
[0125] According to an example, four transform sets can be mapped according to the intra prediction mode, for example, as shown in the following table.
[0126] [Table 2]
[0127] lfnstPredModeIntra lfnstTrSetIdx lfnstPredModeIntra < 0 1 0 <= lfnstPredModeIntra <= 1 0 2 <= lfnstPredModeIntra <= 12 1 13 <= lfnstPredModeIntra <= 23 2 24 <= lfnstPredModeIntra <= 44 3 45 <= lfnstPredModeIntra <= 55 2 56 <= lfnstPredModeIntra <= 80 1 81 <= lfnstPredModeIntra <= 83 0
[0128] As shown in Table 2, any one of the four transform sets, that is, lfnstTrSetIdx, can be mapped to any one of the four indices (i.e., 0 to 3) according to the intra prediction mode.
[0129] When determining a specific set for the non-separable transform, one of the k transform kernels in the specific set can be selected by the non-separable secondary transform index. The encoding device can derive the non-separable secondary transform index indicating the specific transform kernel based on rate distortion (RD) checking, and can signal the non-separable secondary transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable secondary transform index. For example, the lfnst index value 0 can refer to the first non-separable secondary transform kernel, the lfnst index value 1 can refer to the second non-separable secondary transform kernel, and the lfnst index value 2 can refer to the third non-separable secondary transform kernel. Alternatively, the lfnst index value 0 can indicate that the first non-separable secondary transform is not applied to the target block, and the lfnst index values 1 to 3 can indicate three transform kernels.
[0130] The transformer can perform a non-separable quadratic transform based on the selected transform kernel and can obtain modified (quadratic) transform coefficients. As described above, the modified transform coefficients can be derived as the transform coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0131] In addition, as described above, if the quadratic transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as the transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0132] The inverse transformer can perform a series of processes in an order opposite to the order already executed in the above-mentioned transformer. The inverse transformer can receive the (dequantized) transform coefficients, and derive the (primary) transform coefficients by performing a quadratic (inverse) transform (S350), and can obtain the residual block (residual samples) by performing a primary (inverse) transform on the (primary) transform coefficients (S360). In this regard, from the perspective of the inverse transformer, the primary transform coefficients can be referred to as modified transform coefficients. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0133] The decoding device may further include a quadratic inverse transform application determiner (or an element for determining whether to apply the quadratic inverse transform) and a quadratic inverse transform determiner (or an element for determining the quadratic inverse transform). The quadratic inverse transform application determiner can determine whether to apply the quadratic inverse transform. For example, the quadratic inverse transform can be NSST, RST, or LFNST, and the quadratic inverse transform application determiner can determine whether to apply the quadratic inverse transform based on the quadratic transform flag obtained by parsing the bitstream. In another example, the quadratic inverse transform application determiner can determine whether to apply the quadratic inverse transform based on the transform coefficients of the residual block.
[0134] The quadratic inverse transform determiner can determine the quadratic inverse transform. In this case, the quadratic inverse transform determiner can determine the quadratic inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra prediction mode. In an embodiment, the quadratic transform determination method can depend on the primary transform determination method. Various combinations of the primary transform and the quadratic transform can be determined according to the intra prediction mode. In addition, in an example, the quadratic inverse transform determiner can determine the region to which the quadratic inverse transform is applied based on the size of the current block.
[0135] In addition, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, a single (separable) inverse transform can be performed, and a residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0136] In addition, in the present disclosure, a reduced secondary transform (RST) in which the size of the transform matrix (kernel) is reduced can be applied in the concept of NSST, so as to reduce the amount of computation and storage required for the non-separable secondary transform.
[0137] In addition, the transform kernel, the transform matrix, and the coefficients constituting the transform kernel matrix described in the present disclosure, that is, the kernel coefficients or the matrix coefficients, can be represented by 8 bits. This can be a condition implemented in the decoding device and the encoding device, and compared with the existing 9 bits or 10 bits, it can reduce the storage amount required for storing the transform kernel, and can reasonably adapt to the performance degradation. In addition, representing the kernel matrix by 8 bits can allow the use of a small multiplier, and can be more suitable for the single instruction multiple data (SIMD) instruction for optimal software implementation.
[0138] In this specification, the term "RST" may refer to a transform performed on the residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. In the case of performing a reduction transform, due to the reduction in the size of the transform matrix, the amount of computation required for the transform can be reduced. That is, RST can be used to solve the computational complexity problem that occurs in the transform of a large-sized block or a non-separable transform.
[0139] RST can be referred to by various terms such as reduced transform, reduced secondary transform, scaled transform, simplified transform, and simple transform, and the names that RST can be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in the low-frequency region including non-zero coefficients in the transform block, it can be referred to as a low-frequency non-separable transform (LFNST). The transform index can be referred to as the LFNST index.
[0140] In addition, when performing a secondary inverse transform based on RST, the inverse transformers 135 of the encoding device 100 and the inverse transformers 222 of the decoding device 200 can include: an inverse reduced secondary transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients; and an inverse single transformer that derives the residual samples of the target block based on the inverse single transform of the modified transform coefficients. The inverse single transform refers to the inverse transform of the single transform applied to the residual. In the present disclosure, deriving the transform coefficients based on the transform may refer to deriving the transform coefficients by applying the transform.
[0141] Figure 5FIG. 0 is a diagram illustrating an RST according to an embodiment of the present disclosure.
[0142] In the present disclosure, a "target block" may refer to a current block to be encoded, a residual block, or a transform block.
[0143] In an RST according to an example, an N-dimensional vector may be mapped to an R-dimensional vector located in another space, thereby determining a reduction transform matrix, where R is less than N. N may refer to the square of the length of a side of a block to which a transform is applied, or the total number of transform coefficients corresponding to a block to which a transform is applied, and a reduction factor may refer to the R / N value. The reduction factor may be referred to as a reduction factor, a shrinking factor, a simplification factor, a simplifying factor, or various other terms. Further, R may be referred to as a reduction coefficient, but depending on the situation, the reduction factor may refer to R. Further, depending on the situation, the reduction factor may refer to the N / R value.
[0144] In an example, the reduction factor or reduction coefficient may be signaled via a bitstream, but the example is not limited thereto. For example, a predetermined value for the reduction factor or reduction coefficient may be stored in each of the encoding device 100 and the decoding device 200, and in this case, the reduction factor or reduction coefficient may not be signaled separately.
[0145] The size of the reduction transform matrix according to an example may be R×N, which is less than N×N (the size of a conventional transform matrix), and may be defined as in Equation 4 below.
[0146] [Equation 4]
[0147]
[0148] Figure 5 The matrix T in the reduction transform block shown in (a) of may refer to the matrix T of Equation 4 R×N . As Figure 5 shown in (a) of, when the reduction transform matrix T R×N is multiplied by the residual samples of the target block, the transform coefficients of the current block may be derived.
[0149] In an example, if the size of a block to which a transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST according to Figure 5 (a) of may be expressed as a matrix operation shown in Equation 5 below. In this case, the storage and multiplication calculations may be reduced by the reduction factor to approximately 1 / 4.
[0150] In the present disclosure, a matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix set on the left side of the column vector.
[0151] [Equation 5]
[0152]
[0153] In Equation 5, r1 to r 64 can represent the residual samples of the target block, and specifically can be the transform coefficients generated by applying a single transform. As a result of the calculation of Equation 5, the transform coefficients c of the target block can be derived i , and the process of deriving c i can be as shown in Equation 6.
[0154] [Equation 6]
[0155]
[0156] As a result of the calculation of Equation 6, the transform coefficients c1 to c of the target block can be derived R . That is, when R = 16, the transform coefficients c1 to c of the target block can be derived 16 . If a conventional transform is applied instead of RST, and a transform matrix of size 64×64 (N×N) is multiplied by a residual sample of size 64×1 (N×1), then only 16 (R) transform coefficients are derived for the target block because RST is applied, although 64 (N) transform coefficients are derived for the target block. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data sent from the encoding device 100 to the decoding device 200 is reduced, and thus the transmission efficiency between the encoding device 100 and the decoding device 200 can be improved.
[0157] When considered from the perspective of the size of the transform matrix, the size of the conventional transform matrix is 64×64 (N×N), but the size of the reduced transform matrix is reduced to 16×64 (R×N). Therefore, compared with the case of performing a conventional transform, the storage utilization rate in the case of performing RST can reduce the R / N ratio. In addition, when compared with the number of multiplication calculations N×N in the case of using a conventional transform matrix, using the reduced transform matrix can reduce the number of multiplication calculations (R×N) by the R / N ratio.
[0158] In the example, the transformer 132 of the encoding device 100 can derive the transform coefficients of the target block by performing a single transform on the residual samples of the target block and a secondary transform based on RST. These transform coefficients can be transmitted to the inverse transformer of the decoding device 200, and the inverse transformer 222 of the decoding device 200 can derive the modified transform coefficients based on the inverse reduced secondary transform (RST) for the transform coefficients, and can derive the residual samples of the target block based on the inverse single transform for the modified transform coefficients.
[0159] According to the inverse RST matrix T of the example N×Ris N×R, which is smaller than the size of the conventional inverse transform matrix N×N, and is related to the reduced transform matrix T shown in Equation 4 R×N has a transpose relationship.
[0160] Figure 5 The matrix T in the reduced inverse transform block shown in (b) of t may refer to the inverse RST matrix T N×R T (the superscript T refers to the transpose). As Figure 5 shown in (b) of N×R T When the inverse RST matrix T R×N T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. The inverse RST matrix T R×N can be expressed as (T T N×R ).
[0161] More specifically, when the inverse RST is used as a secondary inverse transform, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the residual samples of the target block can be derived.
[0162] In the example, if the size of the block to which the inverse transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST in accordance with Figure 5 (b) can be expressed as the matrix operation shown in Equation 7 below.
[0163] [Equation 7]
[0164]
[0165] In Equation 7, c1 to c 16 can represent the transform coefficients of the target block. As a result of the calculation of Equation 7, r j representing the modified transform coefficients of the target block or the residual samples of the target block can be derived, and the process of deriving r j can be as shown in Equation 8.
[0166] [Equation 8]
[0167]
[0168] As a result of the calculation of Equation 8, r1 to r representing the modified transform coefficients of the target block or the residual samples of the target block can be derived N From the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduction transform matrix is reduced to 64×16 (R×N). Therefore, compared with the case of performing a conventional inverse transform, the storage utilization rate in the case of performing an inverse RST can be reduced by the ratio of R / N. In addition, when compared with the number of multiplication calculations N×N in the case of using a conventional inverse transform matrix, using the inverse reduction transform matrix can reduce the number of multiplication calculations (N×R) by the ratio of R / N.
[0169] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied according to the transform set in Table 2. Since, according to the intra prediction mode, one transform set includes two or three transforms (kernels), it can be configured to select one of up to four transforms including the case where no secondary transform is applied. Among the transforms without applying the secondary transform, applying an identity matrix can be considered. Assuming that indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case of applying the identity matrix, that is, the case where no secondary transform is applied), a transform index or an lfnst index can be signaled as a syntax element for each transform coefficient block, thereby specifying the transform to be applied. That is, for the upper left 8×8 block, through the transform index, 8×8 NSST in the RST configuration can be specified, or 8×8 lfnst can be specified when LFNST is applied. 8×8 lfnst and 8×8 RST refer to the transforms that can be applied to the 8×8 region included in the transform coefficient block when both the W and H of the target block to be transformed are equal to or greater than 8, and the 8×8 region can be the upper left 8×8 region in the transform coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to the transforms that can be applied to the 4×4 region included in the transform coefficient block when both the W and H of the target block are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transform coefficient block.
[0170] According to an embodiment of the present disclosure, for the transform in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 region. Here, "maximum" means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is to say, when performing RST by applying an m×48 transform kernel matrix (m≤16) to an 8×8 region, 48 pieces of data are input, and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is to say, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming the 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on 48 pieces of data constituting the region except for the lower right 4×4 region in the 8×8 region. Here, when performing matrix operations by applying a maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper left 4×4 region according to the scanning order, and the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.
[0171] For the inverse transform in the decoding process, the transpose matrix of the aforementioned transform kernel matrix can be used. That is to say, when performing inverse RST or LFNST in the inverse transform process executed by the decoding device, the input coefficient data for applying inverse RST is configured in a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector can be arranged in a two-dimensional block according to a predetermined arrangement order.
[0172] In summary, in the transform process, when RST or LFNST is applied to an 8×8 region, matrix operations are performed on 48 transform coefficients in the upper left region, upper right region, and lower left region of the 8×8 region except for the lower right region with a 16×48 transform kernel matrix. For matrix operations, 48 transform coefficients are input in a one-dimensional array. When performing matrix operations, 16 modified transform coefficients are derived, and the modified transform coefficients can be arranged in the upper left region of the 8×8 region.
[0173] Conversely, in the inverse transformation process, when applying the inverse RST or LFNST to an 8×8 region, the corresponding 16 transform coefficients in the 8×8 region corresponding to the upper left region of the 8×8 region among the transform coefficients in the 8×8 region can be input as a one-dimensional array according to the scanning order, and can undergo matrix operations with the 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix) * (16×1 transform coefficient vector) = (48×1 modified transform coefficient vector). Here, the n×1 vector can be interpreted as having the same meaning as the n×1 matrix, and thus can be represented as an n×1 column vector. In addition, * represents matrix multiplication. When performing the matrix operation, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left region, upper right region, and lower left region of the 8×8 region except for the lower right region.
[0174] When the second inverse transformation is based on RST, the inverse transformer 135 of the encoding device 100 and the inverse transformer 222 of the decoding device 200 can include an inverse reduced second transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients and an inverse first transformer for deriving the residual samples of the target block based on the inverse first transformation of the modified transform coefficients. The inverse first transformation refers to the inverse transformation of the first transformation applied to the residuals. In the present disclosure, deriving transform coefficients based on a transformation can refer to deriving transform coefficients by applying the transformation.
[0175] The non-separable transformation (LFNST) described above will be described in detail below. The LFNST can include a forward transformation performed by the encoding device and an inverse transformation performed by the decoding device.
[0176] The encoding device receives the result (or a part of the result) derived after applying the first (core) transformation as input, and applies the forward second transformation (second transformation).
[0177] [Equation 9]
[0178] y = G T x
[0179] In Equation 9, x and y are the input and output of the second transformation respectively, G is the matrix representing the second transformation, and the transform basis vectors are composed of column vectors. In the case of the inverse LFNST, when the dimension of the transformation matrix G is represented as [number of rows × number of columns], in the case of the forward LFNST, the transpose of the matrix G becomes the dimension of G T of.
[0180] For the inverse LFNST, the dimensions of matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of 8 transform basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.
[0181] On the other hand, for the forward LFNST, matrix G T has dimensions of [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transform basis vectors from the upper part of the [16×48] matrix and the [16×16] matrix, respectively.
[0182] Therefore, in the case of the forward LFNST, a [48×1] vector or a [16×1] vector can be used as the input x, and a [16×1] vector or an [8×1] vector can be used as the output y. In video coding and decoding, the output of the forward single transform is two-dimensional (2D) data. Therefore, in order to construct a [48×1] vector or a [16×1] vector as the input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data that is the output of the forward transform.
[0183] Figure 6 is a diagram illustrating the order of arranging the output data of the forward single transform into a one-dimensional vector according to the example. Figure 6 The left diagrams of (a) and (b) of illustrate the order for constructing a [48×1] vector, and Figure 6 the right diagrams of (a) and (b) of illustrate the order for constructing a [16×1] vector. In the case of LFNST, a one-dimensional vector x can be obtained by arranging the 2D data in the same order as in Figure 6 the (a) and (b) of.
[0184] The arrangement direction of the output data of the forward single transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in the horizontal direction with respect to the diagonal direction, the output data of the forward single transform can be arranged in the order of Figure 6 the (a) of, and when the intra prediction mode of the current block is in the vertical direction with respect to the diagonal direction, the output data of the forward single transform can be arranged in the order of Figure 6 the (b) of.
[0185] According to the example, an arrangement order different from the Figure 6 arrangement order of (a) and (b) of can be applied, and in order to derive an arrangement order different from the Figure 6For the same result (y vector) in the arrangement orders of (a) and (b), the column vectors of matrix G can be rearranged according to the arrangement order. That is, the column vectors of G can be rearranged such that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0186] Since the output y derived by Equation 9 is a one-dimensional vector, when two-dimensional data is required as input data during the process of using the result of the forward quadratic transform as input (for example, during quantization or residual coding), the output y vector of Equation 9 needs to be appropriately arranged as 2D data again.
[0187] Figure 7 FIG. is a diagram illustrating the order of arranging the output data of the forward quadratic transform into a two-dimensional block according to an example.
[0188] In the case of LFNST, the output values can be arranged in a 2D block according to a predetermined scan order. Figure 7 FIG. (a) shows that when the output y is a [16×1] vector, the output values are arranged at 16 positions in a 2D block according to the diagonal scan order. Figure 7 FIG. (b) shows that when the output y is an [8×1] vector, the output values are arranged at 8 positions in a 2D block according to the diagonal scan order, and the remaining 8 positions are filled with zeros. Figure 7 X in FIG. (b) indicates that it is filled with zeros.
[0189] According to another example, since the order of processing the output vector y during quantization or residual coding can be preset, the output vector y may not be arranged in a 2D block as shown in Figure 7 . However, in the case of residual coding, data coding can be performed in units of 2D blocks (for example, 4×4) (for example, CG (coefficient group)), and in this case, the data is arranged according to a specific order in the diagonal scan order as shown in Figure 7 .
[0190] In addition, the decoding device can configure a one-dimensional input vector y by arranging the two-dimensional data output through the dequantization process according to a preset scan order for inverse transform. The input vector y can be output as an output vector x through the following equation.
[0191] [Equation 10]
[0192] x = Gy
[0193] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.
[0194] The output vector x is arranged in a two-dimensional block according to Figure 6 the order shown in Figure 6 , and is arranged as two-dimensional data, and this two-dimensional data becomes the input data (or a part of the input data) of the inverse first transformation.
[0195] Therefore, the inverse second transformation is overall the opposite of the forward second transformation process, and in the case of the inverse transformation, different from the forward direction, the inverse second transformation is first applied, and then the inverse first transformation is applied.
[0196] In the inverse LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be selected as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.
[0197] In addition, eight matrices can be derived from the four transformation sets shown in Table 2 above, and each transformation set can be composed of two matrices. Which transformation set to use among the four transformation sets is determined according to the intra-frame prediction mode, and more specifically, based on the value of the intra-frame prediction mode extended by considering wide-angle intra-frame prediction (WAIP). Which matrix to select from the two matrices constituting the selected transformation set is derived by index signaling. More specifically, 0, 1, and 2 can be used as the transmitted index values, 0 can indicate that LFNST is not applied, and 1 and 2 can indicate any one of the two transformation matrices constituting the transformation set selected based on the intra-frame prediction mode value.
[0198] Furthermore, as described above, which transformation matrix among the [48×16] matrix and the [16×16] matrix to apply to the LFNST is determined by the size and shape of the transformation target block.
[0199] Figure 8 is a diagram illustrating the block shapes to which the LFNST is applied. Figure 8 (a) of Figure 8 shows a 4×4 block, Figure 8 (b) of Figure 8 shows 4×8 and 8×4 blocks, Figure 8 (c) of Figure 8 shows 4×N or N×4 blocks, where N is 16 or greater, Figure 8 (d) of Figure 8 shows an 8×8 block, Figure 8 and (e) of Figure 8 shows M×N blocks, where M≥8, N≥8 and N>8 or M>8.
[0200] In Figure 8 , the blocks with thick boundaries indicate the regions to which the LFNST is applied. For the blocks of Figure 8 (a) and (b) of Figure 8 , the LFNST is applied to the upper left 4×4 region, and for the block of Figure 8 (c) of Figure 8 , the LFNST is separately applied to two continuously arranged upper left 4×4 regions. InFigure 8 In (a), (b), and (c), since the LFNST is applied in units of 4×4 regions, the LFNST will be referred to as the "4×4 LFNST" hereinafter. Based on the matrix dimensions of G, a [16×16] or [16×8] matrix can be applied.
[0201] More specifically, a [16×8] matrix is applied to the Figure 8 4×4 blocks (4×4 TUs or 4×4 CUs) in (a), and a [16×16] matrix is applied to the Figure 8 blocks in (b) and (c). This is to adjust the worst-case computational complexity to 8 multiplications per sample.
[0202] Regarding Figure 8 (d) and (e), the LFNST is applied to the upper-left 8×8 region, and this LFNST is referred to as the "8×8 LFNST" hereinafter. As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of the forward LFNST, since a [48×1] vector (the X vector in Equation 9) is input as the input data, not all sample values in the upper-left 8×8 region are used as the input values for the forward LFNST. That is, as can be seen from the left-side order in Figure 6 (a) or the left-side order in Figure 6 (b), a [48×1] vector can be constructed based on the samples belonging to the remaining 3 4×4 blocks while leaving the lower-right 4×4 block intact.
[0203] A [48×8] matrix can be applied to the Figure 8 8×8 blocks (8×8 TUs or 8×8 CUs) in (d), and a [48×16] matrix can be applied to the Figure 8 8×8 blocks in (e). This is also to adjust the worst-case computational complexity to 8 multiplications per sample.
[0204] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data (the Y vector in Equation 9, an [8×1] or [16×1] vector) are generated. In the forward LFNST, due to the T characteristics of the matrix G, the number of output data is equal to or less than the number of input data.
[0205] Figure 9 is a diagram illustrating the arrangement of the output data of the forward LFNST according to the example, and shows the blocks in which the output data of the forward LFNST are arranged according to the block shape.
[0206] In Figure 9The shaded area in the upper left of the block shown corresponds to the area where the output data of the forward LFNST is located. The positions marked with 0 indicate samples filled with the value 0, and the remaining area represents the area not changed by the forward LFNST. In the area not changed by the LFNST, the output data of the forward transform remains unchanged.
[0207] As described above, since the size of the applied transformation matrix varies according to the shape of the block, the number of output data also varies. As Figure 9 , the output data of the forward LFNST may not completely fill the upper left 4×4 block. In Figure 9 the cases of (a) and (d), a [16×8] matrix and an A[48×8] matrix are applied to the block indicated by the thick line or a partial area inside the block, respectively, and an [8×1] vector as the output of the forward LFNST is generated. That is, according to Figure 7 the scanning order shown in (b), only 8 output data can be filled, as shown in Figure 9 the cases of (a) and (d), and 0 can be filled in the remaining 8 positions. In Figure 8 the case of the block to which the LFNST in (d) is applied, as shown in Figure 9 the (d), the two 4×4 blocks in the upper right and lower left adjacent to the upper left 4×4 block are also filled with the value 0.
[0208] As described above, basically, by signaling the LFNST index, it is specified whether the LFNST is applied and which transformation matrix is to be applied. As Figure 9 shown, when the LFNST is applied, since the number of output data of the forward LFNST can be equal to or less than the number of input data, areas filled with zero values appear as follows.
[0209] 1) As shown in Figure 9 the (a), samples from the eighth position and subsequent positions in the scanning order in the upper left 4×4 block, that is, samples from the ninth to the sixteenth.
[0210] 2) As shown in Figure 9 the (d) and (e), when a [48×16] matrix or a [48×8] matrix is applied, the two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scanning order.
[0211] Therefore, if non-zero data exists by checking areas 1) and 2), it is determined that the LFNST is not applied, so that signaling of the corresponding LFNST index can be omitted.
[0212] According to an example, for instance, in the case of LFNST adopted in the VVC standard, since signaling of the LFNST index is performed after residual coding, the encoding device can know whether non-zero data (valid coefficients) exists at all positions within a TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform signaling of the LFNST index based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. When non-zero data does not exist in the regions specified in 1) and 2) above, signaling of the LFNST index is performed.
[0213] Since the truncated unary code is applied as the binarization method of the LFNST index, the LFNST index consists of up to two bins, and 0, 10, and 11 are assigned as the binary codes for the possible LFNST index values 0, 1, and 2 respectively. According to an example, context-based CABAC coding can be applied to the first bin (conventional coding), and context-based CABAC coding can also be applied to the second bin. The coding of the LFNST index is shown in the following table.
[0214] [Table 3]
[0215]
[0216] As shown in Table 3, for the first bin (binIdx = 0), context 0 is applied in the case of a single tree, while in the case of a non-single tree, context 1 can be applied. In addition, as shown in Table 3, context 2 can be applied to the second bin (binIdx = 1). That is, two contexts can be assigned to the first bin, one context can be assigned to the second bin, and each context can be distinguished by the ctxInc value (0, 1, 2).
[0217] Here, a single tree means that the luminance component and the chrominance component are encoded using the same coding structure. When the coding unit is divided while having the same coding structure, and the size of the coding unit becomes less than or equal to a specific threshold, and the luminance component and the chrominance component are encoded in separate tree structures, the corresponding coding unit is regarded as a double tree, and thus, the context of the first bin can be determined. That is, as shown in Table 3, context 1 can be assigned.
[0218] Alternatively, when the value of the variable treeType is assigned to SINGLE_TREE for the first bin, context 0 can be used, otherwise context 1 can be used.
[0219] In addition, the following simplified method can be applied to the adopted LFNST.
[0220] (i) According to the example, the number of output data of the forward LFNST can be limited to a maximum of 16.
[0221] In Figure 8 case (c), the 4×4 LFNST can be applied to two adjacent 4×4 regions in the upper left, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to the maximum of 16, in the case of 4×N / N×4 (N≥16) blocks (TU or CU), the 4×4 LFNST is only applied to one 4×4 region in the upper left, and the LFNST can be applied to Figure 8 all blocks at once. By this, the implementation of image coding can be simplified.
[0222] Figure 10 shows that the number of output data of the forward LFNST according to the example is limited to the maximum of 16. As Figure 10 , when the LFNST is applied to the uppermost left 4×4 region in a 4×N or N×4 block (where N is 16 or greater), the output data of the forward LFNST becomes 16.
[0223] (ii) According to the example, the regions to which the LFNST is not applied can be additionally cleared. In this document, clearing can mean filling all positions belonging to a specific region with a value of 0. That is, clearing can be applied to the regions that are not changed due to the LFNST, and the result of the forward single transformation can be maintained. As described above, since the LFNST is divided into 4×4 LFNST and 8×8 LFNST, clearing can be divided into two types as follows ((ii)-(A) and (ii)-(B)).
[0224] (ii)-(A) When the 4×4 LFNST is applied, the regions to which the 4×4 LFNST is not applied can be cleared. Figure 11 is a diagram illustrating clearing in a block to which the 4×4 LFNST is applied according to the example.
[0225] As Figure 11 shown, regarding the block to which the 4×4 LFNST is applied, that is, for Figure 9 all blocks in (a), (b), and (c), the entire region to which the LFNST is not applied can be filled with zeros.
[0226] On the other hand, Figure 11 (d) shows that when the maximum value of the number of output data of the forward LFNST is limited to 16 (as Figure 10 shown), clearing is performed on the remaining blocks to which the 4×4 LFNST is not applied.
[0227] (ii)-(B) When applying the 8×8 LFNST, the areas where the 8×8 LFNST is not applied can be cleared to zero. Figure 12 It is a diagram illustrating the clearing in the block applying the 8×8 LFNST according to the example.
[0228] As Figure 12 shown, regarding the block applying the 8×8 LFNST, that is, for all the blocks in Figure 9 (d) and (e), the entire area where the LFNST is not applied can be filled with zeros.
[0229] (iii) Due to the clearing presented in (ii) above, the area filled with zeros may not be the same as when the LFNST is applied. Therefore, the clearing proposed in (ii) can be performed for a wider area according to the case of comparing the Figure 9 LFNST to check for the presence of non-zero data.
[0230] For example, when (ii)-(B) is applied, after checking whether there is non-zero data in the area filled with zeros in Figure 9 (d) and (e), additionally checking whether there is non-zero data in the area filled with 0 in Figure 12 , the signaling for the LFNST index can be performed only when there is no non-zero data.
[0231] Of course, even when the clearing proposed in (ii) is applied, the presence of non-zero data can be checked in the same way as the existing LFNST index signaling. That is, after checking whether there is non-zero data in the block filled with zeros in Figure 9 , the LFNST index signaling can be applied. In this case, the encoding device only performs the clearing and the decoding device does not assume the clearing, that is, only checking whether non-zero data exists only in the area Figure 9 explicitly marked as 0 in, the LFNST index parsing can be performed.
[0232] Various embodiments of the combination of the simplified methods ((i), (ii)-(A), (ii)-(B), (iii)) for applying the LFNST can be derived. Of course, the combination of the above simplified methods is not limited to the following embodiments, and any combination can be applied to the LFNST.
[0233] Embodiment
[0234] - Limit the number of output data of the forward LFNST to a maximum of 16 → (i)
[0235] - When applying the 4×4 LFNST, all areas where the 4×4 LFNST is not applied are cleared to zero → (II)-(A)
[0236] - When applying 8×8 LFNST, all regions where 8×8 LFNST is not applied are cleared → (II)-(B)
[0237] - After checking whether non-zero data also exists in the existing regions filled with zero values and the regions filled with zero due to additional clearing ((ii)-(A), (ii)-(B)), signal the LFNST index only when no non-zero data exists → (iii).
[0238] In the case of the embodiment, when applying LFNST, the regions where non-cleared data can exist are limited to the inside of the upper left 4×4 region. More specifically, in Figure 11 of (a) and Figure 12 of (a), the eighth position in the scanning order is the last position where non-zero data can exist. In Figure 11 of (b) and (c) and Figure 12 of (b), the sixteenth position in the scanning order (i.e., the position of the lower right edge of the upper left 4×4 block) is the last position where data other than 0 can exist.
[0239] Therefore, after applying LFNST, after checking whether non-zero data exists at positions where the residual coding process does not allow (at positions beyond the last position), it can be determined whether to signal the LFNST index.
[0240] In the case of the clearing method proposed in (ii), due to the quantity of the finally generated data when both a transform and LFNST are applied once, the computational amount required to execute the entire transform process can be reduced. That is, when LFNST is applied, since clearing is applied to the regions where the forward one-time transform output data exists and LFNST is not applied, it is not necessary to generate data for the regions that become cleared during the execution of the forward one-time transform. Therefore, the computational amount required to generate the corresponding data can be reduced. The additional effects of the clearing method proposed in (ii) are summarized as follows.
[0241] First, as described above, reduce the computational amount required to execute the entire transform process.
[0242] In particular, when applying (ii)-(B), the worst-case computational amount is reduced, making the transform process lighter. In other words, generally, a large amount of computation is required to execute a large-size one-time transform. By applying (ii)-(B), the quantity of the data derived as a result of executing the forward LFNST can be reduced to 16 or less. Additionally, as the size of the entire block (TU or CU) increases, the effect of reducing the amount of transform operations further increases.
[0243] Second, the amount of computation required for the entire transformation process can be reduced, thereby reducing the power consumption required to perform the transformation.
[0244] Third, the latency involved in the transformation process is reduced.
[0245] Secondary transformations such as LFNST add computational complexity to existing primary transformations, thus increasing the overall latency involved in performing the transformation. In particular, in the case of intra prediction, since the reconstructed data of adjacent blocks is used during the prediction process, during encoding, the increase in latency due to the secondary transformation leads to an increase in the latency until reconstruction. This can result in an increase in the overall latency of intra prediction coding.
[0246] However, if the zeroing proposed in application (ii) is applied, the latency time for performing the primary transformation can be greatly reduced when LFNST is applied, maintaining or reducing the overall latency of the transformation, such that the encoding device can be implemented more simply.
[0247] In addition, in traditional intra prediction, the block to be currently encoded is regarded as one coding unit, and encoding is performed without division. However, intra sub-partition (ISP) coding means performing intra prediction coding by dividing the block to be currently encoded in the horizontal or vertical direction. In this case, reconstructed blocks can be generated by performing encoding / decoding in units of the divided blocks, and the reconstructed blocks can be used as reference blocks for the next divided block. According to an embodiment, in ISP coding, one coding block can be divided into two or four sub-blocks and encoded, and in ISP, within one sub-block, intra prediction is performed by referring to the reconstructed pixel values of the sub-blocks located adjacent to the left or adjacent to the upper side. Hereinafter, "encoding" can be used as a concept including both encoding performed by an encoding device and decoding performed by a decoding device.
[0248] In addition, the signaling order of the LFNST index and the MTS index will be described below.
[0249] According to an example, the LFNST index signaled in the residual coding can be encoded after the coding position of the last non-zero coefficient position, and the MTS index can be encoded immediately after the LFNST index. In the case of this configuration, the LFNST index can be signaled for each transform unit. Alternatively, even if not signaled in the residual coding, the LFNST index can be encoded after the coding of the last valid coefficient position, and the MTS index can be encoded after the LFNST index.
[0250] The syntax of the residual coding according to an example is as follows.
[0251] [Table 4]
[0252]
[0253]
[0254] The meanings of the main variables shown in Table 4 are as follows.
[0255] 1. cbWidth, cbHeight: The width and height of the current coding block
[0256] 2. log2TbWidth, log2TbHeight: The base-2 logarithms of the width and height of the current transform block, which can be reduced by reflection zeroing to the upper left region where non-zero coefficients can exist.
[0257] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is enabled. If the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in the Sequence Parameter Set (SPS).
[0258] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variables chType and the position (x0, y0). chType can have values of 0 and 1, where 0 indicates the luminance component and 1 indicates the chrominance component. The position (x0, y0) indicates a position on the picture, and MODE_INTRA (intra prediction) and MODE_INTER (inter prediction) can be used as the values of CuPredMode[chType][x0][y0].
[0259] 5. IntraSubPartitionsSplit[x0][y0]: The content at the position (x0, y0) is the same as in item 4 above. It indicates which ISP partition is applied at the position (x0, y0). ISP_NO_SPLIT indicates that the coding unit corresponding to the position (x0, y0) is not divided into sub-blocks.
[0260] 6. intra_mip_flag[x0][y0]: The content at the position (x0, y0) is the same as in item 4 above. intra_mip_flag is a flag indicating whether the matrix-based intra prediction (MIP) prediction mode is applied. If the flag value is 0, it indicates that MIP is not enabled, and if the flag value is 1, it indicates that MIP is enabled.
[0261] 7. cIdx: The value 0 indicates luminance, and the values 1 and 2 indicate Cb and Cr of the chrominance components, respectively.
[0262] 8. treeType: Indicates single tree, dual tree, etc. (SINGLE_TREE: single tree, DUAL_TREE_LUMA: dual tree for luminance component, DUAL_TREE_CHROMA: dual tree for chrominance component)
[0263] 9. tu_cbf_cb[x0][y0]: The content at position (x0, y0) is the same as in item 4. It indicates the coding block flag (CBF) of the Cb component. If its value is 0, it means that there are no non-zero coefficients in the corresponding transform unit of the Cb component, and if its value is 1, it indicates that there are non-zero coefficients in the corresponding transform unit of the Cb component.
[0264] 10. lastSubBlock: It indicates the position of the sub-block (coefficient group (CG)) where the last non-zero coefficient is located in the scanning order. 0 indicates the sub-block containing the DC component, and in the case of being greater than 0, it is not the sub-block containing the DC component.
[0265] 11. lastScanPos: It indicates the position where the last valid coefficient is located in a sub-block in the scanning order. If a sub-block includes 16 positions, values from 0 to 15 can be there.
[0266] 12. lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If not parsed, it is inferred to be the value 0. That is, the default value is set to 0, indicating that LFNST is not applied.
[0267] 13. LastSignificantCoeffX, LastSignificantCoeffY: They indicate the x and y coordinates where the last valid coefficient in the transform block is located. The x coordinate starts from 0 and increases from left to right, and the y coordinate starts from 0 and increases from top to bottom. If the values of both variables are 0, it means that the last valid coefficient is located at DC.
[0268] 14. cu_sbt_flag: A flag indicating whether sub-block transform (SBT) included in the current VVC standard is enabled. If the flag value is 0, it indicates that SBT is not enabled, and if the flag value is 1, it indicates that SBT is enabled.
[0269] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: Flags indicating whether explicit MTS is applied to inter-frame CUs and intra-frame CUs respectively. If the corresponding flag value is 0, it indicates that MTS is not enabled for inter-frame CUs or intra-frame CUs, and if the corresponding flag value is 1, it indicates that MTS is enabled.
[0270] 16. tu_mts_idx[x0][y0]: The MTS index syntax element to be parsed. If not parsed, it is inferred to be the value 0. That is, the default value is set to 0, indicating that DCT-2 is enabled in both the horizontal and vertical directions.
[0271] As shown in Table 4, in the case of a single tree, it is possible to determine whether to signal the LFNST index only using the last significant coefficient position condition for luminance. That is, if the position of the last significant coefficient is not DC and the last significant coefficient exists in the upper left sub-block (CG) (e.g., a 4×4 block), then the LFNST index is signaled. In this case, in the case of 4×4 transform blocks and 8×8 transform blocks, the LFNST index is signaled only when the last significant coefficient exists at positions 0 to 7 in the upper left sub-block.
[0272] In the case of a dual tree, independent of each of luminance and chrominance, the LFNST index is signaled, and in the case of chrominance, the LFNST index can be signaled by applying the last significant coefficient position condition only to the Cb component. For the Cr component, the corresponding condition may not be checked, and if the CBF value of Cb is 0, then the LFNST index can be signaled by applying the last significant coefficient position condition to the Cr component.
[0273] "Min(log2TbWidth, log2TbHeight) >= 2" in Table 4 can be expressed as "Min(tbWidth, tbHeight) >= 4", while "Min(log2TbWidth, log2TbHeight) >= 4" can be expressed as "Min(tbWidth, tbHeight) >= 16".
[0274] In Table 4, log2ZoTbWidth and log2ZoTbHeight respectively mean the base-2 logarithms of the width and height of the upper left region where the last significant coefficient can exist by zeroing.
[0275] As shown in Table 4, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places. The first is before parsing the MTS index or the LFNST index value, and the second is after parsing the MTS index.
[0276] The first update is before parsing the value of the MTS index (tu_mts_idx[x0][y0]), so the log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.
[0277] After parsing the MTS index, set log2ZoTbWidth and log2ZoTbHeigh for MTS indices (DST-7 / DCT-8 combination) greater than 0. When DST-7 / DCT-8 is independently applied to each of the horizontal and vertical directions in a single transform, there can be up to 16 valid coefficients per row or column in each direction. That is, after applying DST-7 / DCT-8 with a length of 32 or greater, up to 16 transform coefficients can be derived for each row or column starting from the left or top. Therefore, in a 2D block, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions, valid coefficients can exist only in the up-left region of up to 16×16.
[0278] In addition, when DCT-2 is independently applied to each of the horizontal and vertical directions in the current single transform, there can be up to 32 valid coefficients per row or column in each direction. That is, when applying DCT-2 with a length of 64 or greater, up to 32 transform coefficients can be derived for each row or column starting from the left or top. Therefore, in a 2D block, when DCT-2 is applied to both the horizontal and vertical directions, valid coefficients can exist only in the up-left region of up to 32×32.
[0279] In addition, when DST-7 / DCT-8 is applied to one side and DCT-2 is applied to the other side for the horizontal and vertical directions, there can be 16 valid coefficients in the former direction and 32 valid coefficients in the latter direction. For example, in the case of a 64×8 transform block, if DCT-2 is applied in the horizontal direction and DST-7 is applied in the vertical direction (which may occur when applying implicit MTS), valid coefficients can exist in the up-left region of up to 32×8.
[0280] If, as shown in Table 4, log2ZoTbWidth and log2ZoTbHeight are updated in two places, that is, before parsing the MTS index, the ranges of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight as shown in the following table.
[0281] [Table 5]
[0282]
[0283] Additionally, in this case, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values during the binarization process of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.
[0284] [Table 6]
[0285]
[0286] According to the example, in the case of applying the ISP mode and LFNST, when applying the signaling of Table 4, the specification text can be configured as shown in Table 7. Compared with Table 4, the condition for signaling the LFNST index only when the ISP mode is not included (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT in Table 4) is removed.
[0287] In a single tree, when the LFNST index sent for the luminance component (cIdx = 0) is reused for the chrominance component, the LFNST index sent for the first ISP partition block with valid coefficients can be applied to the chrominance transform block. Alternatively, even in a single tree, the LFNST index can be signaled for the chrominance component separately from the LFNST index signaled for the luminance component. The descriptions of the variables in Table 7 are the same as those in Table 4.
[0288] [Table 7]
[0289]
[0290] According to the example, the LFNST index and / or the MTS index can be signaled at the coding unit level. As described above, the LFNST index can have three values 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate the first candidate and the second candidate among the two LFNST kernel candidates included in the selected LFNST set, respectively. The LFNST index is encoded by truncated unary binarization, and the values 0, 1, 2 can be encoded as the bin strings 0, 10, 11, respectively.
[0291] According to the example, the LFNST can be applied only when the DCT-2 is applied to both the horizontal and vertical directions in a single transform. Therefore, if the MTS index is signaled after the LFNST index is signaled, the MTS index can be signaled only when the LFNST index is 0, and when the LFNST index is not 0, a single transform can be performed by applying the DCT-2 to both the horizontal and vertical directions without signaling the MTS index.
[0292] The MTS index can have values 0, 1, 2, 3, and 4, where 0, 1, 2, 3, and 4 can indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, DCT-8 / DCT-8 are applied to the horizontal and vertical directions respectively. Additionally, the MTS index can be encoded by truncated unary binarization, and the values 0, 1, 2, 3, 4 can be encoded as bin strings of 0, 10, 110, 1110, 1111 respectively.
[0293] Signaling the LFNST index at the coding unit level can be indicated as shown in the following table. The LFNST index can be signaled in the second half of the coding unit syntax table.
[0294] [Table 8]
[0295]
[0296] The variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag in Table 8 can be set as shown in Table 11 below.
[0297] The variable LfnstDcOnly is equal to 1 when for a transform block with a coding block flag (CBF) of 1 (equal to 0 if there is at least one valid coefficient in the corresponding block, otherwise equal to 0), all the last valid coefficients are located at the DC position (top-left position), and equal to 0 otherwise. Specifically, in the case of dual-tree luminance, the position of the last valid coefficient is checked for a luminance transform block, and in the case of dual-tree chrominance, the position of the last valid coefficient is checked for both the Cb transform block and the Cr transform block. In the case of single-tree, the position of the last valid coefficient can be checked for the luminance, Cb, and Cr transform blocks.
[0298] If there is a valid coefficient at the zeroing position when applying the LFNST, the variable LfnstZeroOutSigCoeffFlag is equal to 0, otherwise equal to 1.
[0299] The lfnst_idx[x0][y0] included in Table 8 and subsequent tables indicates the LFNST index of the corresponding coding unit, while tu_mts_idx[x0][y0] indicates the MTS index of the corresponding coding unit.
[0300] According to the example, in order to code the MTS index continuously after the LFNST index at the coding unit level, the coding unit syntax table can be configured as shown in Table 9.
[0301] [Table 9]
[0302]
[0303] Comparing Table 9 with Table 8, the condition for checking whether the value of tu_mts_idx[x0][y0] is 0 in the condition for signaling lfnst_idx[x0][y0] (i.e., checking whether DCT-2 is applied to both the horizontal and vertical directions) is changed to the condition for checking whether the value of transform_skip_flag[x0][y0] is 0 (!transform_skip_flag[x0][y0]). transform_skip_flag[x0][y0] indicates whether the coding unit is coded in a transform skip mode in which the transform is skipped, and this flag is signaled before the MTS index and the LFNST index. That is, since lfnst_idx[x0][y0] is signaled before the value of tu_mtx_idx[x0][y0] is signaled, the condition regarding the value of transform_skip_flag[x0][y0] can be checked only.
[0304] As shown in Table 9, multiple conditions are checked when coding tu_mts_idx[x0][y0], and as described above, tu_mts_idx[x0][y0] is signaled only when the value of lfnst_idx[x0][y0] is 0.
[0305] tu_cbf_luma[x0][y0] is a flag indicating whether there are valid coefficients for the luminance component, and cbWidth and cbHeight indicate the width and height of the coding unit of the luminance component, respectively.
[0306] In Table 9, (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT) indicates that the ISP mode is not applied, and (!cu_sbt_flag) indicates that SBT is not applied.
[0307] According to Table 9, when both the width and height of the coding unit of the luminance component are 32 or smaller, tu_mts_idx[x0][y0] is signaled, that is, whether to apply MTS is determined by the width and height of the coding unit of the luminance component.
[0308] According to another example, when transform unit (TU) tiling occurs (for example, when the maximum transform size is set to 32, a 64×64 coding unit is divided into 4 32×32 transform blocks and encoded), the MTS index can be signaled based on the size of each transform block. For example, when both the width and height of the transform block are 32 or smaller, the same MTS index value can be applied to all transform blocks in the coding unit, thereby applying the same single transform. Additionally, when transform block tiling occurs, the value of tu_cbf_luma[x0][y0] in Table 9 can be the CBF value of the top-left transform block, or can be set to 1 when the CBF value of even one transform block among all transform blocks is 1.
[0309] According to the example, when the ISP mode is applied to the current block, LFNST can be applied, and in this case Table 9 can be changed as shown in Table 10.
[0310] [Table 10]
[0311]
[0312] As shown in Table 10, even in the ISP mode (IntraSubPartitionsSplitType!= ISP_NO_SPLIT), lfnst_idx[x0][y0] can be configured to be signaled, and the same LFNST index value can be applied to all ISP sub-blocks.
[0313] Furthermore, as shown in Table 10, since tu_mts_idx[x0][y0] is signaled only in modes other than the ISP mode, the MTS index coding part is the same as in Table 9.
[0314] As shown in Tables 9 and 10, when the MTS index is signaled immediately after the LFNST index, information about the single transform cannot be known during residual coding. That is, the MTS index is signaled after residual coding. Therefore, in the residual coding part, the part where zeroing is performed while only retaining 16 coefficients for DST-7 or DCT-8 of length 32 can be changed as shown in Table 11 below.
[0315] [Table 11]
[0316]
[0317]
[0318] As shown in Table 11, in the process of determining log2ZoTbWidth and log2ZoTbHeight (where log2ZoTbWidth and log2ZoTbHeight respectively represent the base-2 logarithms of the width and height of the upper-left region remaining after performing zeroing), it is possible to omit checking the value of tu_mts_idx[x0][y0].
[0319] The binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 11 can be determined based on log2ZoTbWidth and log2ZoTbHeight as shown in Table 6.
[0320] In addition, as shown in Table 11, when determining log2ZoTbWidth and log2ZoTbHeight in residual coding, a condition for checking the sps_mts_enable_flag can be added.
[0321] TR in Table 6 indicates the truncated Rice binarization method, and the last significant coefficient information can be binarized based on cMax and cRiceParam defined in Table 6 according to the method described in the following table.
[0322] [Table 12]
[0323]
[0324]
[0325] In addition, according to another example, the coding unit syntax table, transform unit syntax table, and residual coding syntax table are as follows. According to Table 13, the MTS index moves from the transform unit level to the coding unit level syntax and is signaled after the LFNST index signaling. Also, the constraint that LFNST is not allowed when ISP is applied to the coding unit has been removed. When ISP is applied to the coding unit, the constraint that LFNST is not allowed is removed so that LFNST can be applied to all intra prediction blocks. Also, both the MTS index and the LFNST index are conditionally signaled at the end part of the coding unit level.
[0326] [Table 13]
[0327]
[0328] [Table 14]
[0329]
[0330] [Table 15]
[0331]
[0332] In Table 13, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed in the residual coding of Table 15. When there are valid coefficients in the area filled with 0 by zeroing (LastSignificantCoeffX>15||LastSignificantCoeffY>15), the value of the variable MtsZeroOutSigCoeffFlag changes from 1 to 0. In this case, the MTS index is not signaled, as shown in Table 15.
[0333] In addition, as shown in Table 13, when tu_cbf_luma[x0][y0] is 0, the mts_idx[x0][y0] coding can be omitted. That is, when the CBF value of the luminance component is 0, since no transformation is applied, there is no need to signal the MTS index, so the MTS index coding can be omitted.
[0334] According to the example, the above technical features can be implemented with another conditional syntax. For example, after performing MTS, a variable indicating whether there are valid coefficients in the area other than the DC area of the current block can be derived, and when this variable indicates that there are valid coefficients in the area other than the DC area, the MTS index can be signaled. That is, the presence of valid coefficients in the area other than the DC area of the current block indicates that the value of tu_cbf_luma[x0][y0] is 1, and in this case, the MTS index can be signaled.
[0335] This variable can be represented as MtsDcOnly, and after the variable MtsDcOnly is initially set to 1 at the coding unit level, this value is changed to 0 when it is determined at the residual coding level that there are valid coefficients in the area other than the DC area of the current block. When the variable MtsDcOnly is 0, the image information can be configured to signal the MTS index.
[0336] When tu_cbf_luma[x0][y0] is 0, since the residual coding syntax is not called at the transform unit level of Table 14, the initial value of the variable MtsDcOnly is kept as 1. In this case, since the variable MtsDcOnly is not changed to 0, the image information can be configured so that the MTS index is not signaled. That is, the MTS index is not parsed and signaled.
[0337] In addition, the decoding device may determine the color index cIdx of the transform coefficient to derive the variable MtsZeroOutSigCoeffFlag in Table 15. A color index cIdx of 0 indicates the luminance component.
[0338] According to the example, since MTS can be applied only to the luminance component of the current block, the decoding device may determine whether the color index is luminance when deriving the variable MtsZeroOutSigCoeffFlag used to determine whether to parse the MTS index.
[0339] The variable MtsZeroOutSigCoeffFlag is a variable indicating whether zeroing is performed when applying MTS. It indicates whether there are transform coefficients in the area outside the upper-left area where the last valid coefficient may exist due to zeroing after performing MTS (i.e., in the area outside the upper-left 16×16 area). The variable MtsZeroOutSigCoeffFlag is initially set to 1 at the coding unit level, as shown in Table 13 (MtsZeroOutSigCoeffFlag = 1), and its value may be changed from 1 to 0 at the residual coding level when there are transform coefficients in the area outside the 16×16 area, as shown in Table 15 (MtsZeroOutSigCoeffFlag = 0). When the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.
[0340] As shown in Table 15, at the residual coding level, the non-zeroing area where non-zero transform coefficients may exist may be set according to whether zeroing accompanying MTS is performed. And even in this case, with the color index (cIdx) being 0, the non-zeroing area may be set to the upper-left 16×16 area of the current block.
[0341] Thus, when deriving the variable for determining whether the MTS index is parsed, the color component is determined to be luminance or chrominance. However, since LFNST can be applied to both the luminance component and the chrominance component of the current block, the color component is not determined when deriving the variable for determining whether to parse the LFNST index.
[0342] For example, Table 13 shows the variable LfnstZeroOutSigCoeffFlag, which can indicate that zeroing is performed when applying LFNST. The variable LfnstZeroOutSigCoeffFlag indicates whether there are valid coefficients in the second region of the current block except for the first region in the upper left. This value is initially set to 1, and when there are valid coefficients in the second region, this value can be changed to 0. Only when the value of the initially set variable LfnstZeroOutSigCoeffFlag remains 1 can the LFNST index be parsed. When determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, since LFNST can be applied to both the luminance component and the chrominance component of the current block, the color index of the current block is not determined.
[0343] Figure 13 An example of MIP-based predicted sample generation processing according to an example is illustrated. Referring to Figure 13 , the MIP processing is described as follows.
[0344] 1. Averaging processing
[0345] Among the boundary samples, four samples for the case of W = H = 4 and eight samples for any other case are extracted through averaging processing.
[0346] 2. Matrix-vector multiplication processing
[0347] The matrix-vector multiplication is performed using the averaged samples as input, followed by adding an offset. Through this operation, the reduced predicted samples of the subsampled sample set in the original block can be derived.
[0348] 3. (Linear) interpolation processing
[0349] The predicted samples at the remaining positions are generated based on the predicted samples of the subsampled sample set through linear interpolation, which is single-step linear interpolation in each direction.
[0350] For the matrix, the matrices and offset vectors required to generate the predicted block or predicted samples can be selected from three sets S0, S1, and S2.
[0351] Set S0 can include 16 matrices A0 i , i ∈ {0,..., 15} and 16 offset vectors b0 i , i ∈ {0,..., 15}, and each matrix can include 16 rows and 4 columns. The matrices and offset vectors in set S0 can be used for 4×4 blocks. In another example, set S0 can include 18 matrices.
[0352] Set S1 can include 8 matrices A1i where \(i\in\{0,\ldots,7\}\) and there are eight offset vectors \(\mathbf{b}_1\) i where \(i\in\{0,\ldots,7\}\), and each matrix can include 16 rows and 8 columns. In another example, the set \(S_1\) can include six matrices. The matrices and offset vectors in the set \(S_1\) can be used for \(4\times8\) blocks, \(8\times4\) blocks, and \(8\times8\) blocks. Alternatively, the matrices and offset vectors in the set \(S_1\) can be used for \(4\times H\) blocks or \(W\times4\) blocks.
[0353] Finally, the set \(S_2\) can include six matrices \(\mathbf{A}_2\) i where \(i\in\{0,\ldots,5\}\) and there are six offset vectors \(\mathbf{b}_2\) i where \(i\in\{0,\ldots,5\}\), and each matrix can include 64 rows and 8 columns. The matrices and offset vectors in the set \(S_2\), or some of them, can be used for any blocks with different sizes to which the sets \(S_0\) and \(S_1\) are not applied. For example, the matrices and offset vectors in the set \(S_2\) can be used for operations on blocks with a height and width of 8 or greater.
[0354] The total number of multiplications required to compute the matrix-vector product is always less than or equal to \(4\times W\times H\). That is, in the MIP mode, up to four multiplications per sample are required.
[0355] Figure 14 Illustrates the CCLM applicable when deriving the intra prediction mode of a chrominance block according to an embodiment.
[0356] In this specification, a "reference sample template" may refer to a set of adjacent reference samples of the current chrominance block used to predict the current chrominance block. The reference sample template can be predefined, and information about the reference sample template can be signaled from the encoding device 100 to the decoding device 200.
[0357] Referring to Figure 14 , the set of shaded samples in a single row adjacent to the \(4\times4\) block that is the current chrominance block refers to the reference sample template. The reference sample template is configured as the reference samples of a single row, while the reference sample region in the luminance region corresponding to the reference sample template is configured as two rows, as Figure 14 shown.
[0358] In an embodiment, when performing intra coding of a chrominance image in the Joint Exploration Test Model (JEM) used in the Joint Video Exploration Team (JVET), a Cross-Component Linear Model (CCLM) can be used. CCLM is a method of predicting the pixel values of a chrominance image based on the pixel values of a reconstructed luminance image, and is based on the high correlation between the luminance image and the chrominance image.
[0359] CCLM prediction of the \(C_b\) and \(C_r\) chrominance images can be performed based on the following formula.
[0360] [Formula 11]
[0361] Pred C (i, j) = α · Rec' L (i, j) + β
[0362] Here, Pred c (i, j) represents the Cb or Cr chrominance image to be predicted, Rec L '(i, j) represents the reconstructed luminance image adjusted to the chrominance block size, and (i, j) represents the coordinates of the pixel. In the 4:2:0 color format, since the size of the luminance image is twice that of the chrominance image, RecL' with the chrominance block size needs to be generated by downsampling. Therefore, Rec L (2i, 2j) and adjacent pixels can be used to adopt the pixels of the luminance image to be used for the chrominance image Pred c (i, j). RecL'(i, j) can be called the downsampled luminance sample.
[0363] For example, as shown in the following formula, RecL'(i, j) can be derived using six adjacent pixels.
[0364] [Formula 12]
[0365] Rec' L (x, y) = (2 × Rec L (2x, 2y) + 2 × Rec L (2x, 2y + 1) + Rec L (2x - 1, 2y) + Rec L (2x + 1, 2y) + Rec L (2x - 1, 2y + 1) + Rec L (2x + 1, 2y + 1) + 4) >> 3
[0366] α and β represent the cross - correlation and the average difference between the adjacent templates of the Cb or Cr chrominance block and Figure 14 the adjacent templates of the luminance block in the shaded area in. For example, α and β are represented by Formula 13.
[0367] [Formula 13]
[0368]
[0369] L(n) represents adjacent reference samples and / or left adjacent samples of a luminance block corresponding to a current chrominance image, C(n) represents adjacent reference samples and / or left adjacent samples of a current chrominance block to which encoding is currently applied, and (i, j) represents a pixel position. Additionally, L(n) may represent downsampled upper adjacent samples and / or left adjacent samples of a current luminance block. N may represent the total number of pixel pair (luminance and chrominance) values used to calculate CCLM parameters, and may indicate a value that is twice the smaller value of the width and height of a current chrominance block.
[0370] A picture can be partitioned into a sequence of coding tree units (CTUs). A CTU can correspond to a coding tree block (CTB). Alternatively, a CTU can include a coding tree block of luminance samples and a coding tree block of corresponding chrominance samples. Depending on whether a luminance block and a corresponding chrominance block have a separate partitioning structure, the tree type can be classified as single tree (SINGLE_TREE) or dual tree (DUAL_TREE). A single tree can indicate that the chrominance block has the same partitioning structure as the luminance block, while a dual tree can indicate that the chrominance component block has a partitioning structure different from that of the luminance block.
[0371] When applying LFNST to a chrominance transform block according to an example, information about a collocated luminance transform block needs to be referred to.
[0372] The existing specification text regarding relevant parts is shown in the following table.
[0373] [Table 16]
[0374]
[0375] As shown in Table 16, when the intra prediction mode in the current frame is the CCLM mode, the value of the variable predModeIntra of the chrominance transform block is determined by taking the intra prediction mode value of the collocated chrominance transform block (the italicized part). The intra prediction mode value (predModeIntra value) of the luminance transform block can subsequently be used to determine the LFNST set.
[0376] However, the variables nTbW and nTbH input as input values for this transform process represent the width and height of the current transform block. Thus, when the current block is a luminance transform block, the variables nTbW and nTbH can represent the width and height of the luminance transform block, while when the current block is a chrominance transform block, the variables nTbW and nTbH represent the width and height of the chrominance transform block.
[0377] Here, the variables nTbW and nTbH in the italicized part of Table 16 represent the width and height of the chrominance transform block that do not reflect the color format and thus do not accurately indicate the reference position of the luminance transform block corresponding to the chrominance transform block. Therefore, the italicized part of Table 16 can be modified as shown in the following table.
[0378] [Table 17]
[0379]
[0380] As shown in Table 17, nTbW and nTbH are respectively changed to (nTbW * SubWidthC) / 2 and (nTbH * SubHeightC) / 2. xTbY and yTbY can represent the luminance position in the current picture (the top-left sample of the current luminance transform block relative to the top-left luminance sample of the current picture), while nTbW and nTbH can represent the width and height of the currently encoded transform block (the variable nTbW specifies the width of the current transform block, and the variable nTbH specifies the height of the current transform block).
[0381] When the currently encoded transform block is a chrominance (Cb or Cr) transform block, nTbW and nTbH are respectively the width and height of the chrominance transform block. Therefore, when the currently encoded transform block is a chrominance transform block (cIdx > 0), the width and height of the luminance transform block need to be used to obtain the reference position of the adjacent luminance transform block when obtaining the reference position. In Table 17, SubWidthC and SubHeightC are values set according to the color format (such as 4:2:0, 4:2:2, or 4:4:4). Specifically, they are respectively the width ratio and height ratio between the luminance component and the chrominance component (see Table 18 below). Therefore, in the case of a chrominance transform block, (nTbW * SubWidthC) and (nTbH * SubHeightC) can be respectively the width and height relative to the adjacent luminance transform block.
[0382] Therefore, xTbY + (nTbW * SubWidthC) / 2 and yTbY + (nTbH * SubHeightC) / 2 represent the values of the central position in the adjacent luminance transform block based on the top-left position of the current picture, and thus precisely indicate the adjacent luminance transform block.
[0383] [Table 18]
[0384]
[0385] In Table 17, the variable predModeIntra represents the intra prediction mode value. When the value of the variable predModeIntra is equal to INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM, it indicates that the current transform block is a chrominance transform block. According to the example, in the current VVC standard, INTRA_LT_CCLM, INTRA_L_CCLM, and INTRA_T_CCLM correspond to the mode values 81, 82, and 83 among the intra prediction mode values respectively. Therefore, as shown in Table 17, the values of xTbY+(nTbW*SubWidthC) / 2 and yTbY+(nTbH*SubHeightC) / 2 are required to obtain the reference position of the collocated luma transform block.
[0386] As shown in Table 17, given that both the variable intra_mip_flag[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] and the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] update the value of predModeIntra.
[0387] intra_mip_flag is a variable indicating whether the current transform block (or coding unit) is encoded by the matrix-based intra prediction (MIP) method, and intra_mip_flag[x][y] is a flag value indicating whether MIP is applied to the position corresponding to the coordinates (x, y) based on the luma component when the upper-left position of the current picture is defined as (0, 0). The x and y coordinates increase from left to right and from top to bottom respectively, and when the flag indicating whether MIP is applied is 1, the flag indicates that MIP is applied. When the flag indicating whether MIP is applied is 0, the flag indicates that MIP is not applied. MIP can be applied only to luma blocks.
[0388] According to the modified part of Table 17, when the value of intra_mip_flag[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] in the collocated luma transform block is 1, the value of predModeIntra is set to the planar mode (INTRA_PLANAR).
[0389] The value of the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the prediction mode value corresponding to the coordinates (xTbY+(nTbW*SubWidthC) / 2, yTbY+(nTbH*SubHeightC) / 2) when the upper-left position of the current picture in the luma component is defined as (0, 0). The prediction mode value can have the values MODE_INTRA, MODE_IBC, MODE_PLT, and MODE_INTER, which represent the intra prediction mode, intra block copy (IBC) prediction mode, palette (PLT) coding mode, and inter prediction mode, respectively. According to Table 17, when the value of the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] is MODE_IBC or MODE_PLT, the value of the variable predModeIntra is set to the DC mode. In cases other than these two cases, the value of the variable predModeIntra is set to IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] (the intra prediction mode value corresponding to the central position in the collocated luma transform block).
[0390] According to the example, as shown in the following table, considering whether to perform wide-angle intra prediction, the value of the variable predModeIntra can be updated once again based on the predModeIntra value updated in Table 17.
[0391] [Table 19]
[0392]
[0393] In the mapping process shown in Table 19, the input values of predModeIntra, nTbW, and nTbH are the same as the values of the variable predModeIntra updated in Table 17 and the values of nTbW and nTbH referenced in Table 17, respectively.
[0394] In Table 19, nCbW and nCbH represent the width and height of the coding block corresponding to the transform block, respectively, and the variable IntraSubPartitionsSplitType indicates whether the ISP mode is applied. When IntraSubPartitionsSplitType is equal to ISP_NO_SPLIT, it indicates that the coding unit is not partitioned by ISP (i.e., the ISP mode is not applied). The variable IntraSubPartitionsSplitType that is not equal to ISP_NO_SPLIT indicates that the ISP mode is applied, and thus the coding unit is partitioned into two or four sub-blocks. In Table 19, cIdx is the index indicating the color component. A cIdx value equal to 0 represents a luminance block, and a cIdx value not equal to 0 indicates a chrominance block. The predModeIntra value output through the mapping process in Table 19 is the updated value considering whether the wide-angle intra prediction (WAIP) mode is applied.
[0395] For the predModeIntra value updated through Table 19, the LFNST set can be determined through the mapping relationship shown in the following table.
[0396] [Table 20]
[0397] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 44 3 45 <= predModeIntra <= 55 2 56 <= predModeIntra <= 80 1
[0398] In the above table, lfnstTrSetIdx represents the index indicating the LFNST set and has values from 0 to 3, which indicates that a total of four LFNST sets are configured. Each LFNST set can include two transform kernels, i.e., the LFNST kernel (depending on the region where LFNST is applied and the forward direction, the transform kernel can be a 16×16 matrix or a 16×48 matrix), and the transform kernel to be applied among the two transform kernels can be specified by signaling of the LFNST index. Additionally, whether to apply LFNST can also be specified by the LFNST index. In the current VVC standard, the LFNST index can have values 0, 1, and 2. 0 indicates that LFNST is not applied, while 1 and 2 indicate the two transform kernels respectively.
[0399] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices illustrated in the drawings or the names of specific signals / messages / fields are provided for illustration purposes, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0400] Figure 15 is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.
[0401] Figure 15 Each process disclosed in Figures 3 to 14Some details of the description. Therefore, the description of specific details that overlap with the details described with reference to Figures 3 to 14 will be omitted or will be described schematically.
[0402] According to an embodiment, the decoding device 200 may obtain intra prediction mode information and an LFNST index from a bitstream (S1510).
[0403] The intra prediction mode information may include the intra prediction modes of adjacent blocks (e.g., the left adjacent block and / or the upper adjacent block) of the current block, and an MPM index indicating one of the MPM candidates in the most probable mode (MPM) list derived based on additional candidate modes or remaining intra prediction mode information indicating one of the remaining intra prediction modes not included in the MPM candidates.
[0404] In addition, the intra mode information may include flag information sps_cclm_enabled_flag indicating whether CCLM is applied to the current block and information intra_chroma_pred_mode on the intra prediction mode of the chrominance component.
[0405] The LFNST index information is received as syntax information, and the syntax information is received as a binary bin string including 0 and 1.
[0406] The syntax element of the LFNST index according to the present embodiment may indicate whether to apply inverse LFNST or inverse non-separable transform and any one of the transform kernel matrices included in the transform set, and when the transform set includes two transform kernel matrices, the syntax element of the transform index may have three values.
[0407] That is, according to an embodiment, the value of the syntax element of the LFNST index may include: 0, which indicates that inverse LFNST is not applied to the target block; 1, which indicates the first transform kernel matrix among the transform kernel matrices; and 2, which indicates the second transform kernel matrix among the transform kernel matrices.
[0408] The decoding device 200 may decode information on the quantized transform coefficients of the current block from the bitstream and may derive the quantized transform coefficients of the target block based on the information on the quantized transform coefficients of the current block. The information on the quantized transform coefficients of the target block may be included in the sequence parameter set (SPS) or the slice header and may include at least one of information on whether to apply RST, information on a reduction factor, information on the minimum transform size for applying RST, information on the maximum transform size for applying RST, inverse RST size, and information on a transform index indicating any one of the transform kernel matrices included in the transform set.
[0409] The decoding device 200 may derive transform coefficients by dequantizing the residual information (i.e., quantized transform coefficients) regarding the current block, and may arrange the derived transform coefficients in a predetermined scan order.
[0410] Specifically, the derived transform coefficients may be arranged in 4×4 blocks according to the inverse diagonal scan order, and the transform coefficients in the 4×4 blocks may also be arranged according to the inverse diagonal scan order. That is, the dequantized transform coefficients may be arranged according to the inverse scan order applied in a video codec (such as in VVC or HEVC).
[0411] The transform coefficients derived based on the residual information may be the dequantized transform coefficients as described above, or may be quantized transform coefficients. That is, the transform coefficients may be any data for checking whether there is non-zero data in the current block, regardless of quantization.
[0412] The decoding device may update the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block when the intra prediction mode of the chrominance block is the CCLM mode. In particular, when the intra prediction mode of the luminance block is the MIP mode, the intra prediction mode of the chrominance block may be updated to the intra planar mode (S1520).
[0413] The decoding device may derive the intra prediction mode of the chrominance block as the CCLM mode based on the intra prediction mode information. For example, the decoding device may receive information about the intra prediction mode of the current chrominance block through a bitstream, and may derive the intra prediction mode of the current chrominance block as the CCLM mode based on the intra prediction mode information.
[0414] The CCLM mode may include the upper left CCLM mode, the upper CCLM mode, or the left CCLM mode.
[0415] As described above, the decoding device may derive residual samples by applying the LFNST as an inseparable transform or the MTS as a separable transform, and may perform these transforms based on the LFNST index indicating the LFNST kernel (i.e., the LFNST matrix) and the MTS index indicating the MTS kernel, respectively.
[0416] For the LFNST, it is necessary to determine the LFNST set, and the LFNST set has a mapping relationship with the intra prediction mode of the current block.
[0417] The decoding device may update the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block for the inverse LFNST of the chrominance block.
[0418] According to the example, the updated intra prediction mode can be derived as the intra prediction mode corresponding to a specific position in the luma block, and the specific position can be set based on the color format of the chroma block.
[0419] The specific position can be the central position of the luma block and can be represented by ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)).
[0420] In the central position, xTbY and yTbY represent the upper left coordinates of the luma block, i.e., the upper left position in the luma sample reference of the current transform block, nTbW and nTbH represent the width and height of the chroma block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)) represents the central position of the luma transform block, and IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra prediction mode of the luma block at this position.
[0421] SubWidthC and SubHeightC can be derived as shown in Table 18. That is, when the color format is 4:2:0, SubWidthC and SubHeighC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0422] As shown in Table 17, in order to specify the specific position of the luma block corresponding to the chroma block regardless of the color format, the color format is reflected in the variable indicating the specific position.
[0423] As described above, when the intra prediction mode corresponding to the specific position of the luma block is the matrix-based intra prediction (hereinafter, "MIP") mode, the decoding device can set the updated intra prediction mode to the intra plane mode.
[0424] The MIP mode can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). When MIP is applied to the current block, the predicted samples of the current block can be derived by i) using adjacent reference samples that have undergone averaging processing, ii) performing matrix-vector multiplication processing, and iii) further performing horizontal / vertical interpolation processing.
[0425] Alternatively, according to the example, when the intra prediction mode corresponding to a specific position is the Intra Block Copy (IBC) mode or the palette mode, the decoding device may set the updated intra prediction mode to the intra DC mode.
[0426] The IBC prediction mode or the palette mode can be used to encode content images / videos including games, such as Screen Content Coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction, except that the reference block is derived within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, the values of the samples in the picture can be signaled based on the information about the palette table and the palette index.
[0427] In summary, when the intra prediction mode at the central position is the MIP mode, the IBC mode, or the palette mode, the intra prediction mode of the chrominance block can be updated to a specific mode, such as the intra planar mode or the intra DC mode.
[0428] When the intra prediction mode at the central position is not the MIP mode, the IBC mode, and the palette mode, the intra prediction mode of the chrominance block can be updated to the intra prediction mode of the luma block at the central position in order to reflect the correlation between the chrominance block and the luma block.
[0429] The decoding device may determine an LFNST set including the LFNST matrix based on the updated intra prediction mode (S1530), and may derive the transform coefficients of the chrominance block based on the LFNST matrix derived from the LFNST set (S1540).
[0430] Any one of the multiple LFNST matrices can be selected based on the LFNST set and the LFNST index.
[0431] As shown in Table 20, the LFNST transform set is derived according to the intra prediction mode, and 81 to 83 indicating the CCLM mode in the intra prediction mode are omitted because the LFNST transform set is derived using the intra mode value of the corresponding luma block in the CCLM mode.
[0432] According to the example, as shown in Table 20, any one of the four LFNST sets can be determined according to the intra prediction mode of the current block, and the LFNST set to be applied to the current chrominance block can also be determined.
[0433] The decoding device can perform an inverse RST (e.g., inverse LFNST) by applying the LFNST matrix to the dequantized transform coefficients, thereby deriving the modified transform coefficients of the current chrominance block.
[0434] The decoding device may derive a residual sample from the transform coefficients by a single inverse transform (S1550), and when the current block is a chrominance block, may derive the residual sample of the chrominance block based on the transform coefficients. MTS may be used for the single inverse transform.
[0435] In addition, the decoding device may generate a reconstructed sample based on the residual sample of the current block and the prediction sample of the current block.
[0436] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices illustrated in the drawings or the names of specific signals / messages / fields are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0437] Figure 16 is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.
[0438] Figure 16 Each process disclosed in is based on some details described with reference to Figures 3 to 14 Therefore, specific details overlapping with those described with reference to Figure 1 and Figures 3 to 14 will be omitted or described schematically.
[0439] According to an embodiment, the encoding device 100 may derive a prediction sample of a chrominance block based on the fact that the intra prediction mode of the chrominance block is the CCLM mode (S1610).
[0440] The encoding device may first derive the intra prediction mode of the chrominance block as the CCLM mode.
[0441] For example, the encoding device may determine the intra prediction mode of the current chrominance block based on the rate-distortion (RD) cost (or RDO). Here, the RD cost may be derived based on the sum of absolute differences (SAD). The encoding device may determine the CCLM mode as the intra prediction mode of the current chrominance block based on the RD cost.
[0442] The CCLM mode may include a top-left CCLM mode, a top CCLM mode, or a left CCLM mode.
[0443] The encoding device may encode information about the intra prediction mode of the current chrominance block and may signal the information about the intra prediction mode through a bitstream. The prediction-related information about the current chrominance block may include the information about the intra prediction mode.
[0444] According to an embodiment, the encoding device may derive a residual sample of a chrominance block based on the prediction sample (S1620).
[0445] According to an embodiment, an encoding device may derive transform coefficients of a chrominance block based on a single transformation of residual samples.
[0446] The single transformation may be performed by a plurality of transform kernels, and in this case, a transform kernel may be selected based on an intra prediction mode.
[0447] The encoding device may update the intra prediction mode of the chrominance block for LFNST of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block, and may update the intra prediction mode of the chrominance block to an intra planar mode when the intra prediction mode of the luminance block is a MIP mode (S1630).
[0448] As shown in Table 17, the encoding device may update the CCLM mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block (when predModeIntra is equal to INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM, derive predModeIntra as follows).
[0449] According to an example, the updated intra prediction mode may be derived as an intra prediction mode corresponding to a specific position in the luminance block, and the specific position may be set based on the color format of the chrominance block.
[0450] The specific position may be a central position of the luminance block, and may be represented by ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)).
[0451] In the central position, xTbY and yTbY represent the upper left coordinates of the luminance block, that is, the upper left position in the luminance sample reference of the current transform block, nTbW and nTbH represent the width and height of the chrominance block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)) represents the central position of the luminance transform block, and IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra prediction mode of the luminance block at this position.
[0452] SubWidthC and SubHeightC may be derived as shown in Table 18. That is, when the color format is 4:2:0, SubWidthC and SubHeighC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0453] As shown in Table 17, in order to specify a specific position of a luminance block corresponding to a chrominance block regardless of a color format, the color format is reflected in a variable indicating the specific position.
[0454] As described above, when an intra prediction mode of a luminance block corresponding to a specific position is a matrix-based intra prediction (hereinafter, "MIP") mode, an encoding device may set an updated intra prediction mode to an intra planar mode.
[0455] The MIP mode may be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). When MIP is applied to a current block, prediction samples of the current block may be derived by i) using adjacent reference samples that have undergone averaging processing, ii) performing matrix-vector multiplication processing, and iii) further performing horizontal / vertical interpolation processing.
[0456] Alternatively, according to an example, when an intra prediction mode corresponding to a specific position is an intra block copy (IBC) mode or a palette mode, a decoding device may set an updated intra prediction mode to an intra DC mode.
[0457] The IBC prediction mode or the palette mode may be used to encode a content image / video including a game, such as screen content coding (SCC). The IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction, except that a reference block is derived within the current picture. That is, the IBC may use at least one of the inter prediction techniques described in the present disclosure. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, values of samples in a picture may be signaled based on information about a palette table and a palette index.
[0458] In summary, when an intra prediction mode at a central position is an MIP mode, an IBC mode, or a palette mode, an intra prediction mode of a chrominance block may be updated to a specific mode, such as an intra planar mode or an intra DC mode.
[0459] When an intra prediction mode at a central position is not an MIP mode, an IBC mode, and a palette mode, an intra prediction mode of a chrominance block may be updated to an intra prediction mode of a luminance block at the central position to reflect an association between the chrominance block and the luminance block.
[0460] An encoding device may determine an LFNST set including an LFNST matrix based on the updated intra prediction mode (S1640), and may derive transform coefficients of a chrominance block based on residual samples and the LFNST matrix (S1650).
[0461] The encoding device may determine a transform set based on a mapping relationship according to the intra prediction mode applied to the current block, and may perform LFNST (i.e., non-separable transform) based on any one of the two LFNST matrices included in the transform set.
[0462] As described above, multiple transform sets may be determined according to the intra prediction mode of the transform block to be transformed. The matrix applied to LFNST is the transpose of the matrix used in inverse LFNST.
[0463] In one example, the LFNST matrix may be a non-square matrix with the number of rows less than the number of columns.
[0464] The encoding device may derive quantized transform coefficients by performing quantization based on the modified transform coefficients of the current chroma block, and may encode and output image information including information about the quantized transform coefficients, information about the intra prediction mode, and an LFNST index indicating the LFNST matrix (S1660).
[0465] Specifically, the encoding device 100 may generate information about the quantized transform coefficients and may encode the generated information about the quantized transform coefficients.
[0466] In one example, the information about the quantized transform coefficients may include at least one of information about whether LFNST is applied, information about a reduction factor, information about the minimum transform size for applying LFNST, and information about the maximum transform size for applying LFNST.
[0467] The encoding device may encode flag information indicating whether CCLM is applied to the current block as sps_cclm_enabled_flag and information about the intra prediction mode of the chroma component as intra_chroma_pred_mode as information about the intra mode.
[0468] The information about the CCLM mode as intra_chroma_pred_mode may indicate the top-left CCLM mode, the top CCLM mode, or the left CCLM mode.
[0469] In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of consistent expression.
[0470] In addition, in the present disclosure, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled via the residual coding syntax. The transform coefficients may be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients may be derived by an inverse transform (scaling) of the transform coefficients. The residual samples may be derived based on an inverse transform (transformation) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.
[0471] In the above-described embodiments, the method is explained based on a flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may be performed in an order or steps different from the above order or steps, or a certain step may be performed concurrently with other steps. In addition, those of ordinary skill in the art will understand that the steps shown in the flowchart are not exclusive, and one or more steps in the flowchart may be incorporated or deleted without affecting the scope of the present disclosure.
[0472] The above method according to the present disclosure may be implemented in software form, and the encoding device and / or decoding device according to the present disclosure may be included in devices for image processing such as televisions, computers, smart phones, set-top boxes, and display devices.
[0473] When the embodiments in the present disclosure are implemented by software, the above method may be implemented as modules (steps, functions, etc.) for performing the above functions. These modules may be stored in a memory and may be executed by a processor. The memory may be inside or outside the processor and may be connected to the processor in various well-known ways. The processor may include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure may be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing may be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.
[0474] In addition, the decoding device and encoding device applying the present disclosure may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device (such as video communication), a mobile streaming device, a storage medium, a camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0475] In addition, the processing method applying the present disclosure may be produced in the form of a program executed by a computer and may be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure may also be stored in the computer-readable recording medium. The computer-readable recording medium includes various storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, the bitstream generated by the encoding method may be stored in the computer-readable recording medium or transmitted through a wired or wireless communication network. In addition, the embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer according to the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0476] Figure 17 Examples of a video / image encoding system to which the present disclosure may be applied are illustrated.
[0477] Referring to Figure 17 , the video / image encoding system may include a source device and a receiving device. The source device may transfer the encoded video / image information or data to the receiving device in the form of a file or a stream through a digital storage medium or a network.
[0478] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0479] The video source can obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet computer, and a smart phone, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer or the like. In this case, the video / image capture process can be replaced by a process of generating relevant data.
[0480] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0481] The transmitter can send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include an element for generating a media file in a predetermined file format and can include an element for sending via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to the decoding device.
[0482] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc., corresponding to the operations of the encoding device.
[0483] The renderer can render the decoded video / image. The rendered video / image can be displayed on a display.
[0484] Figure 18 Illustrated is the structure of a content streaming system to which the present disclosure is applied.
[0485] In addition, the content streaming system to which the present disclosure is applied can generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0486] The encoding server is used to compress the content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream, and send it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or the bitstream generation method of the present disclosure. And the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0487] The streaming server sends multimedia data to the user device through the web server based on the user's request. The web server serves as a tool to notify the user of what services exist. When the user requests the service the user wants, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content streaming system can include a separate control server, and in this case, the control server is used to control the commands / responses between the corresponding devices in the content streaming system.
[0488] The streaming server can receive content from the media storage device and / or the encoding server. For example, in the case of receiving content from the encoding server, the content can be received in real time. In this case, in order to smoothly provide the streaming service, the streaming server can store the bitstream for a predetermined time.
[0489] For example, the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigator, a slate PC, a tablet PC, a ultrabook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital sign, etc. Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0490] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined to be implemented or executed in a device, and the technical features of the device claims can be combined to be implemented or executed in a method. In addition, the technical features of the method claims and the device claims can be combined to be implemented or executed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or executed in a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtain intra prediction mode information from the bitstream; Based on the intra prediction mode of the chrominance block being the cross-component linear model (CCLM) mode, update the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block; Determine an LFNST set including an LFNST matrix based on the updated intra prediction mode; Derive transform coefficients for the chrominance block based on the LFNST matrix derived from the LFNST set; And Derive residual samples for the chrominance block based on the transform coefficients, wherein the updated intra prediction mode is derived as the intra prediction mode corresponding to a specific position in the luminance block, wherein based on the intra prediction type corresponding to the specific position being the MIP mode, the intra prediction mode of the chrominance block is updated to the intra planar mode, wherein the specific position is set based on the color format of the chrominance block, wherein the specific position is set to ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)), wherein xTbY and yTbY represent the upper left coordinates of the luminance block, wherein nTbW and nTbH represent the width and height of the chrominance block, and wherein SubWidthC and SubHeightC represent variables corresponding to the color format.
2. The image decoding method according to claim 1, wherein, The specific position is the central position of the luminance block.
3. The image decoding method according to claim 1, wherein, When the color format is 4:2:0, SubWidthC and SubHeightC are 2, and wherein when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
4. The image decoding method according to claim 1, wherein, When the prediction mode corresponding to the specific position is the IBC mode, the intra prediction mode of the chrominance block is updated to the intra DC mode.
5. The image decoding method according to claim 1, wherein, When the prediction mode corresponding to the specific position is the palette mode, the intra prediction mode of the chrominance block is updated to the intra DC mode.
6. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Derive prediction samples for the chrominance block based on the intra prediction mode of the chrominance block being the cross-component linear model CCLM; Derive residual samples for the chrominance block based on the prediction samples; Update the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block; Determine an LFNST set including an LFNST matrix based on the updated intra prediction mode; And Derive modified transform coefficients for the chrominance block based on the residual samples and the LFNST matrix, wherein the updated intra prediction mode is derived as the intra prediction mode corresponding to a specific position in the luminance block, wherein based on the intra prediction type corresponding to the specific position being the MIP mode, the intra prediction mode of the chrominance block is updated to the intra planar mode, wherein the specific position is set based on the color format of the chrominance block, wherein, the specific position is set to ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)), wherein, xTbY and yTbY represent the upper left coordinates of the luminance block, wherein, nTbW and nTbH represent the width and height of the chrominance block, and wherein, SubWidthC and SubHeightC represent variables corresponding to the color format.
7. The image encoding method according to claim 6, wherein, The specific position is the central position of the luminance block.
8. The image encoding method according to claim 6, wherein, When the color format is 4:2:0, SubWidthC and SubHeightC are 2, and wherein, when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
9. The image encoding method according to claim 6, wherein, When the prediction mode corresponding to the specific position is the IBC mode, the intra prediction mode of the chrominance block is updated to the intra DC mode.
10. The image encoding method according to claim 6, wherein, When the prediction mode corresponding to the specific position is the palette mode, the intra prediction mode of the chrominance block is updated to the intra DC mode.
11. A method for transmitting data for image information, the method comprising the following steps: Obtain a bitstream for the image, wherein, the bitstream is generated by: deriving prediction samples for the chrominance block based on the intra prediction mode for the chrominance block being the cross-component linear model CCLM; deriving residual samples for the chrominance block based on the prediction samples; updating the intra prediction mode of the chrominance block based on the intra prediction mode of the luminance block corresponding to the chrominance block; determining an LFNST set including an LFNST matrix based on the updated intra prediction mode; deriving modified transform coefficients for the chrominance block based on the residual samples and the LFNST matrix; and encoding the residual information to generate the bitstream; Transmit the data including the bitstream, wherein, the updated intra prediction mode is derived as the intra prediction mode corresponding to a specific position in the luminance block, wherein, based on the intra prediction type corresponding to the specific position being the MIP mode, the intra prediction mode of the chrominance block is updated to the intra planar mode, wherein, the specific position is set based on the color format of the chrominance block, wherein, the specific position is set to ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)), wherein, xTbY and yTbY represent the upper left coordinates of the luminance block, wherein, nTbW and nTbH represent the width and height of the chrominance block, and wherein, SubWidthC and SubHeightC represent variables corresponding to the color format.
Citation Information
Patent Citations
Image encoding / decoding method and apparatus, and recording medium storing bitstream
CN113545089A