Image encoding apparatus, image decoding apparatus, storage medium therefor, and transmission apparatus
Patent Information
- Application Number
- CN202310620110.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-07
- Filing Date
- 2020-10-05
- Publication Date
- 2026-08-11
- Estimated Expiration
- 2040-10-05
AI Technical Summary
因此,当使用诸如传统有线/无线宽带线这样的介质来发送图像数据或者使用现有存储介质来存储图像/视频数据时,其传输成本和存储成本增加
[0023] According to this disclosure, the overall image/video compression efficiency can be increased.
Smart Images

Figure CN116600115B_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application No. 202080079999.7 (International Application No.: PCT / KR2020 / 013459, Application Date: October 5, 2020, Invention Title: Transformation-Based Image Coding Method and Apparatus). Technical Field
[0002] This disclosure relates to an image coding technique, and more specifically, to a method and apparatus for encoding images based on transformations in an image coding system. Background Technology
[0003] Today, the demand for high-resolution and high-quality images / videos, such as 4K, 8K, or even higher Ultra High Definition (UHD) images / videos, is constantly growing across various fields. As image / video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases compared to traditional image data. Therefore, transmission and storage costs increase when using media such as traditional wired / wireless broadband lines to transmit image data or when using existing storage media to store image / video data.
[0004] In addition, there is increasing interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms, and broadcasting of images / videos with image characteristics that differ from real images such as game images is on the rise.
[0005] Therefore, there is a need for efficient image / video compression techniques to effectively compress, transmit, store, and reproduce information with high resolution and high quality images / videos that have the various characteristics described above. Summary of the Invention
[0006] Technical Purpose
[0007] One aspect of this disclosure is to provide a method and apparatus for increasing image coding efficiency.
[0008] Another technical aspect of this disclosure is to provide a method and apparatus for increasing the efficiency of transform index coding.
[0009] Another technical aspect of this disclosure is to provide an image encoding method and apparatus using LFNST and MTS.
[0010] Another technical aspect of this disclosure is to provide an image encoding method and apparatus for LFNST index and MTS index signaling.
[0011] Technical solution
[0012] According to embodiments of this disclosure, an image decoding method performed by a decoding device is provided. The method may include: performing residual coding to arrange transform coefficients of a current block according to a predetermined scan order by parsing residual information received at a residual coding level; deriving residual samples by applying at least one of LFNST or MTS to the transform coefficients; and generating a reconstructed image based on the residual samples, wherein LFNST may be performed based on an LFNST index indicating an LFNST kernel, and MTS may be performed based on an MTS index indicating an MTS kernel, the LFNST index and the MTS index may be signaled at the coding unit level, and the MTS index may be signaled immediately after the LFNST index is signaled.
[0013] When the tree type of the current block is a single tree, the LFNST index can be parsed after performing residual encoding of the luma and chroma blocks for the current block.
[0014] When the tree type of the current block is a dual-tree type and the chroma components are encoded, the LFNST index can be parsed after performing residual encoding of the Cb and Cr components of the chroma block.
[0015] When the current block is partitioned into multiple sub-partition blocks, the LFNST index can be resolved after performing residual encoding for the multiple sub-partition blocks.
[0016] When the current block is partitioned into multiple sub-blocks, the LFNST index can be resolved regardless of whether there are transform coefficients in the regions other than the DC location in each of the multiple sub-blocks.
[0017] Performing residual encoding may include: deriving a first variable indicating whether a transform coefficient exists in a region other than the DC location in the current block; and deriving a second variable indicating whether a transform coefficient exists in a second region other than the first region in the upper left of the current block or a sub-block partitioned by the current block, and the LFNST index can be parsed when a transform coefficient exists in a region other than the DC location and no transform coefficient exists in the second region.
[0018] Performing residual encoding may include: deriving an indication of whether a third variable of transform coefficients exists in the current block outside the top-left 16×16 region, and resolving the MTS index when no transform coefficients exist in the region outside the 16×16 region.
[0019] According to another embodiment of this disclosure, an image encoding method performed by an encoding device is provided. The method may include: performing a transform coefficient derivation operation by applying at least one of LFNST or MTS to residual samples to derive transform coefficients of the current block and arranging the transform coefficients according to a predetermined scan order; encoding at least one of an LFNST index indicating an LFNST kernel or an MTS index indicating an MTS kernel; and constructing and outputting image information such that the LFNST index and the MTS index are signaled at the encoding unit level, and the MTS index is signaled immediately after the LFNST index is signaled.
[0020] According to another embodiment of the present disclosure, a digital storage medium may be provided that stores image data including a bitstream and encoded image information generated according to an image encoding method performed by an encoding device.
[0021] According to another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including encoded image information and bitstreams to enable a decoding device to perform an image decoding method.
[0022] Technical effect
[0023] According to this disclosure, the overall image / video compression efficiency can be increased.
[0024] According to this disclosure, the efficiency of transformation index encoding can be increased.
[0025] Another technical aspect of this disclosure provides an image encoding method and apparatus using LFNST and MTS.
[0026] Another technical aspect of this disclosure provides an image encoding method and apparatus for LFNST indexing and MTS indexing signaling.
[0027] The effects achievable through the specific examples of this disclosure are not limited to those listed above. For example, various technical effects may exist that can be understood or derived from this disclosure by one of ordinary skill in the art. Therefore, the specific effects of this disclosure are not limited to those expressly described herein, but may include various effects that can be understood or derived from the technical features of this disclosure. Attached Figure Description
[0028] Figure 1 This is a diagram illustrating the configuration of a video / image encoding apparatus to which embodiments of the present disclosure can be applied.
[0029] Figure 2 This is a diagram illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.
[0030] Figure 3 Multiple transformation techniques according to embodiments of the present disclosure are illustrated schematically.
[0031] Figure 4 The intra-frame orientation patterns for 65 predicted directions are schematically shown.
[0032] Figure 5 This is a diagram illustrating an embodiment of the RST according to the present disclosure.
[0033] Figure 6 This is a diagram illustrating the order in which the output data of a forward first transformation is arranged into a one-dimensional vector, based on the example.
[0034] Figure 7 This is a diagram illustrating the order in which the output data of the forward quadratic transform is arranged into two-dimensional blocks, based on the example.
[0035] Figure 8 This is a diagram illustrating the block shape to which LFNST is applied.
[0036] Figure 9 This is a diagram illustrating the arrangement of the output data of the positive LFNST according to the example, and it shows blocks in which the output data of the positive LFNST is arranged according to the example.
[0037] Figure 10 A graph is shown where the amount of output data for the positive LFNST is limited to a maximum of 16, according to the example.
[0038] Figure 11 This is a diagram illustrating the zeroing process in a block of 4×4 LFNST, based on the example.
[0039] Figure 12 This is a diagram illustrating the zeroing process in a block of 8×8 LFNST, based on the example.
[0040] Figure 13 This is a diagram illustrating the zeroing of a block in an 8×8 LFNST application, based on another example.
[0041] Figure 14 This is a diagram illustrating an example of a sub-block segmented from a coded block.
[0042] Figure 15 This is a diagram illustrating another example of sub-blocks segmented from a coded block.
[0043] Figure 16 It is a diagram illustrating the symmetry between the M×2 (M×1) block and the 2×M (1×M) block according to the example.
[0044] Figure 17 This is a diagram illustrating a 2×M block transposed according to the example.
[0045] Figure 18 The scanning order of the 8×2 or 2×8 region is illustrated based on the example.
[0046] Figure 19 This is a diagram used to illustrate the method of decoding an image according to an example.
[0047] Figure 20 This is a diagram used to illustrate the method of encoding images according to the example.
[0048] Figure 21 Examples of video / image coding systems to which this disclosure can be applied are illustrated schematically.
[0049] Figure 22 The structure of a content streaming system applying this disclosure is illustrated. Detailed Implementation
[0050] While this disclosure may be readily modified and includes various embodiments, specific embodiments thereof have been illustrated by way of example in the accompanying drawings and will now be described in detail. However, this is not intended to limit this disclosure to the specific embodiments disclosed herein. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the technical concept of this disclosure. The singular form may include the plural form unless the context clearly indicates otherwise. Terms such as “comprising” and “having” are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and should therefore not be construed as pre-excluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0051] Furthermore, for ease of description of their different features and functions, the components in the accompanying drawings described herein are illustrated independently; however, this does not imply that each component is implemented by a separate piece of hardware or software. For example, any two or more of these components may be combined to form a single component, and any single component may be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of this disclosure, provided they do not depart from the spirit of this disclosure.
[0052] In the following description, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Furthermore, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.
[0053] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Video Coding Universal) standard (ITU-T Rec. H.266), the next generation video / image coding standard after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), EVC (Essential Video Coding) standard, AVS2 standard, etc.).
[0054] This document provides various implementations related to video / image encoding, and these implementations may be combined and performed in combination with each other unless otherwise specified.
[0055] In this document, video can refer to a collection of images over a period of time. Typically, an image is a unit representing a specific time region, while a strip / patch is a unit that constitutes a part of an image. A strip / patch can include one or more coding tree units (CTUs). An image can consist of one or more strips / patches. An image can consist of one or more patch groups. A patch group can include one or more patches.
[0056] A pixel or primitive (pel) can refer to the smallest unit that makes up a picture (or image). Alternatively, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when the pixel value is transformed to the frequency domain, it can refer to the transform coefficients in the frequency domain.
[0057] A unit can represent the basic unit of image processing. A unit may include a specific region and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the context, units and terms such as blocks and regions may be used interchangeably. Typically, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0058] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Furthermore, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A / B / C" can mean "at least one of A, B, and / or C".
[0059] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" could include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0060] In this disclosure, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0061] Furthermore, in this disclosure, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0062] Additionally, the parentheses used in this disclosure can indicate "for example". Specifically, when indicated as "prediction (intra-frame prediction)", it can mean that "intra-frame prediction" is proposed as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" is proposed as an example of "prediction". Furthermore, when indicated as "prediction (i.e., intra-frame prediction)", this can also mean that "intra-frame prediction" is proposed as an example of "prediction".
[0063] The technical features described individually in one of the accompanying drawings of this disclosure may be implemented individually or simultaneously.
[0064] Figure 1 This diagram schematically illustrates the configuration of a video / image encoding apparatus to which embodiments of the present disclosure may be applied. In the following, the term "video encoding apparatus" may include an image encoding apparatus.
[0065] Reference Figure 1The encoding device 100 may include and be configured with an image segmenter 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 may include an inter-frame predictor 121 and an intra-frame predictor 122. The residual processor 130 may include a transformer 132, a quantizer 133, a dequantizer 134, and an inverse transformer 135. The residual processor 130 may further include a subtractor 131. The adder 150 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 110, predictor 120, residual processor 130, entropy encoder 140, adder 150, and filter 160 described above may be constituted by one or more hardware components (e.g., an encoder chipset or a processor). Furthermore, the memory 170 may include a decoded picture buffer (DPB) and may also be constituted by a digital storage medium. The hardware components may further include the memory 170 as an internal / external component.
[0066] Image segmenter 110 can segment an input image (or picture, frame) input to encoding device 100 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit can be recursively segmented from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be segmented into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The encoding process according to this document can be performed based on the final coding unit that is no longer segmented. In this case, based on encoding efficiency according to image characteristics, etc., the LCU can be directly used as the final coding unit, or alternatively, the coding unit can be recursively segmented into deeper coding units such that a coding unit of optimal size can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, as described later. As another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, each of the PU and TU may be segmented or partitioned from the aforementioned final encoding unit. The PU may be a unit for sample prediction, and the TU may be a unit for inducing transform coefficients and / or for inducing residual signals from the transform coefficients.
[0067] In some cases, a unit can be used interchangeably with terms such as block or region. Typically, an M×N block can represent a sample consisting of M columns and N rows, or a set of transform coefficients. A sample can generally represent a pixel or pixel value, and can also represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. A sample can be used as an item corresponding to the pixels or primitives of a picture (or image).
[0068] Encoding device 100 generates a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 121 or intra-frame predictor 122 from the input image signal (original block, original sample array), and sends the generated residual signal to converter 132. In this case, as explained, the unit within encoder 100 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as subtractor 131. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The predictor can generate various information about the prediction, such as prediction mode information, to transmit the generated information to entropy encoder 140, as described later in the description of each prediction mode. The information about the prediction can be encoded by entropy encoder 140 and can be output in the form of a bitstream.
[0069] Intra-predictor 122 can refer to samples within the current image to predict the current block. Depending on the prediction mode, the referenced samples may be located adjacent to the current block or located far from the current block. The prediction modes in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC modes or planar modes. Depending on the granularity of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is exemplary, and more or fewer directional prediction modes than the above numbers can be used depending on the settings. Intra-predictor 122 can also use prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.
[0070] Inter-frame predictor 121 can induce the prediction block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between the motion information of neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially adjacent blocks existing within the current image and temporally adjacent blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally adjacent block may be the same as each other, or they may be different from each other. The temporally adjacent block may be referred to by a name, such as juxtaposed reference block, juxtaposed CU (col CU), or similar, and the reference image including the temporally adjacent block may also be referred to by a name such as juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally adjacent block may be referred to as the juxtaposed image (colPic). For example, the inter-frame predictor 121 can configure a candidate list of motion information based on neighboring blocks and generate information indicating which candidates are used to derive the motion vector of the current block and / or the reference image index. Inter-frame prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the inter-frame predictor 121 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, unlike the merge mode, residual signals may not be sent. The motion vector prediction (MVP) mode can predict the motion vector of the current block by using the motion vectors of neighboring blocks as motion vectors and indicating the motion vector of the current block by signaling the motion vector difference.
[0071] Predictor 120 can generate prediction signals based on various prediction methods described later. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction to predict a block, but also both intra-frame and inter-frame prediction simultaneously. This can be referred to as Combined Inter-Frame and Intra-Frame Prediction (CIIP). Furthermore, the predictor can perform prediction on blocks based on an Intra-Frame Block Copy (IBC) prediction mode or a palette mode. IBC prediction modes or palette modes can be used for content image / video coding, such as Screen Content Coding (SCC), in games, etc. IBC essentially performs prediction within the current frame, but it can be similarly performed for inter-frame prediction because it derives a reference block in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values in the image can be signaled based on information about the palette index and palette table.
[0072] The predicted signal generated by the predictor (including inter-frame predictor 121 and / or intra-frame predictor 122) can be used to generate a reconstructed signal or a residual signal. Transformer 132 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen–Loève Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, GBT represents the transform obtained from the graph. CNT represents the transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform process can be applied to pixel blocks that are squares of the same size, and also to blocks that are not squares but have variable sizes.
[0073] Quantizer 133 can quantize the transform coefficients to send them to entropy encoder 140, and entropy encoder 140 can encode the quantized signal (information about the quantized transform coefficients) to output an encoded quantized signal in bitstream form. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 133 can rearrange the quantized transform coefficients in block form as a one-dimensional vector based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients. Entropy encoder 140 can perform various encoding methods, such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 140 can also encode information required for video / image reconstruction (e.g., values of syntax elements, etc.) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at network abstraction layers (NALs). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may also include general constraint information. In this document, syntax elements and / or information transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded and included in a bitstream through the encoding process described above. The bitstream may be transmitted over a network or stored in a digital storage medium. In this document, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting the signal output from the entropy encoder 140 and / or a storage device (not shown) for storing the signal may be configured as internal / external components of the encoding device 100, or the transmitter may also be included in the entropy encoder 140.
[0074] The quantization transform coefficients output from quantizer 133 can be used to generate a prediction signal. For example, dequantizer 134 and inverse transformer 135 can apply dequantization and inverse transform to the quantization transform coefficients to recover the residual signal (residual block or residual sample). Adder 150 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 121 or intra-frame predictor 122 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If no residual exists in the block to be processed when skip mode is applied, the predicted block can be used as the reconstructed block. Adder 150 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can also be used for inter-frame prediction of the next image by filtering, as described below.
[0075] In addition, Luminance Mapping and Chromaticity Scaling (LMCS) can also be applied during image encoding and / or reconstruction.
[0076] Filter 160 can apply filtering to the reconstructed signal, thereby improving subjective / objective image quality. For example, filter 160 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and store the modified reconstructed image in memory 170, specifically in the DPB of memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information to transmit the generated information to entropy encoder 140, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 140 and can be output as a bitstream.
[0077] The modified reconstructed image sent to memory 170 can be used as a reference image in inter-frame predictor 121. If inter-frame prediction is applied by the inter-frame predictor, the encoding device can avoid prediction mismatch between encoding device 100 and decoding device, and the encoding efficiency can also be improved.
[0078] The DPB of memory 170 can store a modified reconstructed image to be used as a reference image in inter-frame predictor 121. Memory 170 can store motion information of blocks in which motion information within the current image is derived (or encoded) and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 121 as motion information for spatially adjacent blocks or temporally adjacent blocks. Memory 170 can store reconstructed samples of reconstructed blocks within the current image and transmit the reconstructed samples to intra-frame predictor 122.
[0079] Figure 2This is a diagram illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.
[0080] Reference Figure 2 The decoding device 200 may include and be configured with an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The predictor 230 may include an inter-frame predictor 232 and an intra-frame predictor 231. The residual processor 220 may include a dequantizer 221 and an inverse transformer 222. According to embodiments, the entropy decoder 210, residual processor 220, predictor 230, adder 240, and filter 250 described above may be configured by one or more hardware components (e.g., a decoder chipset or processor). Furthermore, the memory 260 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 260 as an internal / external component.
[0081] When an input bitstream includes video / image information, the decoding device 200 can respond to the input bitstream containing... Figure 1 The process of processing video / image information in the encoding device shown reconstructs the image. For example, the decoding device 200 can derive units / blocks based on block segmentation information obtained from the bitstream. The decoding device 200 can perform decoding using processing units applied to the encoding device. Therefore, the processing unit used for decoding can be, for example, an encoding unit, and the encoding unit can be segmented from a CTU or LCU according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the encoding unit. In addition, the reconstructed image signal decoded and output by the decoding device 200 can be reproduced by a reproduction device.
[0082] Decoding device 200 can receive data in bitstream form from... Figure 1The signal output by the encoding device shown, and the received signal, can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can deduce the information required for image reconstruction (or picture reconstruction) (e.g., video / image information) by parsing the bitstream. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information to be signaled / received and / or syntax elements (which will be described later herein) can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 210 can decode the information within the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements required for image reconstruction, as well as the quantized values of the residual correlation transform coefficients. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the syntax elements to be decoded, as well as decoding information of adjacent blocks and the block to be decoded, or information about symbols / bins decoded in previous steps, and generate symbols corresponding to the values of each syntax element by predicting bin generation probabilities based on the determined context model and performing arithmetic decoding on the bins. At this point, the CABAC entropy decoding method can determine the context model and then update the context model using information about the decoded symbols / bins for the next symbol / bin. Prediction-related information from the information decoded by entropy decoder 210 can be provided to predictors (inter-frame predictor 232 and intra-frame predictor 231), and the residual values from entropy decoding performed by entropy decoder 210, i.e., quantization transform coefficients and related parameter information, can be input to residual processor 220. Residual processor 220 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering-related information from the information decoded by entropy decoder 210 can be provided to filter 250. Furthermore, the receiver (not shown) for receiving the signal output from the encoding device can also be configured as an internal / external component of the decoding device 200, or the receiver can also be a component of the entropy decoder 210. Additionally, the decoding device according to this disclosure can be referred to as a video / image / picture decoding device, and the decoding device can also be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 210, and the sample decoder may include at least one of a dequantizer 221, an inverse transformer 222, an adder 240, a filter 250, a memory 260, an inter-frame predictor 232, and an intra-frame predictor 231.
[0083] Dequantizer 221 can dequantize the quantized transform coefficients to output transform coefficients. Dequantizer 221 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. Dequantizer 221 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0084] The inverse transformer 222 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0085] Predictor 230 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from entropy decoder 210, and determine a specific intra-frame / inter-frame prediction mode.
[0086] The predictor can generate a prediction signal based on various prediction methods described later. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction to the prediction of a block, but also simultaneous intra-frame prediction and inter-frame prediction. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Furthermore, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. IBC prediction modes or palette modes can be used for content image / video coding, such as Screen Content Coding (SCC), in games, etc. IBC essentially performs prediction within the current frame, but it can be similarly performed for inter-frame prediction because it derives a reference block in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. A palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be signaled by including it in the video / image information.
[0087] The intra-predictor 231 can refer to samples within the current image to predict the current block. Depending on the prediction mode, the referenced samples may be located adjacent to the current block or located far from the current block. The prediction modes in intra-prediction can include multiple non-directional modes and multiple directional modes. The intra-predictor 231 can also use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.
[0088] Inter-frame predictor 232 can induce the prediction block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially adjacent blocks existing within the current image and temporally adjacent blocks existing in the reference image. For example, inter-frame predictor 232 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.
[0089] Adder 240 can add the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 232 and / or intra-frame predictor 231) to generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array). If the block to be processed has no residual when the skip mode is applied, the prediction block can be used as the reconstruction block.
[0090] Adder 240 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and as described below, it can also be output by filtering or used for inter-frame prediction of the next image.
[0091] In addition, Luminance Mapping and Chromaticity Scaling (LMCS) can also be applied during image decoding.
[0092] Filter 250 can apply filtering to the reconstructed signal, thereby improving the subjective / objective image quality. For example, filter 250 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and send the modified reconstructed image to memory 260, specifically, the DPB of memory 260. Various filtering methods may include, for example, unblocking filtering, sample adaptive shifting, adaptive loop filtering, bidirectional filtering, etc.
[0093] The (modified) reconstructed image stored in the DPB of memory 260 can be used as a reference image in inter-frame predictor 232. Memory 260 can store motion information of blocks in which motion information within the current image is derived (decoded) and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 232 so that it can be used as motion information for spatially adjacent blocks or temporally adjacent blocks. Memory 260 can store reconstructed samples of reconstructed blocks within the current image and transmit the stored reconstructed samples to intra-frame predictor 231.
[0094] In this document, the exemplary embodiments described in the encoding device 100 filter 160, inter-frame predictor 121 and intra-frame predictor 122 can be equally applied to the decoding device 200 filter 250, inter-frame predictor 232 and intra-frame predictor 231.
[0095] As described above, prediction is performed to improve compression efficiency during video encoding. Accordingly, a prediction block can be generated that includes prediction samples for the current block, which is the target block for encoding. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in both the encoding and decoding devices, and the encoding device can improve image encoding efficiency by signaling to the decoding device information about the residual between the original block and the prediction block (residual information), not the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed image including the reconstructed block.
[0096] Residual information can be generated through transformation and quantization processes. For example, an encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients. This allows it to signal the associated residual information to the decoding device (via a bitstream). Here, the residual information can include the value information, position information, transform technique, transform kernel, quantization parameters, etc., of the quantized transform coefficients. The decoding device can perform quantization / dequantization processes based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive the residual block by performing dequantization / inverse transform on the quantized transform coefficients to serve as a reference for inter-frame prediction of the next image, and can generate a reconstructed image based on this.
[0097] Figure 3 Multiple transformation techniques according to embodiments of the present disclosure are illustrated schematically.
[0098] Reference Figure 3 The converter can correspond to the aforementioned Figure 1 The converter in the encoding device, and the inverse converter can correspond to the aforementioned Figure 1 Inverse converter in encoding devices, or Figure 2 The inverse converter in the decoding device.
[0099] The transformer can derive (first) transform coefficients (S410) by performing a first transform based on residual samples (residual sample array) in the residual block. This first transform can be referred to as the core transform. In this paper, the first transform can be based on multiple transform selection (MTS), and when multiple transforms are used as a first transform, it can be referred to as a multi-core transform.
[0100] Multi-core transform can represent a method of performing transforms by additionally using Discrete Cosine Transform (DCT) Type 2 and Discrete Sine Transform (DST) Type 7, DCT Type 8, and / or DST Type 1. In other words, multi-core transform can represent a method of transforming a spatial domain residual signal (or residual block) into frequency domain transform coefficients (or primary transform coefficients) based on multiple transform kernels selected from DCT Type 2, DST Type 7, DCT Type 8, and DST Type 1. In this paper, from the perspective of the transformer, primary transform coefficients can be referred to as temporary transform coefficients.
[0101] In other words, when applying conventional transform methods, transform coefficients can be generated by applying a spatial-to-frequency domain transform to the residual signal (or residual block) based on DCT type 2. In contrast, when applying multi-core transforms, transform coefficients (or single-stage transform coefficients) can be generated by applying a spatial-to-frequency domain transform to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this paper, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transform types, transform kernels, or transform cores. These DCT / DST transform types can be defined based on basis functions.
[0102] When performing a multi-core transform, a vertical transform kernel and a horizontal transform kernel can be selected from the transform kernels for the target block. A vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can indicate the transform of the horizontal components of the target block, and the vertical transform can indicate the transform of the vertical components of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block), including the residual block.
[0103] Furthermore, according to the example, if a transformation is performed by applying an MTS, the mapping relationship of the transformation kernels can be set by setting specific basis functions to predetermined values and combining the basis functions to be applied in the vertical or horizontal transformation. For example, when the horizontal transformation kernel is denoted as trTypeHor and the vertical transformation kernel is denoted as trTypeVer, a value of 0 for trTypeHor or trTypeVer can be set to DCT2, a value of 1 for trTypeHor or trTypeVer can be set to DST7, and a value of 2 for trTypeHor or trTypeVer can be set to DCT8.
[0104] In this scenario, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transform cores. For example, MTS index 0 can indicate that both trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that trTypeHor is 2 and trTypeVer is 1, MTS index 3 can indicate that trTypeHor is 1 and trTypeVer is 2, and MTS index 4 can indicate that both trTypeHor and trTypeVer values are 2.
[0105] In one example, the transformation kernel set based on MTS index information is shown in the table below.
[0106] [Table 1]
[0107]
[0108] The transformer can perform a quadratic transformation based on the (first) transform coefficients to derive modified (second) transform coefficients (S420). A first transform is a transformation from the spatial domain to the frequency domain, while a quadratic transform refers to transforming to a more compact representation using the correlations existing between the (first) transform coefficients. Quadratic transforms can include inseparable transforms. In this case, the quadratic transform can be called an inseparable quadratic transform (NSST) or a mode-dependent inseparable quadratic transform (MDNSST). NSST can represent a transform based on an inseparable transform matrix, performing a quadratic transform on the (first) transform coefficients derived from the first transform to generate modified transform coefficients (or quadratic transform coefficients) for the residual signal. Here, based on the inseparable transform matrix, the transform can be applied first to the (first) transform coefficients without separating the vertical and horizontal transforms (or applying the horizontal / vertical transforms independently). In other words, NSST is not applied solely to (first-order) transform coefficients in the vertical and horizontal directions, but can represent, for example, a transform method that rearranges a two-dimensional signal (transform coefficients) into a one-dimensional signal through a specific predetermined direction (e.g., row-first or column-first) and then generates modified transform coefficients (or second-order transform coefficients) based on an inseparable transform matrix. For example, row-first order is for M×N blocks arranged in the order of first row, second row, ..., and Nth row, while column-first order is for M×N blocks arranged in the order of first column, second column, ..., and Mth column. NSST can be applied to the upper left region of a block containing (first-order) transform coefficients (hereinafter referred to as a transform coefficient block). For example, when both the width W and height H of the transform coefficient block are 8 or greater, an 8×8 NSST can be applied to the upper left 8×8 region of the transform coefficient block. Furthermore, when both the width (W) and height (H) of the transform coefficient block are 4 or greater, and the width (W) or height (H) of the transform coefficient block is less than 8, a 4×4 NSST can be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block. However, the implementation is not limited to this. For example, even if only the condition that the width W or height H of the transform coefficient block is 4 or greater is met, a 4×4 NSST can be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block.
[0109] Specifically, for example, if a 4×4 input block is used, the inseparable quadratic transformation can be performed as follows.
[0110] A 4×4 input block X can be represented as follows.
[0111] [Formula 1]
[0112]
[0113] If X is represented as a vector, then the vector It can be represented as follows.
[0114] [Equation 2]
[0115]
[0116] In Equation 2, the vector It is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to the row priority order.
[0117] In this case, the inseparable quadratic transformation can be calculated as follows.
[0118] [Formula 3]
[0119]
[0120] In this formula, represents the transformation coefficient vector, while T represents the 16×16 (inseparable) transformation matrix.
[0121] Using Equation 3 above, the 16×1 transformation coefficient vector can be derived. And the vector can be scanned in order (horizontal, vertical, and diagonal, etc.). Reorganize into 4×4 blocks. However, the above calculation is an example, and the hypercube-Givens transform (HyGT) and similar methods can also be used to calculate inseparable quadratic transformations in order to reduce the computational complexity of inseparable quadratic transformations.
[0122] Furthermore, in inseparable quadratic transforms, the transform kernel (or transform type) can be selected as mode-dependent. In this case, the mode can include intra-frame prediction mode and / or inter-frame prediction mode.
[0123] As described above, an inseparable quadratic transformation can be performed based on an 8×8 transformation or a 4×4 transformation determined by the width (W) and height (H) of the transform coefficient block. An 8×8 transformation is a transformation applicable to an 8×8 region contained within the transform coefficient block when both W and H are equal to or greater than 8, and this 8×8 region can be the top-left 8×8 region within the transform coefficient block. Similarly, a 4×4 transformation is a transformation applicable to a 4×4 region contained within the transform coefficient block when both W and H are equal to or greater than 4, and this 4×4 region can be the top-left 4×4 region within the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, while the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0124] Here, to select mode-dependent transform kernels, two inseparable quadratic transform kernels can be configured for each transform set of inseparable quadratic transforms for both 8×8 and 4×4 transforms, and there can be four transform sets. That is, four transform sets can be configured for 8×8 transforms, and four transform sets can be configured for 4×4 transforms. In this case, each transform set in the four transform sets for 8×8 transforms can include two 8×8 transform kernels, and each transform set in the four transform sets for 4×4 transforms can include two 4×4 transform kernels.
[0125] However, as the size of the transformation (i.e., the size of the region to which the transformation is applied) can be, for example, a size other than 8×8 or 4×4, the number of sets can be n, and the number of transformation kernels in each set can be k.
[0126] The transform set can be referred to as the NSST set or the LFNST set. A specific set within the transform set can be selected, for example, based on the intra-prediction mode of the current block (CU or sub-block). The Low-Frequency Inseparable Transform (LFNST) can be an example of a reduced inseparable transform, which will be described later, and represents an inseparable transform for low-frequency components.
[0127] For reference, for example, intra-prediction modes may include two non-directional (or non-angular) intra-prediction modes and 65 directional (or angular) intra-prediction modes. Non-directional intra-prediction modes may include planar intra-prediction mode number 0 and DC intra-prediction mode number 1, and directional intra-prediction modes may include 65 intra-prediction modes numbered 2 through 66. However, this is an example, and this document can be applied even if the number of intra-prediction modes differs. Furthermore, in some cases, intra-prediction mode number 67 may be used, and intra-prediction mode number 67 may represent a linear model (LM) mode.
[0128] Figure 4 The intra-frame orientation patterns for 65 predicted directions are schematically shown.
[0129] Reference Figure 4 Based on the intra-prediction mode 34 with a left-top diagonal prediction direction, intra-prediction modes can be divided into intra-prediction modes with horizontal directionality and intra-prediction modes with vertical directionality. Figure 4In the diagram, H and V denote horizontal and vertical orientation, respectively, and the numbers -32 to 32 indicate a displacement of 1 / 32 unit at the sample grid position. These numbers can represent the offset for the mode index value. Intra-prediction modes 2 to 33 are horizontally oriented, and intra-prediction modes 34 to 66 are vertically oriented. Strictly speaking, intra-prediction mode 34 can be considered neither horizontal nor vertical, but it can be classified as horizontally oriented when determining the transform set of the quadratic transform. This is because the input data is transposed for a vertical orientation mode symmetric to intra-prediction mode 34, and the input data alignment method for the horizontal mode is used for intra-prediction mode 34. Transposing the input data means switching the rows and columns of the two-dimensional M×N block data to N×M data. Intra-prediction modes 18 and 50 can represent the horizontal and vertical intra-prediction modes, respectively, and intra-prediction mode 2 can be called the upper-right diagonal intra-prediction mode because it has a left reference pixel and performs prediction in the upper-right direction. Similarly, intra-prediction mode 34 can be referred to as the bottom-right diagonal intra-prediction mode, while intra-prediction mode 66 can be referred to as the bottom-left diagonal intra-prediction mode.
[0130] Based on the example, four transform sets can be mapped according to the intra-frame prediction mode, as shown in the table below.
[0131] [Table 2]
[0132]
[0133] As shown in Table 2, any one of the four transform sets, i.e., lfnstTrSetIdx, can be mapped to any one of the four indices (i.e., 0 to 3) according to the intra-frame prediction mode.
[0134] When a specific set is determined to be used for an inseparable quadratic transform, one of the k transform kernels in that set can be selected using the inseparable quadratic transform index. The encoding device can derive the inseparable quadratic transform index indicating the specific transform kernel based on rate-distortion (RD) check and can signal the inseparable quadratic transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the inseparable quadratic transform index. For example, lfnst index 0 can refer to the first inseparable quadratic transform kernel, lfnst index 1 can refer to the second inseparable quadratic transform kernel, and lfnst index 2 can refer to the third inseparable quadratic transform kernel. Alternatively, lfnst index 0 can indicate that the first inseparable quadratic transform is not applied to the target block, and lfnst indexes 1 through 3 can indicate three transform kernels.
[0135] The converter can perform an inseparable quadratic transform based on the selected transform core and obtain modified (quadratic) transform coefficients. As mentioned above, the modified transform coefficients can be derived as transform coefficients quantized by a quantizer and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse converter in the encoding device.
[0136] Furthermore, as mentioned above, if the second transformation is omitted, the (first) transformation coefficients, which are the output of the first (separable) transformation, can be derived as the transformation coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0137] The inverse transformer can perform a series of processes in the reverse order of those already executed in the aforementioned transformers. The inverse transformer can receive (dequantized) transform coefficients and derive (first) transform coefficients by performing a second (inverse) transform (S350), and obtain residual blocks (residual samples) by performing a first (inverse) transform on the (first) transform coefficients (S360). In this regard, from the perspective of the inverse transformer, the first transform coefficients can be referred to as modified transform coefficients. As described above, the encoding and decoding devices can generate reconstructed blocks based on the residual blocks and the prediction blocks, and can generate reconstructed images based on the reconstructed blocks.
[0138] The decoding device may also include a second-order inverse transform application determiner (or a component for determining whether to apply the second-order inverse transform) and a second-order inverse transform determiner (or a component for determining the second-order inverse transform). The second-order inverse transform application determiner can determine whether to apply the second-order inverse transform. For example, the second-order inverse transform can be NSST, RST, or LFNST, and the second-order inverse transform application determiner can determine whether to apply the second-order inverse transform based on a second-order transform flag obtained by parsing the bitstream. In another example, the second-order inverse transform application determiner can determine whether to apply the second-order inverse transform based on the transform coefficients of the residual block.
[0139] A secondary inverse transform determiner can determine the secondary inverse transform. In this case, the secondary inverse transform determiner can determine the secondary inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra-prediction mode. In implementations, the secondary transform determination method can be determined depending on the primary transform determination method. Various combinations of primary and secondary transforms can be determined based on the intra-prediction mode. Furthermore, in the example, the secondary inverse transform determiner can determine the region where the secondary inverse transform is applied based on the size of the current block.
[0140] Furthermore, as mentioned above, if the second (inverse) transform is omitted, the (dequantized) transform coefficients can be received, a first (separable) inverse transform can be performed, and a residual block (residual sample) can be obtained. As mentioned above, the encoding and decoding devices can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed image based on the reconstructed block.
[0141] Furthermore, in this disclosure, a reduced quadratic transformation (RST) in which the size of the transformation matrix (kernel) is reduced can be applied to the concept of NSST in order to reduce the computational and storage requirements of the inseparable quadratic transformation.
[0142] Furthermore, the transform kernel, transform matrix, and coefficients constituting the transform kernel matrix described in this disclosure, i.e., kernel coefficients or matrix coefficients, can be represented in 8 bits. This is feasible in decoding and encoding devices, and compared to existing 9-bit or 10-bit representations, it reduces the amount of storage required to store the transform kernel and can reasonably accommodate performance degradation. Additionally, representing the kernel matrix in 8 bits allows for the use of smaller multipliers and is more suitable for Single Instruction Multiple Data (SIMD) instructions for optimal software implementation.
[0143] In this specification, the term "RST" can refer to a transformation performed on the residual samples of a target block based on a transformation matrix whose size is reduced according to a reduction factor. When performing a reduction transformation, the computational cost required for the transformation can be reduced due to the smaller size of the transformation matrix. In other words, RST can be used to address computational complexity issues that arise when transforming large blocks or when transforming indivisible blocks.
[0144] RST can be referred to by various terms such as reduced transform, reduced quadratic transform, reduced transform, simplified transform, and simple transform, and the names that RST can be called are not limited to the examples listed. Alternatively, since RST is performed primarily in the low-frequency region of the transform block that includes non-zero coefficients, it can be called low-frequency inseparable transform (LFNST). The transform index can be called the LFNST index.
[0145] Furthermore, when performing a second inverse transform based on RST, the inverse transformer 135 of the encoding device 100 and the inverse transformer 222 of the decoding device 200 may include: an inverse reduced second transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients; and an inverse first transformer that derives the residual samples of the target block based on the inverse first transform of the modified transform coefficients. An inverse first transform refers to the inverse transform of a first transform applied to the residuals. In this disclosure, deriving transform coefficients based on a transform can mean deriving the transform coefficients by applying a transform.
[0146] Figure 5This is a diagram illustrating an embodiment of the RST according to the present disclosure.
[0147] In this disclosure, "target block" may refer to the current block, residual block, or transform block to be encoded.
[0148] In the example RST, an N-dimensional vector can be mapped to an R-dimensional vector in another space, thus determining the reduced transformation matrix, where R is less than N. N can refer to the square of the length of the side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the reduction factor can refer to the R / N value. The reduction factor can be called a reduction factor, shrinkage factor, simplification factor, or other various terms. Furthermore, R can be called a reduction coefficient, but depending on the situation, the reduction factor can refer to R. Additionally, depending on the situation, the reduction factor can refer to the N / R value.
[0149] In this example, the reduction factor or reduction coefficient can be signaled via a bitstream, but the example is not limited to this. For instance, a predetermined value for the reduction factor or reduction coefficient can be stored in each of the encoding device 100 and the decoding device 200, and in this case, the reduction factor or reduction coefficient does not need to be signaled separately.
[0150] The size of the reduced transformation matrix, as shown in the example, can be less than N×N (the size of the regular transformation matrix) and can be limited as shown in Equation 4 below.
[0151] [Formula 4]
[0152]
[0153] Figure 5 The matrix T in the reduced transformation block shown in (a) can refer to the matrix T in Equation 4. R×N .like Figure 5 As shown in (a), when the reduced transformation matrix T R×N By multiplying by the residual sample of the target block, the transformation coefficients of the current block can be derived.
[0154] In the example, if the size of the block to which the transformation is applied is 8×8 and R=16 (i.e., R / N = 16 / 64 = 1 / 4), then according to Figure 5 The RST of (a) can be represented as the matrix operation shown in Equation 5. In this case, the storage and multiplication computations can be reduced to approximately 1 / 4 by a reduction factor.
[0155] In this disclosure, matrix operations can be understood as operations on column vectors obtained by multiplying a column vector by a matrix placed to the left of the column vector.
[0156] [Formula 5]
[0157]
[0158] In Equation 6, r1 to r 64 The residual samples of the target block can be represented, and specifically, they can be the transformation coefficients generated by applying a single transformation. As a result of the calculation in Equation 5, the transformation coefficients c of the target block can be derived. i And derive c i The process can be shown in Equation 6.
[0159] [Formula 6]
[0160]
[0161] As a result of Equation 6, the transformation coefficients c1 to c of the target block can be derived. R In other words, when R = 16, the transformation coefficients c1 to c of the target block can be derived. 16 If a conventional transform is applied instead of an RST, and a 64×64 (N×N) transform matrix is multiplied by a 64×1 (N×1) residual sample, only 16(R) transform coefficients are derived for the target block because of the application of the RST, even though 64(N) transform coefficients are derived for the target block. Since the total number of transform coefficients used for the target block is reduced from N to R, the amount of data sent from the encoding device 100 to the decoding device 200 is reduced, thus improving the transmission efficiency between the encoding device 100 and the decoding device 200.
[0162] When considering the size of the transformation matrix, the size of a regular transformation matrix is 64×64 (N×N), but the size of a reduced transformation matrix is reduced to 16×64 (R×N). Therefore, compared to performing a regular transformation, the storage usage ratio of performing an RST can be reduced. Furthermore, compared to the number of multiplications (N×N) when using a regular transformation matrix, using a reduced transformation matrix can reduce the number of multiplications (R×N) by the R / N ratio.
[0163] In the example, the transformer 132 of the encoding device 100 can derive the transform coefficients of the target block by performing a first transform and an RST-based second transform on the residual samples of the target block. These transform coefficients can be passed to the inverse transformer of the decoding device 200, and the inverse transformer 222 of the decoding device 200 can derive the modified transform coefficients based on the inverse reduced second transform (RST) for the transform coefficients, and can derive the residual samples of the target block based on the inverse first transform for the modified transform coefficients.
[0164] Based on the example inverse RST matrix T N×RIts size is N×R, which is larger than the size of the conventional inverse transformation matrix N×N, and is the same as the reduced transformation matrix T shown in Equation 4. R×N It has a transpose relationship.
[0165] Figure 5 The matrix T in the reduced inverse transform block shown in (b) t It can refer to the inverse RST matrix T N×R T (The superscript T indicates transpose). For example... Figure 5 As shown in (b), when the inverse RST matrix T N×R T Multiplying by the transform coefficients of the target block allows for the derivation of the modified transform coefficients of the target block or the residual samples of the target block. The inverse RST matrix T R×N T It can be represented as (T) R×N ) T N×R .
[0166] More specifically, when the inverse RST is used as a second inverse transformation, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. Furthermore, the inverse RST can be used as the inverse first-order transform, and in this case, when the inverse RST matrix T... N×R T When multiplied by the transformation coefficients of the target block, the residual sample of the target block can be derived.
[0167] In the example, if the size of the block to which the inverse transform is applied is 8×8 and R=16 (i.e., R / N = 16 / 64 = 1 / 4), then according to Figure 5 The RST of (b) can be represented as the matrix operation shown in Equation 7.
[0168] [Formula 7]
[0169]
[0170] In Equation 7, c1 to c 16 This can represent the transformation coefficients of the target block. As a result of the calculation in Equation 7, the transformation coefficients representing the modifications to the target block or the r of the residual samples of the target block can be derived. j And derive r j The process can be shown in Equation 8.
[0171] [Formula 8]
[0172]
[0173] As a result of Equation 8, the transformation coefficients representing the modification of the target block or the residual samples of the target block, r1 to r2, can be derived. N From the perspective of the size of the inverse transformation matrix, the size of the regular inverse transformation matrix is 64×64 (N×N), but the size of the inverse reduced transformation matrix is reduced to 64×16 (R×N). Therefore, compared with performing the regular inverse transformation, the storage utilization rate of performing the inverse RST can be reduced by the R / N ratio. In addition, when comparing the number of multiplications N×N when using the regular inverse transformation matrix, using the inverse reduced transformation matrix can reduce the number of multiplications (N×R) by the R / N ratio.
[0174] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied based on the transform sets in Table 2. Since a transform set includes two or three transforms (kernels) depending on the intra-prediction mode, it can be configured to select one of up to four transforms, including those without applying a secondary transform. In the transforms without applying a secondary transform, the application of an identity matrix can be considered. Assuming indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case where the identity matrix is applied, i.e., without applying a secondary transform), the transform index or lfnst index, which is used as a syntax element, can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the top-left 8×8 block, the 8×8 NSST in the RST configuration can be specified via the transform index, or the 8×8 lfnst can be specified when applying LFNST. 8×8 lfnst and 8×8 RST refer to transformations of 8×8 regions within a transform coefficient block when both W and H of the target block are equal to or greater than 8, and the 8×8 region can be the top-left 8×8 region within the transform coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to transformations of 4×4 regions within a transform coefficient block when both W and H of the target block are equal to or greater than 4, and the 4×4 region can be the top-left 4×4 region within the transform coefficient block.
[0175] According to embodiments of this disclosure, for the transformation during the encoding process, only 48 data points can be selected, and a maximum 16×48 transformation kernel matrix can be applied to them, instead of applying a 16×64 transformation kernel matrix to the 64 data points forming an 8×8 region. Here, "maximum" means that m has a maximum value of 16 in the m×48 transformation kernel matrix to generate m coefficients. That is, when performing RST by applying an m×48 transformation kernel matrix (m≤16) to an 8×8 region, 48 data points are input, and m coefficients are generated. When m is 16, 48 data points are input, and 16 coefficients are generated. That is, assuming 48 data points form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied sequentially, thereby generating a 16×1 vector. Here, the 48 data points forming the 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on 48 data points constituting the region other than the lower right 4×4 region within the 8×8 region. Here, when matrix operations are performed by applying a maximum 16×48 transformation kernel matrix, 16 modified transformation coefficients are generated. These 16 modified transformation coefficients can be arranged in the upper left 4×4 region according to the scan order, and the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.
[0176] For the inverse transform in the decoding process, the transpose of the aforementioned transform kernel matrix can be used. That is, when performing inverse RST or LFNST during the inverse transform performed by the decoding device, the input coefficient data for applying inverse RST is arranged in a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector with the corresponding inverse RST matrix to the left of the one-dimensional vector is arranged in a two-dimensional block according to a predetermined arrangement order.
[0177] In summary, during the transformation process, when RST or LFNST is applied to an 8×8 region, matrix operations are performed on the 48 transformation coefficients in the upper left, upper right, and lower left regions of the 8×8 region (excluding the lower right region) with a 16×48 transformation kernel matrix. For matrix operations, the 48 transformation coefficients are input as a one-dimensional array. When performing matrix operations, 16 modified transformation coefficients are derived, and these modified coefficients can be arranged in the upper left region of the 8×8 region.
[0178] Conversely, in the inverse transform process, when the inverse RST or LFNST is applied to an 8×8 region, the 16 transform coefficients corresponding to the upper left region of the 8×8 region can be input as a one-dimensional array according to the scan order, and matrix operations can be performed with a 48×16 transform kernel matrix. That is, the matrix operation can be represented as a (48×16 matrix). (16×1 transformation coefficient vector) = (48×1 modified transformation coefficient vector). Here, an n×1 vector can be interpreted as having the same meaning as an n×1 matrix, and therefore can be represented as an n×1 column vector. Furthermore, This represents matrix multiplication. When performing matrix operations, 48 modified transformation coefficients can be derived and arranged in the upper left, upper right, and lower left regions of an 8×8 region, excluding the lower right region.
[0179] When the inverse quadratic transform is based on the Regression-Simplified Transform (RST), the inverse transformer 135 of the encoding device 100 and the inverse transformer 222 of the decoding device 200 may include an inverse reduced quadratic transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients, and an inverse first-order transformer for deriving residual samples of the target block based on the inverse first-order transform of the modified transform coefficients. The inverse first-order transform refers to the inverse transform applied to the first-order transform of the residuals. In this disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.
[0180] The Non-Separate Transform (LFNST) described above will be described in detail below. LFNST may include a forward transform performed by the encoding device and an inverse transform performed by the decoding device.
[0181] The encoding device receives the result (or part of the result) derived after applying a first (core) transform as input and applies a forward second transform (second transform).
[0182] [Formula 9]
[0183]
[0184] In Equation 9, x and y are the input and output of the quadratic transformation, respectively, and G is the matrix representing the quadratic transformation, with the transformation basis vectors consisting of column vectors. In the case of inverse LFNST, when the dimension of the transformation matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the transpose of matrix G becomes G... T Dimensions.
[0185] For the inverse LFNST, the dimensions of matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of the eight transformed basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.
[0186] On the other hand, for a positive LFNST, matrix G TThe dimensions are [16×48], [8×48], [16×16], and [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transformation basis vectors from the upper part of the [16×48] matrix and the [16×16] matrix, respectively.
[0187] Therefore, in the case of forward LFNST, a [48×1] vector or a [16×1] vector can be used as input x, and a [16×1] vector or an [8×1] vector can be used as output y. In video encoding and decoding, the output of the forward first transform is two-dimensional (2D) data, so in order to construct a [48×1] vector or a [16×1] vector as input x, it is necessary to construct a one-dimensional vector by properly arranging the 2D data as the output of the forward transform.
[0188] Figure 6 This is a diagram illustrating the order in which the output data of a forward first transformation is arranged into a one-dimensional vector, based on the example. Figure 6 The left figures of (a) and (b) show the order used to construct the [48×1] vector, and Figure 6 The right figures (a) and (b) illustrate the order used to construct the [16×1] vector. In the case of LFNST, this can be achieved by combining 2D data with... Figure 6 Arrange the same order in (a) and (b) sequentially to obtain a one-dimensional vector x.
[0189] The orientation of the output data for the forward first transform can be determined based on the intra-prediction mode of the current block. For example, when the intra-prediction mode of the current block is horizontal relative to the diagonal direction, the orientation can be determined by... Figure 6 The output data of the forward first transform are arranged in the order of (a), and when the intra-prediction mode of the current block is perpendicular to the diagonal direction, it can be arranged according to... Figure 6 The output data of the first forward transformation are arranged in the order of (b).
[0190] Based on the example, different methods can be applied. Figure 6 The arrangement order of (a) and (b), and for derivation and application Figure 6 The arrangement order of (a) and (b) results in the same outcome (y vector), and the column vectors of matrix G can be rearranged according to the arrangement order. That is, the column vectors of G can be rearranged such that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0191] Since the output y derived by Equation 9 is a one-dimensional vector, when two-dimensional data is required as input data in the process of using the result of the forward quadratic transform as input (e.g., in the process of performing quantization or residual coding), the output y vector of Equation 9 needs to be properly arranged as 2D data again.
[0192] Figure 7 This is a diagram illustrating the order in which the output data of the forward quadratic transform is arranged into two-dimensional blocks, based on the example.
[0193] In the case of LFNST, the output values can be arranged in 2D blocks according to a predetermined scan order. Figure 7 (a) shows how the output values are arranged at 16 positions in a 2D block according to the diagonal scan order when the output y is a [16×1] vector. Figure 7 (b) shows that when the output y is an [8×1] vector, the output values are arranged in 8 positions of the 2D block according to the diagonal scan order, and the remaining 8 positions are filled with zeros. Figure 7 In (b), X indicates that it is filled with zeros.
[0194] According to another example, since the order in which the output vector y is processed during quantization or residual coding can be preset, the output vector y does not need to be arranged as shown in the example. Figure 7 In the 2D block shown. However, in the case of residual coding, data encoding can be performed in 2D block (e.g., 4×4) cells (e.g., CG (coefficient group)), and in this case, according to as Figure 7 The data is arranged in a specific order within the diagonal scanning sequence.
[0195] Furthermore, the decoding device can configure the one-dimensional input vector y by arranging the two-dimensional data output from the dequantization process according to a preset scan order used for the inverse transform. The input vector y can be output as the output vector x using the following formula.
[0196] [Formula 10]
[0197]
[0198] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.
[0199] The output vector x is based on Figure 6 The sequence shown is arranged in a two-dimensional block and is arranged as two-dimensional data, which becomes the input data (or part of the input data) for the inverse first transformation.
[0200] Therefore, the inverse quadratic transform is the opposite of the forward quadratic transform process in general, and in the case of the inverse transform, unlike in the forward direction, the inverse quadratic transform is applied first, followed by the inverse first transform.
[0201] In the inverse LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be chosen as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.
[0202] Additionally, eight matrices can be derived from the four transform sets shown in Table 2 above, and each transform set can consist of two matrices. The choice of which of the four transform sets to use is determined based on the intra-prediction mode, and more specifically, based on the values of the intra-prediction mode extended by taking into account wide-angle intra-prediction (WAIP). The selection of which matrix from the two matrices constituting the chosen transform set is derived via index signaling. More specifically, 0, 1, and 2 can be used as transmit index values; 0 can indicate that LFNST is not applied, and 1 and 2 can indicate either of the two transform matrices constituting the transform set selected based on the intra-prediction mode values.
[0203] Furthermore, as mentioned above, the size and shape of the target block determine which transformation matrix, either the [48×16] matrix or the [16×16] matrix, will be applied to the LFNST.
[0204] Figure 8 This is a diagram illustrating the block shape to which LFNST is applied. Figure 8 (a) shows a 4×4 block. Figure 8 (b) shows 4×8 blocks and 8×4 blocks. Figure 8 (c) shows a 4×N block or an N×4 block, where N is 16 or greater. Figure 8 (d) shows an 8×8 block. Figure 8 (e) shows an M×N block where M≥8, N≥8 and N>8 or M>8.
[0205] exist Figure 8 In the diagram, blocks with thick boundaries indicate the area where LFNST is applied. For Figure 8 For blocks (a) and (b), LFNST is applied to the top-left 4×4 region, and for Figure 8 Block (c) is individually applied to two consecutively arranged top-left 4×4 regions. Figure 8 In (a), (b), and (c), since the LFNST is applied in units of 4×4 regions, this LFNST will be referred to as "4×4 LFNST" in the following text. Based on the matrix dimension of G, a [16×16] or [16×8] matrix can be applied.
[0206] More specifically, a [16×8] matrix is applied to Figure 8 (a) 4×4 blocks (4×4 TU or 4×4 CU), and a [16×16] matrix is applied to Figure 8 The blocks in (b) and (c) are used to adjust the worst-case computational complexity to 8 multiplications per sample.
[0207] about Figure 8 In (d) and (e), LFNST is applied to the top-left 8×8 region, and this LFNST is referred to as "8×8 LFNST" below. As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of the forward LFNST, since the [48×1] vector (the X vector in Equation 9) is input as input data, not all sample values from the top-left 8×8 region are used as input values for the forward LFNST. That is, as can be obtained from... Figure 6 The left-hand order of (a) or Figure 6 As can be seen from the left-hand order of (b), the [48×1] vector can be constructed based on the samples belonging to the other three 4×4 blocks while leaving the bottom right 4×4 block as is.
[0208] A [48×8] matrix can be applied to Figure 8 The 8×8 blocks (8×8 TU or 8×8 CU) in (d) and the [48×16] matrix can be applied Figure 8 The 8×8 blocks in (e). This is also to adjust the worst-case computational complexity to 8 multiplications per sample.
[0209] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data (the Y vector in Equation 9, [8×1] or [16×1] vectors) are generated. In the forward LFNST, due to matrix G... T Due to its characteristic, the amount of output data is equal to or less than the amount of input data.
[0210] Figure 9 This is a diagram illustrating the arrangement of the output data of the forward LFNST according to an example, and showing the blocks in which the output data of the forward LFNST is arranged according to the block shape.
[0211] exist Figure 9 The shaded area in the upper left corner of the block shown corresponds to the region where the output data of the forward LFNST is located. The positions marked with 0 indicate samples filled with a value of 0, and the remaining areas represent regions that were not altered by the forward LFNST. In regions not altered by LFNST, the output data of the first forward transform remains unchanged.
[0212] As mentioned above, since the size of the applied transformation matrix varies depending on the shape of the block, the amount of output data also varies. Figure 9The output data of a forward LFNST may not completely fill the top-left 4×4 block. Figure 9 In cases (a) and (d), the [16×8] matrix and the A[48×8] matrix are applied to the block indicated by the thick line or a portion of the area inside the block, respectively, and an [8×1] vector is generated as the output of the positive LFNST. That is, according to Figure 7 The scan order shown in (b) can fill only 8 output data, such as Figure 9 As shown in (a) and (d), zeros can be filled in the remaining 8 positions. Figure 8 In the case of the LFNST application block of (d), such as Figure 9 As shown in (d), the two 4×4 blocks adjacent to the top-left 4×4 block, the top-right and bottom-left blocks, are also filled with the value 0.
[0213] As described above, essentially, by signaling the LFNST index, it is specified whether to apply LFNST and the transformation matrix to be applied. Figure 9 As shown, when LFNST is applied, since the number of output data of the positive LFNST can be equal to or less than the number of input data, the following area filled with zero values appears.
[0214] 1) such as Figure 9 As shown in (a), the samples are from the eighth position and the subsequent positions in the scanning order of the top left 4×4 block, that is, from the ninth to the sixteenth position.
[0215] 2) such as Figure 9 As shown in (d) and (e), when applying a [48×16] matrix or a [48×8] matrix, the two 4×4 blocks adjacent to the top left 4×4 block or the second and third 4×4 blocks in the scan order.
[0216] Therefore, if non-zero data is found in regions 1) and 2), it is determined that LFNST has not been applied, so the signaling for the corresponding LFNST index can be omitted.
[0217] Based on the example, such as in the case of LFNST used in the VVC standard, since the signaling for the LFNST index is executed after residual coding, the encoding device can know from the residual coding whether non-zero data (valid coefficients) exists at all locations within the TU or CU block. Therefore, the encoding device can determine whether to execute signaling regarding the LFNST index based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. The signaling for the LFNST index is executed when non-zero data does not exist in the areas specified in 1) and 2) above.
[0218] Because truncated unary codes are used as the binarization method for the LFNST index, the LFNST index consists of up to two bins, and 0, 10, and 11 are assigned as binary codes for possible LFNST index values 0, 1, and 2, respectively. In the current case of LFNST used for VVC, context-based CABAC encoding is applied to the first bin (regular encoding), and bypass encoding is applied to the second bin. The total number of contexts in the first bin is 2. When (DCT-2, DCT-2) is applied for a single transform pair in the horizontal and vertical directions, and the luma and chroma components are encoded in a dual-tree type, one context is assigned, and the other context is applied for the rest. The encoding of the LFNST index is shown in the table below.
[0219] [Table 3]
[0220]
[0221] In addition, the following simplification method can be applied to the LFNST used.
[0222] (i) As shown in the example, the number of output data for a positive LFNST can be limited to a maximum of 16.
[0223] exist Figure 8 In case (c), the 4×4 LFNST can be applied to two adjacent 4×4 regions to the upper left, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data for the forward LFNST is limited to a maximum of 16, in the case of 4×N / N×4 (N≥16) blocks (TU or CU), the 4×4 LFNST is applied only to one 4×4 region to the upper left, and the LFNST can be applied only to... Figure 8 All blocks are processed at once. This simplifies the implementation of image encoding.
[0224] Figure 10 The example shows that the amount of output data for a positive LFNST is limited to a maximum of 16. (See example.) Figure 10 When LFNST is applied to the top left 4×4 region of a 4×N or N×4 block (where N is 16 or greater), the output data of the forward LFNST becomes 16.
[0225] (ii) As in the example, zeroing can be additionally applied to regions where LFNST has not been applied. In this document, zeroing can mean filling all positions belonging to a particular region with a value of 0. That is, zeroing can be applied to regions that have not changed due to LFNST and maintain the result of a positive first transformation. As mentioned above, since LFNST is divided into 4×4 LFNST and 8×8 LFNST, zeroing can be divided into two types as follows ((ii)-(A) and (ii)-(B)).
[0226] (ii)-(A) When 4×4 LFNST is applied, the area where 4×4 LFNST is not applied can be zeroed. Figure 11 This is a diagram illustrating the zeroing process in a block of 4×4 LFNST, based on the example.
[0227] like Figure 11 As shown, regarding the block that applied 4×4 LFNST, that is, for Figure 9 All blocks in (a), (b) and (c) where LFNST is not applied can be filled with zeros.
[0228] on the other hand, Figure 11 (d) shows that when the maximum number of output data for the positive LFNST is limited to 16 (e.g. Figure 10 When (as shown), zeroing is performed on the remaining blocks that have not applied 4×4 LFNST.
[0229] (ii)-(B) When 8×8 LFNST is applied, areas where 8×8 LFNST is not applied can be zeroed. Figure 12 This is a diagram illustrating the zeroing process in a block of 8×8 LFNST, based on the example.
[0230] like Figure 12 As shown, regarding the application of 8×8 LFNST to the block, that is, for Figure 9 In all blocks in (d) and (e), the entire area where LFNST is not applied can be filled with zeros.
[0231] (iii) Due to the zeroing presented in (ii) above, the zero-filled area may not be the same as when LFNST was applied. Therefore, it can be determined by comparison. Figure 9 In the case of LFNST, a wider area is used to perform zeroing as proposed in (ii) to check for the presence of non-zero data.
[0232] For example, when (ii)-(B) is applied, in the examination Figure 9 After checking whether there is non-zero data in the zero-filled regions in (d) and (e), additional checks are performed. Figure 12The presence of non-zero data in the zero-filled region can be used to perform signaling for the LFNST index only if no non-zero data exists.
[0233] Of course, even with the zeroing proposed in application (ii), the existence of non-zero data can be checked in the same way as existing LFNST index signaling. That is, when checking... Figure 9 After confirming the presence of non-zero data within the zero-padded block, LFNST index signaling can be applied. In this case, the encoding device only performs zeroing and the decoding device does not assume zeroing; that is, it only checks whether non-zero data exists within the zero-padded block. Figure 9 In regions explicitly marked as 0, LFNST index resolution can be performed.
[0234] Alternatively, according to another example, the following can be performed: Figure 13 The reset is shown. Figure 13 This is a diagram illustrating the zeroing of a block in an 8×8 LFNST application, based on another example.
[0235] like Figure 11 and Figure 12 As shown, zeroing can be applied to all areas except the area where LFNST is applied, or it can be applied only to a local area, such as... Figure 13 As shown. Zeroing only applies to items other than... Figure 13 Zeroing the area outside the top-left 8x8 area does not apply to the bottom-right 4x4 block within the top-left x8 area.
[0236] Various implementations of the simplified methods for applying LFNST (combinations of (i), (ii)-(A), (ii)-(B), (iii)) can be derived. Of course, the combinations of the above simplified methods are not limited to the following implementations, and any combination can be applied to LFNST.
[0237] Implementation
[0238] - Limit the number of output data for the forward LFNST to a maximum of 16. (i)
[0239] - When 4×4 LFNST is applied, all areas where 4×4 LFNST is not applied are zeroed out. (II)-(A)
[0240] - When 8×8 LFNST is applied, all areas where 8×8 LFNST is not applied are zeroed out. (II)-(B)
[0241] - After checking whether non-zero data also exists in existing areas filled with zero values and areas filled with zeros due to additional zeroing ((ii)-(A), (ii)-(B)), the LFNST index is signaled only if no non-zero data exists. (iii).
[0242] In the implementation scenario, when LFNST is applied, the area containing non-zeroed data is limited to the upper left 4×4 area. More specifically, in Figure 11 (a) and Figure 12 In case (a), the eighth position in the scan order is the last position where non-zero data can exist. Figure 11 (b) and (c) and Figure 12 In case (b), the sixteenth position in the scan order (i.e., the position at the bottom right edge of the top left 4×4 block) is the last position in which data other than 0 can exist.
[0243] Therefore, after applying LFNST, and after checking whether non-zero data exists at a position that is not allowed in the residual encoding process (at a position beyond the last position), it can be determined whether to signal the LFNST index.
[0244] In the case of the zeroing method proposed in (ii), the computational cost required to perform the entire transformation process can be reduced because of the amount of data ultimately generated when both the first transformation and LFNST are applied. That is, when LFNST is applied, since zeroing is applied to regions where the output data of the forward first transformation exists without LFNST, it is not necessary to generate data for regions that are zeroed during the forward first transformation. Therefore, the computational cost required to generate the corresponding data can be reduced. The additional effects of the zeroing method proposed in (ii) are summarized below.
[0245] First, as mentioned above, reduce the amount of computation required to perform the entire transformation process.
[0246] Specifically, when (ii)-(B) is applied, the worst-case computational cost is reduced, making the transformation process lighter. In other words, generally, a large amount of computation is required to perform a single transformation of a large size. By applying (ii)-(B), the amount of data derived as a result of performing a forward LFNST can be reduced to 16 or less. Furthermore, as the size of the entire block (TU or CU) increases, the effect of reducing the number of transformation operations further increases.
[0247] Secondly, it can reduce the amount of computation required for the entire transformation process, thereby reducing the power consumption required to perform the transformation.
[0248] Third, it reduces the delay involved in the transformation process.
[0249] Secondary transforms, such as LFNST, add computational complexity to existing primary transforms, thus increasing the overall latency involved in performing the transform. Specifically, in the case of intra-frame prediction, the increased latency due to secondary transforms during encoding leads to an increase in latency until reconstruction because reconstructed data from adjacent blocks is used during prediction. This can result in an increase in the overall latency of intra-frame predictive coding.
[0250] However, if the zeroing proposed in application (ii) is applied, the delay time for performing a single transformation can be greatly reduced when LFNST is applied, maintaining or reducing the delay time of the entire transformation, making it easier to implement the encoding device.
[0251] In traditional intra-frame prediction, the block to be encoded is treated as a single coding unit and encoding is performed without segmentation. However, Intra-Frame Sub-Partition (ISP) coding means performing intra-frame prediction coding by dividing the block to be encoded horizontally or vertically. In this case, reconstructed blocks can be generated by performing encoding / decoding on a block-by-block basis, and the reconstructed blocks can be used as reference blocks for the next block. According to implementations, in ISP coding, a coding block can be divided into two or four sub-blocks and encoded, and in ISP, within a sub-block, intra-frame prediction is performed with reference to the reconstructed pixel values of the adjacent left or upper sub-block. Hereinafter, "encoding" can be used as a concept encompassing both encoding performed by an encoding device and decoding performed by a decoding device.
[0252] Table 4 shows the number of sub-blocks divided according to the block size when applying ISP, and the sub-partitions divided according to ISP can be called transform blocks (TU).
[0253] [Table 4]
[0254]
[0255] ISP divides blocks within a predicted luma frame into two or four sub-partitions in the vertical or horizontal direction, based on the block size. For example, the minimum block size that can be applied to ISP is 4×8 or 8×4. When the block size is larger than 4×8 or 8×4, the block is divided into 4 sub-partitions.
[0256] Figure 14 and Figure 15 An example of how a coded block is divided into sub-blocks is given, and more specifically, Figure 14 Examples of coded blocks (width (W) × height (H)) being divided into 4×8 blocks or 8×4 blocks are shown, and Figure 15 Examples of partitioning are shown for cases where the coded block is not a 4×8 block, 8×4 block, or 4×4 block.
[0257] When applying ISP, sub-blocks are encoded sequentially from left to right or top to bottom (e.g., horizontally or vertically) according to the partitioning type. After reconstruction processing via inverse transform and intra-prediction for one sub-block, the encoding of the next sub-block can be performed. For the leftmost or topmost sub-block, reconstructed pixels of already encoded blocks are referenced, as in conventional intra-prediction methods. Furthermore, when each side of a subsequent internal sub-block is not adjacent to the previous sub-block, reconstructed pixels of already encoded adjacent blocks are referenced to derive reference pixels adjacent to the corresponding side, as in conventional intra-prediction methods.
[0258] In ISP coding mode, all sub-blocks can be encoded using the same intra-prediction mode, and signals can be sent indicating whether ISP coding is used and whether the sub-blocks are divided in the direction (horizontal or vertical). For example... Figure 14 and Figure 15 As shown, the number of sub-blocks can be adjusted to 2 or 4 depending on the shape of the block. When the size (width × height) of a sub-block is less than 16, it can be restricted so that it is not allowed to be divided into corresponding sub-blocks or the ISP encoding itself is not applied.
[0259] In ISP prediction mode, a coding unit is divided into two or four partition blocks (i.e., sub-blocks) and prediction is performed, and the same intra-prediction mode is applied to the two or four partition blocks.
[0260] As described above, in terms of partitioning direction, both the horizontal direction (when M×N coding units with horizontal and vertical lengths of M and N are partitioned horizontally, if an M×N coding unit is divided into two, then the M×N coding unit is divided into M×(N / 2) blocks, and if an M×N coding unit is divided into four blocks, then the M×N coding unit is divided into M×(N / 4) blocks) and the vertical direction (when M×N coding units are partitioned vertically, if an M×N coding unit is divided into two, then the M×N coding unit is divided into (M / 2)×N blocks, and if an M×N coding unit is divided into four, then the M×N coding unit is divided into (M / 4)×N blocks) are possible. When M×N coding units are partitioned horizontally, the partition blocks are encoded in a top-to-bottom order, and when M×N coding units are partitioned vertically, the partition blocks are encoded in a left-to-right order. In the case of horizontal (vertical) division, the reconstructed pixel values of the upper (left) partition can be referenced to predict the current encoded partition.
[0261] Transforms can be applied to residual signals generated in blocks using the ISP prediction method. Multiple transform selection (MTS) techniques based on the DST-7 / DCT-8 combination and the existing DCT-2 can be applied to a forward-based single transform (core transform), and forward low-frequency non-separable transform (LFNST) can be applied to the transform coefficients generated from the single transform to generate the final modified transform coefficients.
[0262] In other words, LFNST can be applied to partitions divided by applying the ISP prediction mode, and the same intra-prediction mode is applied to the partitioned partitions, as described above. Therefore, when selecting an LFNST set derived based on the intra-prediction mode, the derived LFNST set can be applied to all partitions. That is, because the same intra-prediction mode is applied to all partitions, the same LFNST set can be applied to all partitions.
[0263] According to the implementation, LFNST can be applied only to transform blocks with both horizontal and vertical lengths of 4 or greater. Therefore, when the horizontal or vertical length of a partition block divided according to the ISP prediction method is less than 4, LFNST is not applied and no LFNST index is signaled. Furthermore, when applying LFNST to each partition block, the corresponding partition block can be considered as a transform block. When the ISP prediction method is not applied, LFNST can be applied to the coded block.
[0264] The method of applying LFNST to each partition block will be described in detail.
[0265] According to the implementation method, after applying the forward LFNST to each partition block, only a maximum of 16 (8 or 16) coefficients are left in the upper left 4×4 region in the order of scanning the transform coefficients, and then zeroing can be applied, in which the remaining positions and regions are all filled with 0.
[0266] Alternatively, according to the implementation, when the length of one side of the partition block is 4, LFNST is applied only to the upper left 4×4 region, and when the length (i.e., width and height) of all sides of the partition block is 8 or greater, LFNST can be applied to the remaining 48 coefficients in the upper left 8×8 region except for the lower right 4×4 region.
[0267] Alternatively, according to the implementation method, in order to adjust the worst-case computational complexity to 8 multiplications per sample, when each partition is 4×4 or 8×8, only 8 transformation coefficients can be output after applying the forward LFNST. That is, when the partition is 4×4, an 8×16 matrix can be used as the transformation matrix, and when the partition is 8×8, an 8×48 matrix can be used as the transformation matrix.
[0268] In the current VVC standard, LFNST index signaling is executed on a unit-by-unit basis. Therefore, in ISP prediction mode, and when LFNST is applied to all partition blocks, the same LFNST index value can be applied to the corresponding partition block. That is, when an LFNST index value is sent once at the unit-level, the corresponding LFNST index can be applied to all partition blocks within that unit. As mentioned above, the LFNST index value can have values of 0, 1, and 2, where 0 indicates no LFNST application, and 1 and 2 represent two transform matrices existing in a set of LFNSTs when LFNST is applied.
[0269] As mentioned above, the LFNST set is determined by the intra-prediction mode, and in the case of ISP prediction mode, since all partition blocks in the coding unit are predicted in the same intra-prediction mode, the partition blocks can refer to the same LFNST set.
[0270] As another example, LFNST index signaling is still performed on a unit-by-unit basis. However, in ISP predictive mode, it is uncertain whether LFNST is applied uniformly to all blocks. For each block, the application of the LFNST index value signaled at the unit-by-unit level and the application of LFNST can be determined by separate conditions. Here, separate conditions can be signaled via a bitstream in the form of flags for each block. When the flag value is 1, the LFNST index value signaled at the unit-by-unit level is applied, and when the flag value is 0, LFNST may not be applied.
[0271] In an encoding unit that uses ISP mode, an example of applying LFNST when the length of one side of the partition block is less than 4 is described below.
[0272] First, when the size of the partition block is N×2 (2×N), LFNST can be applied to the upper left M×2 (2×M) region (where M≤N). For example, when M=8, the upper left region becomes 8×2 (2×8), so the region with 16 residual signals can be the input of the forward LFNST, and an R×16 (R≤16) forward transformation matrix can be applied.
[0273] Here, the forward LFNST matrix can be a separate, additional matrix besides those included in the current VVC standard. Furthermore, for worst-case complexity control, an 8×16 matrix, where only the top 8 rows of the 16×16 matrix are sampled, can be used for the transformation. The complexity control method will be described in detail later.
[0274] Secondly, when the size of the partition block is N×1 (1×N), LFNST can be applied to the upper left M×1 (1×M) region (where M≤N). For example, when M=16, the upper left region becomes 16×1 (1×16), so the region with 16 residual signals can be the input of the forward LFNST, and an R×16 (R≤16) forward transformation matrix can be applied.
[0275] Here, the corresponding forward LFNST matrix can be a separate additional matrix besides those included in the current VVC standard. Furthermore, to control worst-case complexity, an 8×16 matrix, where only the top 8 rows of the 16×16 matrix are sampled, can be used for the transformation. The complexity control method will be described in detail later.
[0276] The first and second embodiments can be applied simultaneously, or either one of the two embodiments can be applied. Specifically, in the case of the second embodiment, because a transformation is considered in the LFNST, experiments have shown that the compression performance improvement obtainable in the existing LFNST is relatively small compared to the LFNST index signaling cost. However, in the case of the first embodiment, a compression performance improvement similar to that obtainable from the conventional LFNST is observed. That is, in the case of ISP, the contribution of applying 2×N and N×2 LFNSTs to the actual compression performance can be observed experimentally.
[0277] In the current VVC's LFNST, symmetry is applied between intra-prediction modes. The same set of LFNSTs is applied to two directional modes set around mode 34 (prediction in the 45-degree diagonal direction at the bottom right corner), for example, the same set of LFNSTs is applied to mode 18 (horizontal directional prediction mode) and mode 50 (vertical directional prediction mode). However, in modes 35 through 66, when a forward LFNST is applied, the input data is transposed before the LFNST is applied.
[0278] VVC supports Wide-Angle Intra-Prediction (WAIP) mode. Considering WAIP mode, the LFNST set is derived based on the modified intra-prediction mode. For modes extended by WAIP, the LFNST set is determined using symmetry, just as in general intra-prediction directional modes. For example, because mode-1 is symmetric to mode 67, the same LFNST set is applied, and because mode-14 is symmetric to mode 80, the same LFNST set is applied. Modes 67 through 80 apply the LFNST transform after transposing the input data before applying the forward LFNST.
[0279] When applying LFNST to the top-left M×2 (M×1) block, symmetry with respect to LFNST cannot be applied because the block to which LFNST is applied is not square. Therefore, instead of applying symmetry based on intra-prediction mode, as shown in Table 2 for LFNST, symmetry between M×2 (M×1) and 2×M (1×M) blocks can be applied.
[0280] Figure 16 This is a diagram illustrating the symmetry between M×2 (M×1) blocks and 2×M (1×M) blocks according to the implementation method.
[0281] like Figure 16 As shown, since pattern 2 in the M×2 (M×1) block can be considered symmetric to pattern 66 in the 2×M (1×M) block, the same LFNST set can be applied to both the 2×M (1×M) block and the M×2 (M×1) block.
[0282] In this case, in order to apply the LFNST set applied to the M×2 (M×1) block to the 2×M (1×M) block, the LFNST set is selected based on mode 2 instead of mode 66. That is, the LFNST can be applied after transposing the input data of the 2×M (1×M) block before applying the forward LFNST.
[0283] Figure 17 This is a diagram illustrating an example of a transposed 2×M block according to an embodiment.
[0284] Figure 17 (a) is a diagram illustrating how LFNST can be applied by reading 2×M blocks of input data in column-major order. Figure 17 (b) is a graph illustrating how LFNST can be applied by reading the input data of an M×2 (M×1) block in row-major order. The method for applying LFNST to the top-left M×2 (M×1) or 2×M (M×1) block is described below.
[0285] 1. First, as Figure 17 As shown in (a) and (b), the input data is arranged into an input vector that constitutes a positive LFNST. For example, refer to Figure 16 For an M×2 block predicted using mode 2, follow Figure 17 In the order of (b), for a 2×M block predicted in pattern 66, the input data is in the following order: Figure 17 The sequential arrangement of (a) can then be applied to LFNST set for mode 2.
[0286] 2. For an M×2 (M×1) block, considering WAIP, the LFNST set is determined based on the modified intra-prediction mode. As mentioned above, a preset mapping relationship is established between the intra-prediction mode and the LFNST set, which can be represented by the mapping table shown in Table 2.
[0287] For a 2×M (1×M) block, taking into account WAIP, a symmetric mode around the prediction mode (mode 34 in the case of the VVC standard) can be obtained from the modified intra-prediction mode, moving downwards along a 45-degree diagonal. The LFNST set is then determined based on the corresponding symmetric mode and the mapping table. The symmetric mode (y) around mode 34 can be derived using the following formula. The mapping table will be described in more detail below.
[0288] [Equation 11]
[0289] If 2 ≤ x ≤ 66, then y = 68 - x.
[0290] Otherwise (x≤-1 or x≥67), y=66-x
[0291] 3. When applying forward LFNST, the transform coefficients can be derived by multiplying the input data prepared in process 1 by the LFNST kernel. The LFNST kernel can be selected based on the LFNST set determined in process 2 and the predetermined LFNST index.
[0292] For example, when M=8 and a 16×16 matrix is used as the LFNST kernel, 16 transform coefficients can be generated by multiplying the matrix by 16 input data. The generated transform coefficients can be arranged in the upper left 8×2 or 2×8 region according to the scan order used in the VVC standard.
[0293] Figure 18 The scanning sequence of 8×2 or 2×8 regions according to the implementation method is illustrated.
[0294] All regions except the top-left 8×2 or 2×8 region can be filled with zero values (cleared), or existing transformation coefficients that have undergone a single transformation can be left as is. The predefined LFNST index can be one of the LFNST index values (0, 1, 2) that are tried when calculating the RD cost while changing the LFNST index value during programming processing.
[0295] In cases where the worst-case computational complexity is tuned to a certain level or lower (e.g., 8 multiplications / sample), for example, after generating only 8 transformation coefficients by multiplying by an 8×16 matrix that takes only the top 8 rows of a 16×16 matrix, the transformation coefficients can be... Figure 18 The scan order can be set, and zeroing can be applied to the remaining coefficient regions. Worst-case complexity control will be described later.
[0296] 4. When applying the inverse LFNST, a preset number (e.g., 16) of transform coefficients are set as the input vector, and the LFNST set obtained from process 2 and the LFNST kernel (e.g., a 16×16 matrix) derived from the selected parse LFNST index are selected. The output vector can then be derived by multiplying the LFNST kernel with the corresponding input vector.
[0297] In the case of M×2 (M×1) blocks, the output vector can be... Figure 17 The row priority setting in (b) is used, while in the case of 2×M (1×M) blocks, the output vector can be set to... Figure 17 The column priority setting for (a).
[0298] Except for the regions where the corresponding output vectors are set in the upper left M×2 (M×1) or 2×M (M×2) region, the remaining regions in the partition block except for the upper left M×2 (M×1) or 2×M (M×2) region (the M×2 region in the partition block) can all be cleared to have zero values, or can be configured to retain the reconstructed transform coefficients as is through residual coding and inverse quantization.
[0299] When constructing the input vector, as in point 3, the input data can be constructed according to... Figure 18 The scanning order can be arranged, and in order to keep the worst-case computational complexity to a certain extent or lower, the input vector can be constructed by reducing the number of input data (e.g., 8 instead of 16).
[0300] For example, when M=8, if 8 input data are used, the leftmost 16×8 matrix can be taken from the corresponding 16×16 matrix and multiplied to obtain 16 output data. Worst-case complexity control will be described later.
[0301] In the above implementation, when applying LFNST, the case of applying symmetry between M×2 (M×1) blocks and 2×M (1×M) blocks is shown. However, according to another example, different sets of LFNST can be applied to each of the two block shapes.
[0302] The following sections will describe various examples of mapping methods using intra-prediction mode and LFNST set configurations using ISP mode.
[0303] In ISP mode, the LFNST set configuration can differ from the existing LFNST set. In other words, a different core than the existing LFNST core can be applied, and a different mapping table can be applied than the mapping table used between the intra-prediction mode index and the LFNST set in the current VVC standard. The mapping table used in the current VVC standard can be the same as the mapping table in Table 2.
[0304] In Table 2, the preModeIntra value represents the intra-prediction mode value that has changed to take WAIP into account, and the lfnstTrSetIdx value is the index value indicating a specific LFNST set. Each LFNST set is configured with two LFNST cores.
[0305] When applying the ISP prediction mode, if both the horizontal and vertical lengths of each partition block are equal to or greater than 4, the same kernel as the LFNST kernel used in the current VVC standard can be applied, and the mapping table can be applied as is. Alternatively, mapping tables and LFNST kernels different from those in the current VVC standard can be applied.
[0306] When applying the ISP prediction mode, if the horizontal or vertical length of each block is less than 4, a mapping table and LFNST core different from the current VVC standard can be applied. In the following text, Tables 5 to 7 show the mapping table between intra-prediction mode values (intra-prediction mode values changed to take into account WAIP) and LFNST sets, which can be applied to M×2 (M×1) blocks or 2×M (1×M) blocks.
[0307] [Table 5]
[0308]
[0309] [Table 6]
[0310]
[0311] [Table 7]
[0312]
[0313] The first mapping table in Table 5 is configured with seven LFNST sets, the mapping table in Table 6 is configured with four LFNST sets, and the mapping table in Table 7 is configured with two LFNST sets. As another example, when it is configured with one LFNST set, the lfnstTrSetIdx value can be fixed to 0 relative to the preModeIntra value.
[0314] The following section describes a method for maintaining the worst-case computational complexity when applying LFNST to the ISP pattern.
[0315] In ISP mode, when applying LFNST, the number of multiplications per sample (or per coefficient, per position) may be limited to a certain value or less. Depending on the size of the partition block, the number of multiplications per sample (or per coefficient, per position) can be kept to 8 or less by applying LFNST as follows.
[0316] 1. When both the horizontal and vertical lengths of the partition block are 4 or greater, the same computational complexity control method as the worst-case method for LFNST in the current VVC standard can be applied.
[0317] In other words, when the partition block is 4×4, an 8×16 matrix obtained by sampling the top 8 rows of a 16×16 matrix can be applied instead of a 16×16 matrix in the forward direction, and a 16×8 matrix obtained by sampling the left 8 columns of a 16×16 matrix can be applied in the reverse direction. Furthermore, when the partition block is 8×8, in the forward direction, instead of a 16×48 matrix, an 8×48 matrix obtained by sampling the top 8 rows of a 16×48 matrix is applied, and in the reverse direction, instead of a 48×16 matrix, a 48×8 matrix obtained by sampling the left 8 columns of a 48×16 matrix can be applied.
[0318] In the case of 4×N or N×4 (N>4) blocks, when performing the forward transformation, the 16 coefficients generated after applying the 16×16 matrix only to the top-left 4×4 block can be set in the top-left 4×4 region, and other regions can be filled with values of 0. Conversely, when performing the inverse transformation, the 16 coefficients in the top-left 4×4 block are set in scan order to form the input vector, and then 16 output data points are generated by multiplying by the 16×16 matrix. The generated output data can be set in the top-left 4×4 region, and the remaining regions can be filled with values of 0.
[0319] In the case of 8×N or N×8 (N>8) blocks, when performing the forward transformation, the 16 coefficients generated after applying a 16×48 matrix to only the ROI region within the top-left 8×8 block (excluding the remaining regions from the bottom-right 4×4 block within the top-left 8×8 block) can be set in the top-left 4×4 region, and all other regions can be filled with values of 0. Furthermore, when performing the inverse transformation, the 16 coefficients located in the top-left 4×4 region are set in scan order to form the input vector, which can then be multiplied by a 48×16 matrix to generate 48 output data points. The generated output data can be filled in the ROI regions, and all other regions can be filled with values of 0.
[0320] 2. When the size of the partition block is N×2 or 2×N and LFNST is applied to the top left M×2 or 2×M region (M≤N), a matrix sampled according to the value of N can be applied.
[0321] With M=8, for N=8 partition blocks, i.e., 8×2 or 2×8 blocks, in the case of forward transformation, an 8×16 matrix obtained by sampling the top 8 rows of a 16×16 matrix can be applied instead of a 16×16 matrix, and in the case of inverse transformation, a 16×8 matrix obtained by sampling the left 8 columns of a 16×16 matrix can be applied instead of a 16×16 matrix.
[0322] When N is greater than 8, in the forward transformation, the 16×16 matrix applied to the top-left 8×2 or 2×8 block generates 16 output data points, which are then placed within that block, with the remaining areas filled with values of 0. In the inverse transformation, the 16 coefficients in the top-left 8×2 or 2×8 block are arranged in scan order to form the input vector, which is then multiplied by the 16×16 matrix to generate 16 output data points. These output data points can also be placed within the top-left 8×2 or 2×8 block, with all remaining areas filled with values of 0.
[0323] 3. When the size of the partition block is N×1 or 1×N and LFNST is applied to the top left M×1 or 1×M region (M≤N), a matrix sampled according to the value of N can be applied.
[0324] When M=16, for partitioned blocks of N=16, i.e., 16×1 or 1×16 blocks, in the case of forward transformation, an 8×16 matrix obtained by sampling the top 8 rows of the 16×16 matrix can be applied instead of a 16×16 matrix, and in the case of inverse transformation, a 16×8 matrix obtained by sampling the left 8 columns of the 16×16 matrix can be applied instead of a 16×16 matrix.
[0325] When N is greater than 16, in the forward transformation, the 16 output data generated by applying a 16×16 matrix to the top-left 16×1 or 1×16 block can be set within the top-left 16×1 or 1×16 block, and the remaining areas can be filled with values of 0. In the inverse transformation, the 16 coefficients located in the top-left 16×1 or 1×16 block can be set in scan order to form the input vector, and then multiplied by the 16×16 matrix to generate 16 output data. The generated output data can be set within the top-left 16×1 or 1×16 block, and all remaining areas can be filled with values of 0.
[0326] As another example, to keep the number of multiplications per sample (or per coefficient, per position) at a certain value or less, the number of multiplications per sample (or per coefficient, per position) can be kept to 8 or less based on the ISP coding unit size rather than the ISP block size. When only one block in the ISP blocks satisfies the conditions for applying LFNST, the worst-case complexity of LFNST can be calculated based on the corresponding coding unit size rather than the block size. For example, if the luma coding block of a particular coding unit is encoded using ISP by being divided into four 4×4 blocks, and there are no non-zero transform coefficients for two of these blocks, it can be configured such that 16 transform coefficients are generated for each of the other two blocks instead of 8.
[0327] The following describes a method for signaling the LFNST index in ISP mode.
[0328] As described above, the LFNST index can have values 0, 1, and 2, where 0 indicates that no LFNST is applied, and 1 and 2 indicate that it is included in one of the two LFNST kernel matrices in the selected LFNST set. In the current VVC standard, the LFNST index is sent as described below.
[0329] 1. The LFNST index can be sent once for each coding unit (CU). In the case of a dual-tree system, a separate LFNST index can be signaled for each of the luma and chroma blocks.
[0330] 2. When no signal is sent to the LFNST index, the LFNST index is inferred to be 0, which is the default value. The LFNST index value is inferred to be 0 in the following cases.
[0331] A. When in a mode where no transform is applied (e.g., transform skipping, BDPCM, lossless coding, etc.)
[0332] B. When a transformation is not DCT-2 (when it is DST7 or DCT8), that is, when the horizontal or vertical transformation is not DCT-2.
[0333] C. LFNST is not applicable when the horizontal or vertical length of the luminance block of the coding unit exceeds the size of the maximum convertible luminance transform, for example, when the size of the luminance block of the coding unit is 128×16 when the size of the maximum convertible luminance transform is 64.
[0334] In the dual-tree case, it is determined whether each of the coding units for the luma component and the chroma component exceeds the maximum luma transform size. That is, it is checked whether the luma block exceeds the maximum transformable luma transform size, and whether the chroma block exceeds the horizontal / vertical length of the corresponding luma block and the maximum transformable luma transform size for the color format. For example, when the color format is 4:2:0, each of the horizontal / vertical lengths of the corresponding luma block is twice the horizontal / vertical length of the chroma block, and the transform size of the corresponding luma block is twice the transform size of the chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical lengths of the corresponding luma block are the same as the horizontal / vertical lengths of the chroma block.
[0335] A 64-length transformation or a 32-length transformation can represent a transformation applied horizontally or vertically with a length of 64 or 32, respectively, and "transformation size" can represent the corresponding length of 64 or 32.
[0336] In the case of a single tree, check whether the horizontal or vertical length of the brightness block exceeds the size of the maximum transformable brightness block, and if it does, the LFNST index signaling can be skipped.
[0337] D. The LFNST index may be sent only when both the horizontal and vertical lengths of the encoding unit are greater than or equal to 4.
[0338] In the case of a two-tree system, the LFNST index can be signaled only if the horizontal and vertical lengths of the corresponding component (i.e., the luminance or chrominance component) are both greater than or equal to 4.
[0339] In the case of a single tree, the LFNST index can be signaled only when both the horizontal and vertical lengths of the luminance component are greater than or equal to 4.
[0340] E. If the last non-zero coefficient position is not a DC position (top left position of the block), when it is a dual-tree type luma block, send an LFNST index if the last non-zero coefficient position is not a DC position. When it is a dual-tree type chroma block, send an LNFST index even if one of the last non-zero coefficient positions of Cb and Cr is not a DC position.
[0341] In the case of a single-tree type, if the last non-zero coefficient position is not a DC position even in one of the luminance component, Cb component, and Cr component, the LFNST index is sent.
[0342] In this paper, if the Code Block Flag (CBF) indicating the presence of transform coefficients for a transform block is 0, the last non-zero coefficient position of the transform block is not checked to determine whether to signal the LFNST index. That is, if the CBF value is 0, the last non-zero coefficient position can be ignored when checking the conditions for LFNST index signaling since the transform is not applied to the block.
[0343] For example, 1) in the case of dual-tree type and luma component, if the corresponding CBF value is 0, no signal is sent to the LFNST index; 2) in the case of dual-tree type and chrominance component, if the CBF value of Cb is 0 and the CBF value of Cr is 1, the LFNST index is sent by checking only the last non-zero coefficient position of Cr; and 3) in the case of single-tree type, for all of luma, Cb and Cr, the last non-zero coefficient position is checked only for the component with the corresponding CBF value of 1.
[0344] F. When transform coefficients are identified at locations other than where LNFSF transform coefficients can be present, the LFNST index signaling can be skipped. In the case of 4×4 and 8×8 transform blocks, LFNST transform coefficients can exist at 8 locations starting from the DC position according to the transform coefficient scan order in the VVC standard, with all remaining locations padded with 0. Additionally, when the transform block is not a 4×4 or 8×8 transform block, LFNST transform coefficients can exist at 16 locations starting from the DC position according to the transform coefficient scan order in the VVC standard, with all remaining locations padded with 0.
[0345] Therefore, after performing residual coding, the LFNST index signaling can be skipped when non-zero transform coefficients are present in regions to be filled with 0.
[0346] Furthermore, the ISP mode can be applied only to luma blocks or to both luma and chroma blocks. As described above, when applying ISP prediction, prediction is performed by dividing the corresponding coding unit into 2 or 4 partition blocks, and the transform can be applied to each partition block. Therefore, when determining the conditions for signaling the LFNST index based on the coding unit, the fact that LFNST applies to each partition block should also be considered. Additionally, when applying the ISP prediction mode only to a specific component (e.g., luma block), the LFNST index should be signaled by considering the fact that component partitioning is only implemented for the component. Possible LFNST index signaling schemes when in ISP mode are summarized below.
[0347] 1. The LFNST index can be sent once for each coding unit (CU). In the case of a dual-tree system, a separate LFNST index can be signaled for each of the luma and chroma blocks.
[0348] 2. When no signal is sent to the LFNST index, the LFNST index is inferred to be 0, which is the default value. The LFNST index value is inferred to be 0 in the following cases.
[0349] A. When in a mode where no transform is applied (e.g., transform skipping, BDPCM, lossless coding, etc.)
[0350] B. LFNST is not applicable when the horizontal or vertical length of the luminance block of the coding unit exceeds the size of the maximum convertible luminance transform, for example, when the size of the luminance block of the coding unit is 128×16 in the case that the size of the maximum convertible luminance transform is 64.
[0351] The LFNST index signaling can also be determined based on the size of the partition block rather than the coding unit. That is, when the horizontal or vertical length of the partition block corresponding to the luma block exceeds the size of the maximum luma change that can be transformed, the LFNST index signaling can be skipped, and it can be inferred that the LFNST index is 0.
[0352] In the dual-tree case, it is determined whether each of the coding units or blocks for the luma component and the coding units or blocks for the chroma component exceeds the maximum transform block size. That is, each of the horizontal and vertical lengths of the luma coding units or blocks is compared to the maximum luma transform size, and if either of them is greater than the maximum luma transform size, LFNST is not applied. In the case of chroma coding units or blocks, the horizontal / vertical length of the corresponding luma block for the color format is compared to the maximum transformable luma transform size. For example, when the color format is 4:2:0, each of the horizontal / vertical lengths of the corresponding luma block is twice the horizontal / vertical length of the chroma block, and the transform size of the corresponding luma block is twice the transform size of the chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical lengths of the corresponding luma block are the same as the horizontal / vertical lengths of the chroma block.
[0353] In the case of a single tree, check whether the luminance block (encoding unit or partition block) exceeds the size of the maximum luminance transform block, and if it exceeds this size, the LFNST index signaling can be skipped.
[0354] C. If LFNST, which is included in the current VVC standard, is applied, the LFNST index can be sent only if both the horizontal and vertical lengths of the partition block are greater than or equal to 4.
[0355] If up to LFNST for 2×M (1×M) or M×2 (M×1) blocks is applied in addition to the LFNST included in the current VVC standard, then the LFNST index is sent only if the partition block has a size greater than or equal to 2×M (1×M) or M×2 (M×1) blocks. Here, P×Q blocks greater than or equal to R×S blocks means P≥R and Q≥S.
[0356] In summary, the LFNST index can be sent only if the partition block has a size greater than or equal to the minimum size applicable to LFNST. In the two-tree case, the LFNST index can be signaled only if the partition block of the luma or chroma component has a size greater than or equal to the minimum size applicable to LFNST. In the single-tree case, the LFNST index can be signaled only if the partition block of the luma component has a size greater than or equal to the minimum size applicable to LFNST.
[0357] In this document, an M×N block equal to or greater than a K×L block means that M is equal to or greater than K and N is equal to or greater than L. An M×N block greater than a K×L block means that M is equal to or greater than K and N is equal to or greater than L, where M is greater than K or N is greater than L. An M×N block less than or equal to a K×L block means that M is less than or equal to K and N is less than or equal to L. An M×N block less than a K×L block means that M is less than or equal to K and N is less than or equal to L, where M is less than K or N is less than L.
[0358] D. If the last non-zero coefficient position is not the DC position (top-left position of the block), when it is a dual-tree type luma block, LFNST transmission can be performed if the corresponding last non-zero coefficient position is not the DC position in any or all of the partition blocks. When it is a dual-tree type chroma block, the corresponding LNFST index can be transmitted if either the non-zero coefficient positions of all partition blocks of Cb (here, when the ISP mode is not applied to the chroma component, the number of partition blocks is considered to be 1) or the last non-zero coefficient positions of all partition blocks of Cr (here, when the ISP mode is not applied to the chroma component, the number of partition blocks is considered to be 1) is not the DC position.
[0359] In the case of a single-tree type, if the last non-zero position is not a DC position even in any of the partition blocks of the luminance component, Cb component, and Cr component, the corresponding LNFST index can be sent.
[0360] In this paper, if the Coded Block Flag (CBF) indicating whether transform coefficients exist for each partition block is 0, the position of the last non-zero coefficient in the corresponding partition block is not checked to determine whether to signal the LFNST index. That is, if the CBF value is 0, the position of the last non-zero coefficient is not considered when checking the conditions for LFNST index signaling because the transform is not applied to the block.
[0361] For example, 1) in the case of a dual-tree type and luma component, when determining whether to signal the LFNST index for each partition block if the corresponding CBF value is 0, 2) in the case of a dual-tree type and chroma component, if the CBF value of Cb is 0 and the CBF value of Cr is 1, it can be determined whether to signal the LFNST index by checking only the non-zero coefficient position of Cr in each partition block, and 3) in the case of a single-tree type, it can be determined whether to signal the LFNST index by checking only the last non-zero coefficient position of the blocks with a CBF value of 1 in all partition blocks of the luma component, Cb component, and Cr component.
[0362] In ISP mode, image information can be configured not to check the position of the last non-zero coefficient, and this can be implemented as follows.
[0363] i. In ISP mode, LFNST index signaling can be allowed by skipping the check for non-zero coefficient positions of both luma and chroma blocks. That is, LFNST index signaling can be allowed even if the last non-zero coefficient position is a DC position for all partitions or the corresponding CBF value is 0.
[0364] ii. In ISP mode, based on the aforementioned scheme, the check for non-zero coefficient positions can be skipped only for luma blocks, while the check for non-zero coefficient positions can be performed for chroma blocks. For example, in the case of dual-tree type and luma blocks, LFNST index signaling is allowed without checking for non-zero coefficient positions, and in the case of dual-tree type and chroma blocks, whether to signal the LFNST index can be determined by checking whether a DC position exists for the last non-zero coefficient position according to the aforementioned scheme.
[0365] iii. In ISP mode and single-tree type, scheme i or ii can be applied. That is, in ISP mode and single-tree type, when scheme i is applied, LFNST index signaling can be allowed by skipping the check for the last non-zero coefficient position of both the luma and chroma blocks. Alternatively, scheme ii can be applied to determine the signaling of the LFNST index by skipping the check for the non-zero coefficient position for the luma component partition block and performing the check for the non-zero coefficient position for the chroma component partition block as described above (when ISP is not applied to the chroma component, the number of partition blocks can be considered as 1).
[0366] E. If, even for one of all partition blocks, the transform coefficients are identified as existing in a location other than where LFNST transform coefficients could exist, then the LFNST index signaling can be skipped.
[0367] For example, in the case of 4×4 and 8×8 blocks, the LFNST transform coefficients can exist at 8 positions starting from the DC position, according to the transform coefficient scan order in the VVC standard, with all remaining positions padded with 0. Additionally, when the transform block is larger than or equal to 4×4 and is not a 4×4 or 8×8 block, the LFNST transform coefficients can exist at 16 positions starting from the DC position, according to the transform coefficient scan order in the VVC standard, with all remaining positions padded with 0.
[0368] Therefore, after performing residual coding, the LFNST index signaling can be skipped when non-zero transform coefficients are present in regions to be filled with 0.
[0369] If LFNST also applies when the partition block has a size of 2×M (1×M) or M×2 (M×1), the region containing the LFNST transform coefficients can be specified as follows. When assuming LFNST is applied, if the region outside the region containing the transform coefficients can be filled with 0, and if there are non-zero transform coefficients in the region to be filled with 0, the LFNST index signaling can be skipped.
[0370] i. If LFNST applies to 2×M or M×2 blocks, and if M=9, then 8 LFNST transform coefficients can be generated only for 2×8 or 8×2 partitioned blocks. When according to Figure 18 When arranging the transformation coefficients according to the scanning order, eight transformation coefficients can be arranged from the DC position according to the scanning order, and the remaining eight positions can be filled with 0.
[0371] When 16 LFNST transform coefficients can be generated for a 2×N or N×2 (N>8) partition block, and when according to Figure 18When arranging transform coefficients according to the scan order, 16 transform coefficients can be arranged from the DC position according to the scan order, and the remaining positions can be filled with 0. That is, in a 2×N or N×2 (N>8) partition block, the area except for the upper left 2×8 or 8×2 block can be filled with 0. Instead of 8 LFNST transform coefficients, 16 transform coefficients can also be generated for a 2×8 or 8×2 partition block. In this case, there is no area to be filled with 0. As mentioned above, when applying LFNST, if it is determined that there are non-zero transform coefficients in the area determined to be filled with 0 in a partition block, the LFNST index signaling can be skipped, and the LFNST index can be inferred to be 0.
[0372] ii. If LFNST applies to 1×M or M×1 blocks, and if M = 16, then only 8 LFNST transform coefficients can be generated for 1×16 or 16×1 partition blocks. When the transform coefficients are arranged according to the scan order from left to right or from top to bottom, 8 transform coefficients can be arranged from the DC position according to the corresponding scan order, and the remaining 8 positions can be filled with 0.
[0373] When 16 LFNST transform coefficients can be generated for 1×N or N×1 (N>16) partitions, and when the transform coefficients are arranged according to the scan order from left to right or from top to bottom, the 16 transform coefficients can be arranged from the DC position according to the corresponding scan order, and the remaining positions can be filled with 0. That is, the area in the 1×N or N×1 (N>16) partition, except for the top left 1×16 or 16×1 block, can be filled with 0.
[0374] Instead of 8 LFNST transform coefficients, 16 transform coefficients can also be generated for a 1×16 or 16×1 partition block. In this case, there is no region to be filled with 0. As mentioned above, when applying LFNST, if it is determined that there are non-zero transform coefficients in a region determined to be filled with 0 even within a partition block, the LFNST index signaling can be skipped, and the LFNST index can be inferred to be 0.
[0375] Furthermore, in ISP mode, instead of DCT-2, DST-7 is applied in the absence of signaling for the MTS index in the current VVC standard by independently considering length conditions for each of the horizontal and vertical directions. It is determined whether the horizontal or vertical length is greater than or equal to 4 or greater than or equal to 16, and a transformation kernel is determined based on the result. Therefore, when in ISP mode and when LFNST is applicable, the transformation combination can be configured as follows.
[0376] 1. When the LFNST index is 0 (including cases where the LFNST index is inferred to be 0), the conditions for determining a first transformation are met when the ISP is included in the current VVC standard. That is, it is checked whether the length condition (length greater than or equal to 4 and less than or equal to 16) is independently satisfied for each of the horizontal and vertical directions. If the condition is satisfied, DST-7 is applied instead of DCT-2 for a first transformation; if the condition is not satisfied, DCT-2 is applied.
[0377] 2. When the LFNST index is greater than 0, the following two configurations are possible for a single transformation.
[0378] A. DCT-2 is applicable to both horizontal and vertical directions.
[0379] B. When ISP is included in the current VVC standard, it can meet the conditions for determining a transformation. That is, check whether the length condition (length greater than or equal to 4 and less than or equal to 16) is met independently for each of the horizontal and vertical directions, and if the condition is met, DST-7 is applied instead of DCT-2, and if the condition is not met, DCT-2 can be applied.
[0380] In ISP mode, image information can be configured so that the LFNST index is sent per partition block instead of per coding unit. In this case, it can be determined whether to signal the LFNST index by assuming that only one partition block exists in the unit that sends the LFNST index in the aforementioned LFNST index signaling scheme.
[0381] In addition, the signaling order of the LFNST index and the MTS index will be described below.
[0382] According to the example, the LFNST index, which is signaled in residual coding, can be encoded after the coding position of the last non-zero coefficient position, and the MTS index can be encoded immediately after the LFNST index. In this configuration, the LFNST index can be signaled for each transform unit. Alternatively, even if no signal is given in residual coding, the LFNST index can be encoded after the coding of the last valid coefficient position, and the MTS index can be encoded after the LFNST index.
[0383] The syntax for residual coding based on the example is as follows.
[0384] [Table 8]
[0385]
[0386]
[0387] The meanings of the main variables shown in Table 8 are as follows.
[0388] 1. cbWidth, cbHeight: The width and height of the current encoding block.
[0389] 2. log2TbWidth, log2TbHeight: The base-2 logarithmic values of the width and height of the current transform block, which can be reduced to the upper left region where non-zero coefficients can exist by reflecting zeroing.
[0390] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is enabled. If the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in the Sequence Parameter Set (SPS).
[0391] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variable chType and the position (x0, y0). chType can have values of 0 and 1, where 0 indicates the luma component and 1 indicates the chroma component. The position (x0, y0) indicates the location on the image, and MODE_INTRA (intra-frame prediction) and MODE_INTER (inter-frame prediction) can be used as the values of CuPredMode[chType][x0][y0].
[0392] 5. IntraSubPartitionsSplit[x0][y0]: The content at position (x0, y0) is the same as in item 4. It indicates which ISP partition was applied at position (x0, y0), and ISP_NO_SPLIT indicates that the coding unit corresponding to position (x0, y0) was not divided into a partition block.
[0393] 6. intra_mip_flag[x0][y0]: The content at position (x0, y0) is the same as in point 4 above. intra_mip_flag is a flag indicating whether matrix-based intra-frame prediction (MIP) prediction mode is applied. If the flag value is 0, it indicates that MIP is not enabled; if the flag value is 1, it indicates that MIP is enabled.
[0394] 7. cIdx: A value of 0 indicates luminance, and values of 1 and 2 indicate the Cb and Cr of the chromaticity components, respectively.
[0395] 8. treeType: Indicates whether it is a single tree or a dual tree (SINGLE_TREE: single tree, DUAL_TREE_LUMA: dual tree for the luminance component, DUAL_TREE_CHROMA: dual tree for the chrominance component).
[0396] 9. tu_cbf_cb[x0][y0]: The content at position (x0, y0) is the same as in item 4. It indicates the coded block flag (CBF) of the Cb component. If its value is 0, it means that there are no non-zero coefficients in the corresponding transform unit of the Cb component, and if its value is 1, it indicates that there are non-zero coefficients in the corresponding transform unit of the Cb component.
[0397] 10. lastSubBlock: This indicates the position of the subblock (coefficient group (CG)) containing the last non-zero coefficient in the scan order. 0 indicates a subblock containing the DC component, and a value greater than 0 indicates a subblock that does not contain the DC component.
[0398] 11. lastScanPos: This indicates the position of the last valid coefficient within a sub-block in scan order. If a sub-block contains 16 positions, it can have values from 0 to 15.
[0399] 12. lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If not parsed, it is inferred to be 0. That is, the default value is set to 0, indicating that LFNST is not applied.
[0400] 13. LastSignificantCoeffX, LastSignificantCoeffY: These indicate the x and y coordinates of the last significant coefficient in the transform block. The x-coordinate starts at 0 and increases from left to right, and the y-coordinate starts at 0 and increases from top to bottom. If both variables are 0, it means the last significant coefficient is located at DC.
[0401] 14. cu_sbt_flag: A flag indicating whether Subblock Transformation (SBT) included in the current VVC standard is enabled. If the flag value is 0, it indicates that SBT is not enabled, and if the flag value is 1, it indicates that SBT is enabled.
[0402] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: These flags indicate whether explicit MTS is applied to inter-frame CUs and intra-frame CUs, respectively. If the corresponding flag value is 0, it indicates that MTS is not enabled for inter-frame or intra-frame CUs; if the corresponding flag value is 1, it indicates that MTS is enabled.
[0403] 16. tu_mts_idx[x0][y0]: The MTS index syntax element to be parsed. If not parsed, it is inferred to be 0. That is, the default value is set to 0, indicating that DCT-2 is enabled in both the horizontal and vertical directions.
[0404] As shown in Table 8, in the case of a single tree, the location condition of the last valid coefficient for luminance can be used only to determine whether to signal the LFNST index. That is, if the location of the last valid coefficient is not DC and the last valid coefficient exists in the top-left sub-block (CG) (e.g., a 4×4 block), then the LFNST index is signaled. In this case, for both 4×4 and 8×8 transform blocks, the LFNST index is signaled only if the last valid coefficient exists at positions 0 to 7 in the top-left sub-block.
[0405] In the case of a dual-tree system, the LFNST index is signaled independently of each of the luminance and chrominance components. In the case of chrominance, the LFNST index can be signaled by applying the last valid coefficient position condition only to the Cb component. For the Cr component, the corresponding condition is not checked, and if the CBF value of Cb is 0, the LFNST index can be signaled by applying the last valid coefficient position condition to the Cr component.
[0406] In Table 8, “Min(log2TbWidth, log2TbHeight)>=2” can be represented as “Min(tbWidth,tbHeight)>=4”, and “Min(log2TbWidth,log2TbHeight)>=4” can be represented as “Min(tbWidth,tbHeight)>=16”.
[0407] In Table 8, log2ZoTbWidth and log2ZoTbHeight represent the logarithmic values of the width and height of the top-left region, which can contain the last valid coefficients by clearing them to zero, with a base of 2 (base-2).
[0408] As shown in Table 8, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places. The first is before parsing the MTS index or LFNST index value, and the second is after parsing the MTS index.
[0409] The first update occurs before parsing the MTS index (tu_mts_idx[x0][y0]) value, so log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.
[0410] After resolving the MTS index, log2ZoTbWidth and log2ZoTbHeigh are set for MTS indexes greater than 0 (DST-7 / DCT-8 combination). When DST-7 / DCT-8 is applied independently in each of the horizontal and vertical directions in a single transformation, there can be up to 16 valid coefficients per row or column in each direction. That is, after applying DST-7 / DCT-8 of length 32 or greater, up to 16 transformation coefficients can be derived for each row or column, starting from the left or top. Therefore, in a 2D block, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions, there can be valid coefficients in only up to a 16×16 top-left region.
[0411] Furthermore, when DCT-2 is applied independently in each of the horizontal and vertical directions in the current transformation, there can be up to 32 valid coefficients per row or column in each direction. That is, when applying a DCT-2 of length 64 or greater, up to 32 transformation coefficients can be derived for each row or column, starting from the left or top. Therefore, in a 2D block, when DCT-2 is applied to both the horizontal and vertical directions, valid coefficients can exist in only a maximum of 32×32 in the upper left region.
[0412] Furthermore, when DST-7 / DCT-8 is applied to one side for both the horizontal and vertical directions and DCT-2 is applied to the other side, there can be 16 effective coefficients in the forward direction and 32 effective coefficients in the backward direction. For example, in the case of a 64×8 transform block, if DCT-2 is applied in the horizontal direction and DST-7 is applied in the vertical direction (which may occur when implicit MTS is applied), there can be effective coefficients in up to a 32×8 region in the upper left.
[0413] If, as shown in Table 8, log2ZoTbWidth and log2ZoTbHeight are updated in two places, that is, before resolving the MTS index, the ranges of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight, as shown in the table below.
[0414] [Table 9]
[0415]
[0416] Additionally, in this case, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values during the binarization process of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.
[0417] [Table 10]
[0418]
[0419] According to the example, when applying ISP mode and LFNST, the canonical text can be configured as shown in Table 11 when applying the signaling in Table 8. Compared with Table 8, the condition that the LFNST index is signaled only when ISP mode is not included has been removed (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT in Table 8).
[0420] In a single tree, when the LFNST index sent for the luma component (cIdx=0) is reused for the chroma component, the LFNST index sent for the first ISP block with valid coefficients can be applied to the chroma transform block. Alternatively, even in a single tree, the LFNST index can be signaled for the chroma component separately from the LFNST index signaled for the luma component. The variables in Table 11 are described the same as those in Table 8.
[0421] [Table 11]
[0422]
[0423] According to another example, in Table 11, when the last valid coefficient is allowed to be located only in the DC location of all partition blocks in the ISP, the conditions for resolving the LFNST index can be varied as follows.
[0424] [Table 12]
[0425]
[0426] As shown in the example, the LFNST index and / or MTS index can be signaled at the encoding unit level. As mentioned above, the LFNST index can have three values: 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate that the selected LFNST set includes the first and second candidates of the two LFNST kernel candidates, respectively. The LFNST index is encoded using truncated univariate binarization, and the values 0, 1, and 2 can be encoded as bin strings of 0, 10, and 11, respectively.
[0427] Based on the example, LFNST can be applied only when DCT-2 is applied to both the horizontal and vertical directions in a single transformation. Therefore, if the MTS index is signaled after the LFNST index is signaled, the MTS index can be signaled only when the LFNST index is 0, and when the LFNST index is not 0, a single transformation can be performed by applying DCT-2 to both the horizontal and vertical directions without signaling the MTS index.
[0428] The MTS index can have values 0, 1, 2, 3, and 4, where 0, 1, 2, 3, and 4 indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, and DCT-8 / DCT-8 are applied in the horizontal and vertical directions, respectively. Additionally, the MTS index can be encoded using truncated unary binarization, and the values 0, 1, 2, 3, and 4 can be encoded as bin strings of 0, 10, 110, 1110, and 1111, respectively.
[0429] Signaling the LFNST index at the coding unit level can be indicated as shown in the table below. Signaling the LFNST index can also be done in the latter part of the coding unit syntax table.
[0430] [Table 13]
[0431]
[0432] The variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag in Table 13 can be set as shown in Table 16 below.
[0433] The variable LfnstDcOnly is equal to 1 for a transform block with a coded block flag (CBF) of 1 (0 if at least one valid coefficient exists in the corresponding block, otherwise 0), and all last valid coefficients are located at the DC position (top left position); otherwise, it is equal to 0. Specifically, in the case of dual-tree luma, the position of the last valid coefficient is checked for a luma transform block, and in the case of dual-tree chroma, the position of the last valid coefficient is checked for both the Cb and Cr transform blocks. In the case of single-tree, the position of the last valid coefficient can be checked for luma, Cb, and Cr transform blocks.
[0434] If a valid coefficient exists at the zeroing position when applying LFNST, the variable LfnstZeroOutSigCoeffFlag equals 0; otherwise, it equals 1.
[0435] In Table 13 and subsequent tables, lfnst_idx[x0][y0] indicates the LFNST index of the corresponding coding unit, while tu_mts_idx[x0][y0] indicates the MTS index of the corresponding coding unit.
[0436] According to the example, in order to encode the MTS index consecutively after the LFNST index at the coding unit level, the coding unit syntax table can be configured as shown in Table 14.
[0437] [Table 14]
[0438]
[0439] Comparing Table 14 with Table 13, the condition used to check if the value of tu_mts_idx[x0][y0] is 0 (i.e., checking if DCT-2 is applied in both the horizontal and vertical directions) in the condition used to signal lfnst_idx[x0][y0] is changed to a condition used to check if the value of transform_skip_flag[x0][y0] is 0 (!transform_skip_flag[x0][y0]). transform_skip_flag[x0][y0] indicates whether the encoding unit is encoded in a transform skip mode where the transform is skipped, and this flag is signaled before the MTS and LFNST indices. In other words, since lfnst_idx[x0][y0] is signaled before the value of tu_mtx_idx[x0][y0] is signaled, the condition regarding the value of transform_skip_flag[x0][y0] can be checked only.
[0440] As shown in Table 14, multiple conditions are checked when encoding tu_mts_idx[x0][y0], and as mentioned above, tu_mts_idx[x0][y0] is only signaled when the value of lfnst_idx[x0][y0] is 0.
[0441] tu_cbf_luma[x0][y0] is a flag indicating whether there are valid coefficients for the luminance component, and cbWidth and cbHeight indicate the width and height of the coding unit of the luminance component, respectively.
[0442] In Table 14, (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT) indicates that ISP mode is not applied, and (!cu_sbt_flag) indicates that SBT is not applied.
[0443] According to Table 14, when both the width and height of the coding unit of the luminance component are 32 or less, a signal is sent to tu_mts_idx[x0][y0]. In other words, whether to apply MTS is determined by the width and height of the coding unit of the luminance component.
[0444] According to another example, when transform block (TU) tiling occurs (e.g., when the maximum transform size is set to 32, a 64×64 coding unit is divided into four 32×32 transform blocks and encoded), the MTS index can be signaled based on the size of each transform block. For example, when both the width and height of the transform block are 32 or less, the same MTS index value can be applied to all transform blocks in the coding unit, thus applying the same transform. Additionally, when transform block tiling occurs, the value of tu_cbf_luma[x0][y0] in Table 14 can be the CBF value of the top-left transform block, or it can be set to 1 even if the CBF value of one of the transform blocks is 1.
[0445] According to the example, when the ISP mode is applied to the current block, LFNST can be applied, in which case Table 14 can be changed as shown in Table 15.
[0446] [Table 15]
[0447]
[0448] As shown in Table 15, even in ISP mode (IntraSubPartitionsSplitType!=ISP_NO_SPLIT), lfnst_idx[x0][y0] can be configured to be signaled, and the same LFNST index value can be applied to all ISP partition blocks.
[0449] Furthermore, as shown in Table 15, since the signal notification for tu_mts_idx[x0][y0] is only sent in modes other than ISP mode, the MTS index encoding part is the same as in Table 14.
[0450] As shown in Tables 14 and 15, when the MTS index is signaled immediately after the LFNST index, information about the first transformation is unavailable during residual coding. In other words, the MTS index is signaled after residual coding. Therefore, the part of the residual coding section that performs zeroing while retaining only 16 coefficients for a 32-coefficient DST-7 or DCT-8 can be modified as shown in Table 16 below.
[0451] [Table 16]
[0452]
[0453]
[0454] As shown in Table 16, in the process of determining log2ZoTbWidth and log2ZoTbHeight (where log2ZoTbWidth and log2ZoTbHeight represent the base-2 logarithmic values of the width and height of the upper left region after the zeroing is performed), the value of tu_mts_idx[x0][y0] can be omitted.
[0455] The binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 16 can be determined based on log2ZoTbWidth and log2ZoTbHeight as shown in Table 10.
[0456] In addition, as shown in Table 16, when determining log2ZoTbWidth and log2ZoTbHeight in the residual coding, a condition for checking sps_mts_enable_flag can be added.
[0457] The TR indicator in Table 10 is the truncated Rice binarization method, and the final effective coefficient information can be binarized based on cMax and cRiceParam as defined in Table 10 according to the method described in the table below.
[0458] [Table 17]
[0459]
[0460] In another example, the coding unit syntax table and the residual coding syntax table are as follows.
[0461] [Table 18]
[0462]
[0463] [Table 19]
[0464]
[0465] In Table 18, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed in the residual coding of Table 19. When there are valid coefficients in the region to be filled with 0 by zeroing (LastSignificantCoeffX>15||LastSignificantCoeffY>15), the value of the variable MtsZeroOutSigCoeffFlag changes from 1 to 0. In this case, no signal is sent to the MTS index, as shown in Table 19.
[0466] Based on the example, the MTS index encoding portion in Table 18 can be changed as shown in the following table.
[0467] [Table 20]
[0468]
[0469] Unlike Table 18, the variable MtsZeroOutSigCoeffFlag in Table 20 is initialized to 0 instead of 1 (MtsZeroOutSigCoeffFlag=0). When there are valid coefficients in the region to be filled with 0 by zeroing (LastSignificantCoeffX>15||LastSignificantCoeffY>15), the value of the variable MtsZeroOutSigCoeffFlag remains 0, and in this case, no signal is sent to the MTS index.
[0470] The following figures are provided to illustrate specific examples of this specification. Since the names of particular devices or signals / messages / fields described in the figures are presented by way of example, the technical features of this specification are not limited to the specific names used in the following figures.
[0471] Figure 19 This is a flowchart illustrating the operation of a video decoding device according to an embodiment of this document.
[0472] Figure 19 Each step disclosed in the document is based on the above. Figures 2 to 18 Some content described above. Therefore, omissions or simplifications will be made. Figures 2 to 18 The description repeats specific content.
[0473] According to the implementation, the decoding device 200 can perform residual coding based on the residual information received from the bit stream, and specifically can parse the residual information received at the residual coding level, and can arrange the transform coefficients of the current block according to a predetermined scan order (S1910).
[0474] Decoding device 200 can decode information about the quantization transform coefficients of the current block from the bitstream, and can deduce the quantization transform coefficients of the target block based on the information about the quantization transform coefficients of the current block. The information about the quantization transform coefficients of the target block can be incorporated into the Sequence Parameter Set (SPS) or the stripe header, and can include at least one of the following: information about whether a reduced transform (RST) is applied, information about the reduction factor, information about the minimum transform size for applying the reduced transform, information about the maximum transform size for applying the reduced transform, and information about the transform index indicating either the transform kernel matrix included in the transform set or the size of the simplified inverse transform.
[0475] In addition, the decoding device can also receive information about the intra-prediction mode of the current block and information about whether ISP encoding or ISP mode is applied to the current block. The decoding device can deduce whether the current block is divided into a predetermined number of sub-partition transform blocks by receiving and parsing flag information indicating whether ISP encoding or ISP mode is applied. Here, the current block can be a coded block. Furthermore, the decoding device can deduce the size and number of sub-partition blocks by flag information indicating the direction in which the current block will be divided.
[0476] The decoding device 200 can derive the transform coefficients by dequantizing the residual information about the current block (i.e., the quantized transform coefficients), and can arrange the derived transform coefficients according to a predetermined scan order.
[0477] More specifically, the derived transform coefficients can be arranged in 4×4 blocks according to the inverse diagonal scan order, and the transform coefficients within the 4×4 blocks can also be arranged according to the inverse diagonal scan order. That is to say, the transform coefficients that have been dequantized can be arranged according to the inverse scan order used in video codecs such as VVC or HEVC.
[0478] The transform coefficients derived from this residual information can be either dequantized transform coefficients as described above, or quantized transform coefficients. In other words, the transform coefficients can be any data that can be checked to see if it is non-zero data in the current block, regardless of whether it has been quantized.
[0479] Decoding devices can derive residual samples by applying an inverse transform to the quantization transform coefficients.
[0480] As described above, the decoding device can derive residual samples by applying LFNST as an inseparable transform or MTS as a separable transform, and such transforms can be performed based on the LFNST index indicating the LFNST kernel (i.e., the LFNST matrix) and the MTS index indicating the MTS kernel, respectively.
[0481] The decoding device can receive and parse at least one of the LFNST index or MTS index at the encoding unit level, and can parse the LFNST index indicating the LFNST core before (i.e. immediately before) the MTS index indicating the MTS core (S1920).
[0482] The MTS index can be resolved based on specific conditions of the LFNST index, for example, when the LFNST index is 0.
[0483] The residual coding level can include syntax for the last valid coefficient position information, and the LFNST index can be parsed after parsing the last valid coefficient position information.
[0484] According to the example, the MTS index can be resolved when the current block is a luma block and the LFNST index is 0. In other words, if the LFNST index is greater than 0 when the current block is a luma block, the MTS index does not need to be resolved.
[0485] According to the example, when the tree type of the current block is a dual tree, the LFNST index for each of the luma and chroma blocks can be resolved.
[0486] Furthermore, when deriving the transformation coefficients, the width and height of the upper left region in the current block, which can contain the last valid coefficients by clearing them to zero, can be derived, and the width and height of the upper left region can be derived before parsing the MTS index.
[0487] Furthermore, the position of the last effective coefficient can be derived from the width and height of the upper left region, and the position information of the last effective coefficient can be binarized based on the width and height of the upper left region.
[0488] Alternatively, according to the example, when the tree type of the current block is a single tree type, the decoding device can perform residual encoding on the luma and chroma blocks of the current block, and then parse the LFNST index.
[0489] When resolving the LFNST index at the coding unit level (rather than at the transform block level or residual coding level) after performing residual coding, it can receive the zeroing information needed during the transform process and the LFNST index reflecting the complete transform coefficient positions of the luma and chroma blocks, rather than transform coefficient information about the luma or chroma blocks.
[0490] When the current block is a dual-tree type and the chroma components are encoded, the decoding device can perform residual encoding on the Cb and Cr components of the chroma block, and then parse the LFNST index.
[0491] When resolving the LFNST index at the coding unit level (rather than at the transform block level or residual coding level) after residual coding is performed, the zeroing information required during the transform process and the LFNST index reflecting the position of the complete transform coefficients of the Cb and Cr components of the chroma block can be received, rather than transform coefficient information about either the Cb or Cr components of the chroma block.
[0492] When the current block is partitioned into multiple sub-partition blocks, the decoding device can perform residual encoding on the multiple sub-partition blocks and then parse the LFNST index.
[0493] Similar to the above description, when resolving the LFNST index at the coding unit level (rather than at the transform block level or residual coding level) after performing residual coding, one can receive the zeroing information required during the transform process and the LFNST index reflecting the position of the complete transform coefficients of all sub-blocks, rather than transform coefficient information about some or individual sub-blocks.
[0494] According to the example, when the current block is partitioned into multiple sub-blocks, the LFNST index can be resolved regardless of whether transform coefficients exist in the region other than the DC location in each of the multiple sub-blocks. In other words, if the ISP allows the last valid coefficients to be located only in the DC locations of all sub-blocks when applied to the current block, then signaling the LFNST index can be permitted.
[0495] During residual coding, the decoding device can deduce whether a first variable, indicating whether a transform coefficient exists in the region of the current block other than the DC position, and can deduce whether a second variable, indicating whether a transform coefficient exists in the second region of the current block or the sub-blocks divided by the current block other than the first region in the upper left, exists.
[0496] The decoding device can resolve the LFNST index when there are transform coefficients in the region other than the DC position and no transform coefficients in the second region.
[0497] Specifically, in order to determine whether to resolve the LFNST index, the decoding device can deduce a first variable indicating whether transform coefficients (i.e., valid coefficients) exist in the region of the current block other than the DC position.
[0498] The first variable can be LfnstDcOnly, which can be derived during residual coding. When the index of the sub-block containing the last valid coefficient in the current block is 0 and the position of the last valid coefficient in the sub-block is greater than 0, the first variable can be derived to 0, and the LFNST index can be resolved when the first variable is 0. A sub-block refers to a 4×4 block used as a coding unit in residual coding, also known as a coefficient group (CG). Sub-block index 0 indicates the top-left 4×4 sub-block.
[0499] The first variable can be initially set to 1, and depending on whether there is a valid coefficient in the region other than the DC location, it can remain at 1 or be changed to 0.
[0500] The variable LfnstDcOnly indicates whether there are non-zero coefficients at the non-DC component positions of at least one transform block within a coding unit. It can be 0 when there are non-zero coefficients at the non-DC component positions of at least one transform block within a coding unit, and it can be 1 when there are no non-zero coefficients at the non-DC component positions of all transform blocks within a coding unit.
[0501] The decoding device can deduce whether a second variable with a valid coefficient exists in the second region of the current block, excluding the first region in the upper left, thereby checking whether the second region has been cleared.
[0502] The second variable can be LfnstZeroOutSigCoeffFlag, which indicates whether zeroing was performed when LFNST was applied. The second variable can be initially set to 1 and can be changed to 0 when a valid coefficient exists in the second region.
[0503] If the index of a sub-block with the last non-zero coefficient is greater than 0 and both the width and height of the transform block are greater than or equal to 4, or if the last position of the non-zero coefficient in a sub-block with the last non-zero coefficient is greater than 7 and the size of the transform block is 4×4 or 8×8, then the variable LfnstZeroOutSigCoeffFlag can be deduced to be 0. A sub-block represents a 4×4 block used on the basis of the code in the case of residual coding and can be referred to as a coefficient group (CG). When the index of a sub-block is not 0, it indicates the top-left 4×4 sub-block.
[0504] In other words, if a non-zero coefficient is derived in a region outside the upper left region where the LFNST transform coefficients exist in the transform block, or if a non-zero coefficient exists outside the eighth position in the scan order for 4×4 and 8×8 blocks, the variable LfnstZeroOutSigCoeffFlag is set to 0.
[0505] According to the example, when ISP is applied to the coding unit, if it is identified that a transform coefficient exists in a location other than where LFNST transform coefficients may exist, even for one sub-block of all sub-blocks, the LFNST index signaling can be skipped. That is, if zeroing is not performed in a sub-block and a valid coefficient exists in the second region, the LFNST index is not signaled.
[0506] Furthermore, the first region can be derived based on the size of the current block.
[0507] For example, if the size of the current block is 4×4 or 8×8, the first region can extend from the top left position of the current block to the eighth sample position in the scan direction. When the current block is segmented, if the size of the sub-block is 4×4 or 8×8, the first region can extend from the top left position of the sub-block to the eighth sample position in the scan direction.
[0508] If the current block size is 4×4 or 8×8, since 8 data lines are output via the forward LFNST, the 8 transform coefficients received in the decoding device can be arranged up to the eighth sample position in the scanning direction of the current block, such as... Figure 11 (a) and Figure 12 As shown in (a).
[0509] Additionally, if the current block size is not 4×4 or 8×8, the first region can be a 4×4 region in the upper left position of the current block. If the current block size is not 4×4 or 8×8, since 16 data lines are output via the forward LFNST, the 16 transform coefficients received in the decoding device can be arranged in the 4×4 region in the upper left position of the current block, such as... Figure 11 (b) to (d) and Figure 12 As shown in (b).
[0510] In addition, it can be based on, for example Figure 7 The transformation coefficients can be arranged in the first region by means of the diagonal scanning direction shown.
[0511] As described above, when the current block is divided into sub-blocks, the decoding device can resolve the LFNST index if no transform coefficient exists in any of the corresponding second regions of the multiple sub-blocks. If a transform coefficient exists in the second region for any sub-block, the LFNST index is not resolved.
[0512] As described above, LFNST can be applied to sub-partition blocks in which the width and height are greater than or equal to 4, and the LFNST index of the current block as the encoding block can be applied to multiple sub-partition blocks.
[0513] Furthermore, since the zeroing of LFNST (including every zeroing that may involve the application of LFNST) is also directly applied to the sub-blocks, the first region is also applied to the sub-blocks. That is, if the segmented sub-blocks are 4×4 or 8×8 blocks, LFNST is applied to the transform coefficients of the eighth transform coefficient from the top left position of the sub-block to the scan direction, and if the sub-blocks are not 4×4 or 8×8 blocks, LFNST can be applied to the transform coefficients of the top left 4×4 region of the sub-block.
[0514] In addition, during residual coding, the decoding device can deduce whether a third variable indicating the presence of transform coefficients exists in the current block outside the top-left 16×16 region, and can resolve the MTS index when no transform coefficients exist in the region outside the 16×16 region.
[0515] The third variable can be MtsZeroOutSigCoeffFlag, which indicates whether zeroing is performed when MTS is applied. MtsZeroOutSigCoeffFlag indicates whether transform coefficients exist in the region outside the top-left region where the last valid coefficients can exist after MTS is applied (i.e., outside the top-left 16×16 region). The value of the third variable can be initially set to 1 and can be changed from 1 to 0 when transform coefficients exist in the region outside the top-left 16×16 region. When the value of the third variable is 0, no signal is sent to the MTS index.
[0516] The decoding device can derive residual samples by applying at least one of LFNST performed based on the LFNST index or MTS performed based on the MTS index (S1930).
[0517] Subsequently, the decoding device 200 can generate a reconstructed sample based on the residual sample of the current block and the predicted sample of the current block (S1940).
[0518] The following figures are provided to illustrate specific examples of this specification. Since the names of specific devices or signals / messages / fields described in the figures are presented by way of example, the technical features of this specification are not limited to the specific names used in the following figures.
[0519] Figure 20 This is a flowchart illustrating the operation of a video encoding device according to an embodiment of this document.
[0520] Figure 20 Each step disclosed in the document is based on the above. Figures 3 to 18 Some of the content described above. Therefore, the above will be omitted or briefly explained. Figures 3 to 18 Explanation of the specific content that is repeated in the description.
[0521] According to the implementation method, the encoding device 100 can derive the prediction sample of the current block based on the intra-prediction mode applied to the current block (S2010).
[0522] When the ISP is applied to the current block, the encoding device can perform prediction for each sub-partition transform block.
[0523] The encoding device can determine whether to apply ISP encoding or ISP mode to the current block (i.e., the encoding block), and based on the determination result, it can determine the direction in which the current block should be divided, and deduce the size and number of sub-blocks to be divided.
[0524] The encoding device 100 can derive the residual sample of the current block based on the predicted sample (S2020).
[0525] The encoding device 100 can derive the transform coefficients of the current block by applying at least one of LFNST or MTS to the residual sample, and can arrange the transform coefficients according to a predetermined scan order (S2030).
[0526] A transformation can be performed using multiple transform kernels, such as MTS, and in this case, the transform kernel can be selected based on the intra-frame prediction mode.
[0527] Furthermore, the encoding device 100 can determine whether to perform a quadratic transformation or an inseparable transformation (specifically, LFNST) on the transform coefficients of the current block, and can apply LFNST to the transform coefficients to derive the modified transform coefficients.
[0528] Unlike a single transform that separates and transforms the coefficients as the target of the transform in the vertical or horizontal direction, LFNST is an inseparable transform that applies the transform without separating the coefficients in a specific direction. Such an inseparable transform can be a low-frequency inseparable transform that applies the transform only to the low-frequency region rather than to the entire target block as the target of the transform.
[0529] When applying ISP to the current block, the encoding device can determine whether LFNST can be applied to the height and width of the divided sub-blocks.
[0530] The encoding device can determine whether LFNST can be applied to the height and width of the divided sub-blocks. In this case, the decoding device can resolve the LFNST index if the height and width of the sub-blocks are equal to or greater than 4.
[0531] The encoding device can encode at least one of the LFNST index indicating the LFNST core and the MTS index indicating the MTS core (S2040).
[0532] The MTS index can be encoded based on specific conditions of the LFNST index (e.g., when the LFNST index is 0).
[0533] According to the example, when the current block is a luma block and the LFNST index indicates 0, the encoding device can encode the MTS index.
[0534] According to the example, when the tree type of the current block is a dual-tree, the encoding device can encode the LFNST index for each of the luma and chroma blocks.
[0535] According to the example, when deriving transform coefficients, the encoding device can derive the width and height of the upper left region in the current block where the last valid coefficient can exist by clearing it to zero. It can derive the position of the last valid coefficient based on the width and height of the upper left region and can binarize the position information of the last valid coefficient.
[0536] Based on the example, the width and height of the top-left region can be derived before signaling the MTS index.
[0537] Alternatively, according to the example, when the tree type of the current block is a single tree type, the encoding device can derive all the transform coefficients of the luma and chroma blocks of the current block, and then the LFNST index can be encoded at the encoding unit level.
[0538] When encoding the LFNST index at the coding unit level (rather than at the transform block level or residual coding level) after all transform coefficients have been derived, the LFNST index that reflects the position of the complete transform coefficients of the luma and chroma blocks can be encoded instead of the transform coefficient information of the luma or chroma blocks.
[0539] When the tree type of the current block is a dual-tree type and the chrominance components are encoded, the encoding device can derive all the transform coefficients of the Cb and Cr components of the chrominance block, and then the LFNST index can be encoded at the encoding unit level.
[0540] When encoding the LFNST index at the coding unit level (rather than at the transform block level or residual coding level) after all transform coefficients have been derived, the LFNST index that reflects the position of the complete transform coefficients of the Cb and Cr components of the chroma block can be encoded instead of the transform coefficient information about either the Cb or Cr components of the chroma block.
[0541] When the current block is divided into multiple sub-blocks, the encoding device can derive all the transform coefficients of the multiple sub-blocks, and then the LFNST index can be encoded at the coding unit level.
[0542] Similar to the above description, when resolving the LFNST index at the coding unit level (rather than at the transform block level or residual coding level) after all transform coefficients have been derived, the zeroing information required during the transform process and the LFNST index reflecting the position of the complete transform coefficients of all sub-blocks can be encoded, instead of encoding the transform coefficient information of some or individual sub-blocks.
[0543] When the current block is partitioned into multiple sub-blocks, the encoding device can encode the LFNST index regardless of whether transform coefficients exist in regions other than DC locations in each of the multiple sub-blocks. In other words, signaling for the LFNST index is permitted if the last valid coefficients are allowed to reside only in the DC locations of all sub-blocks when applying the ISP to the current block.
[0544] In the process of deriving transform coefficients, the encoding device can derive a first variable indicating whether transform coefficients exist in the region of the current block other than the DC position, and a second variable indicating whether transform coefficients exist in the second region of the current block or in the sub-blocks divided by the current block other than the first region in the upper left.
[0545] The encoding device can encode the LFNST index when there are transform coefficients in the region other than the DC position and no transform coefficients in the second region.
[0546] Specifically, the first variable can be the variable LfnstDcOnly, and it can be deduced to be 0 when the index of the sub-block containing the last valid coefficient in the current block is 0 and the position of the last valid coefficient in the sub-block is greater than 0. When the first variable is 0, the LFNST index can be encoded.
[0547] The first variable can be initially set to 1, and depending on whether there are valid coefficients in the region other than the DC location, it can remain at 1 or be changed to 0.
[0548] The variable LfnstDcOnly indicates whether there are non-zero coefficients at the non-DC component positions of at least one transform block within a coding unit. It can be 0 when there are non-zero coefficients at the non-DC component positions of at least one transform block within a coding unit, and can be 1 when there are no non-zero coefficients at the non-DC component positions of all transform blocks within a coding unit.
[0549] The encoding device can zero out the second region of the current block that does not have modified transform coefficients after performing LFNST, and can deduce a second variable indicating whether transform coefficients exist in the second region.
[0550] like Figure 11 and Figure 12 As shown, the remaining regions of the current block's transform coefficients that have not been modified can be zeroed out. This zeroing reduces the computational load and overall operational complexity required for the transform process, thereby reducing power consumption. Furthermore, it reduces latency during the transform process, thus increasing image coding efficiency.
[0551] The second variable can be LfnstZeroOutSigCoeffFlag, which indicates that zeroing should be performed when LFNST is applied. The second variable can be initially set to 1 and can be changed to 0 when a valid coefficient exists in the second region.
[0552] The variable LfnstZeroOutSigCoeffFlag can be deduced to be 0 when the index of the sub-block containing the last non-zero coefficient is greater than 0 and both the width and height of the transform block are equal to or greater than 4, or when the position of the last non-zero coefficient in the sub-block containing the last non-zero coefficient is greater than 7 and the size of the transform block is 4×4 or 8×8.
[0553] In other words, when a non-zero coefficient is derived in a region other than the upper left region where LFNST transform coefficients can exist in the transform block, or when a non-zero coefficient exists outside the eighth position in the scan order for 4×4 and 8×8 blocks, the variable LfnstZeroOutSigCoeffFlag is set to 0.
[0554] The description of the first region and the zeroing process when applying ISP are essentially the same as those described in the decoding method, so redundant descriptions will be omitted.
[0555] According to the example, the encoding device can perform zeroing when applying the MTS in a single transform of the current block. The encoding device can perform zeroing by filling the area of the current block or sub-block with 0s, except for the top-left 16×16 area, and can encode the MTS index according to a third variable indicating whether transform coefficients exist in the zeroing area.
[0556] The third variable can be MtsZeroOutSigCoeffFlag, which indicates whether zeroing is performed when MTS is applied. MtsZeroOutSigCoeffFlag indicates whether transform coefficients exist in the region outside the top-left region (i.e., outside the top-left 16×16 region) where the last valid coefficient can exist after zeroing following MTS execution. The value of the third variable can be initially set to 1 and can be changed from 1 to 0 if transform coefficients exist in the region outside the top-left 16×16 region. When the value of the third variable is 0, the MTS index is not encoded or signaled.
[0557] The encoding device can construct and output image information such that at least one of the LFNST index and MTS index is signaled at the encoding unit level and the MTS index is signaled immediately after the LFNST index is signaled (S2050).
[0558] Additionally, the encoding device can derive quantized transform coefficients by performing quantization based on the transform coefficients of the current block or modified transform coefficients, and can encode and output image information including information about the quantized transform coefficients.
[0559] The encoding device can generate residual information that includes information about the quantization transform coefficients. The residual information can include the transform-related information / syntax elements mentioned above. The encoding device can encode image / video information including the residual information and output the encoded image / video information as a bitstream.
[0560] More specifically, the encoding device can generate information about the quantization transform coefficients and can encode the information about the generated quantization transform coefficients.
[0561] In this disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation may be omitted. When quantization / dequantization is omitted, the quantization transformation coefficients may be referred to as transformation coefficients. When transformation / inverse transformation is omitted, the transformation coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency, may still be referred to as transformation coefficients.
[0562] Furthermore, in this disclosure, quantization transform coefficients and transform coefficients can be referred to as transform coefficients and scaling transform coefficients, respectively. In this case, residual information can include information about the transform coefficients, and this information can be signaled via residual coding syntax. Transform coefficients can be derived based on residual information (or information about transform coefficients), and scaling transform coefficients can be derived through the inverse transform (scaling) of the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaling transform coefficients. These details can also be applied / expressed in other parts of this disclosure.
[0563] In the above embodiments, the method is explained based on a flowchart using a series of steps or blocks. However, this disclosure is not limited to the order of the steps, and a step may be performed in a different order or sequence than described above, or a step may be performed concurrently with other steps. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of this disclosure.
[0564] The methods described above according to this disclosure can be implemented in software form, and the encoding and / or decoding devices according to this disclosure can be included in devices for image processing such as televisions, computers, smartphones, set-top boxes, and display devices.
[0565] When the embodiments of this disclosure are implemented by software, the above methods can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor in various well-known ways. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0566] Furthermore, the decoding and encoding devices using this disclosure can include multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices (such as video communication), mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0567] Furthermore, the processing methods of this disclosure can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices for storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks. Additionally, embodiments of this disclosure can be implemented as computer program products by program code, and the program code can be executed on a computer according to embodiments of this disclosure. The program code can be stored on a computer-readable medium.
[0568] Figure 21 Examples of video / image coding systems to which this disclosure can be applied are illustrated.
[0569] Reference Figure 21 A video / image encoding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0570] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0571] Video sources can be obtained through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, the video / image capture process can be replaced by a process that generates related data.
[0572] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0573] A transmitter can send encoded video / image information or data, output in bitstream form, to a receiver in a receiving device via a digital storage medium or network, either as a file or a stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to a decoding device.
[0574] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0575] The renderer can render decoded video / images. The rendered video / images can then be displayed on a monitor.
[0576] Figure 22 The structure of a content streaming system applying this disclosure is illustrated.
[0577] Furthermore, the content streaming system using this disclosure can generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0578] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then sends it to a streaming server. As another example, in cases where the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. Furthermore, the streaming server can temporarily store the bitstream during the sending or receiving process.
[0579] The streaming server sends multimedia data to the user's device via a web server based on the user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case controls the commands / responses between the corresponding devices within the content streaming system.
[0580] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0581] For example, user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, board-type PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. The servers in the content streaming system can operate as distributed servers, and in this case, data received by each server can be processed in a distributed manner.
[0582] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims can be combined to be implemented or performed in a device, and the technical features of the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features of the method claims and the device claims can be combined to be implemented or performed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or performed in a method.
Claims
1. A decoding device for image decoding, the decoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: The intra-frame prediction mode of the current block is derived based on the prediction information; Based on the intra-frame prediction mode of the current block, derive prediction samples; The residual coding is performed by parsing the residual information received at the residual coding level to arrange the transform coefficients for the current block according to a predetermined scan order; The residual samples are derived by applying at least one of LFNST or MTS to the transform coefficients; and A reconstructed image is generated based on the residual samples and the predicted samples. Specifically, the LFNST is executed based on the LFNST index associated with the LFNST kernel. Specifically, the MTS is executed based on the MTS index associated with the MTS core. The LFNST index and the MTS index are signaled at the coding unit level. The MTS index is signaled immediately after the LFNST index is signaled. Wherein, based on the fact that the tree type of the current block is a single tree type, the LFNST index is parsed after the residual encoding of the luminance block and chrominance block for the current block is performed.
2. The decoding device according to claim 1, wherein, Based on the fact that the current block is a dual-tree type and the chroma components are encoded, the LFNST index is parsed after the residual encoding of the Cb and Cr components of the chroma block is performed.
3. The decoding device according to claim 1, wherein, Based on the fact that the current block is partitioned into multiple sub-partition blocks, the LFNST index is parsed after residual encoding is performed on the multiple sub-partition blocks.
4. The decoding device according to claim 3, wherein, The LFNST index is resolved based on the fact that the current block is partitioned into the plurality of sub-partition blocks, regardless of whether there are transform coefficients in the region of each of the plurality of sub-partition blocks except for the DC location.
5. The decoding device according to claim 1, wherein, The at least one processor is further configured to: derive a first variable relating to whether the transformation coefficient exists in the region of the current block other than the DC location; And derive a second variable related to whether the transformation coefficient exists in a second region other than the first region in the upper left of the current block or the sub-blocks partitioned by the current block, and Specifically, the LFNST index is parsed when the transformation coefficient exists in the region other than the DC location and does not exist in the second region.
6. The decoding device according to claim 1, wherein, The at least one processor is further configured to: derive a third variable relating to the existence of the transform coefficients in the region of the current block other than the top-left 16×16 region, and Specifically, the MTS index is parsed in response to the absence of the transformation coefficients in regions other than the 16×16 region.
7. An encoding device for image encoding, the encoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: Determine the intra-prediction mode for the current block; The prediction samples of the current block are derived based on the intra-prediction mode of the current block, wherein the prediction information for specifying the intra-prediction mode of the current block is encoded into the bitstream. The residual sample of the current block is derived based on the predicted sample; Perform a transform coefficient derivation operation by applying at least one of LFNST or MTS to the residual sample to derive the transform coefficients of the current block and arranging the transform coefficients according to a predetermined scan order; Encode at least one of the LFNST index associated with the LFNST core or the MTS index associated with the MTS core; and Image information is constructed and output such that the LFNST index and the MTS index are signaled at the coding unit level, and the MTS index is signaled immediately after the LFNST index is signaled. In the case where the tree type of the current block is a single tree, the LFNST index is signaled after performing residual encoding of the luminance block and chrominance block for the current block.
8. An apparatus for transmitting data for an image, the apparatus comprising: At least one processor is configured to obtain a bitstream, wherein the bitstream is generated by: determining an intra-prediction mode for a current block; deriving prediction samples for the current block based on the intra-prediction mode, wherein prediction information specifying the intra-prediction mode for the current block is encoded into the bitstream; deriving residual samples for the current block based on the prediction samples; performing a transform coefficient derivation operation by applying at least one of LFNST or MTS to the residual samples to derive transform coefficients for the current block and arranging the transform coefficients according to a predetermined scan order; encoding at least one of an LFNST index associated with an LFNST kernel or an MTS index associated with an MTS kernel; and constructing and outputting image information such that the LFNST index and the MTS index are signaled at the coding unit level and the MTS index is signaled immediately after the LFNST index is signaled; and A transmitter configured to transmit the data, including the bitstream of the image information. In the case where the tree type of the current block is a single tree, the LFNST index is signaled after performing residual encoding of the luminance block and chrominance block for the current block.