Video coding method and apparatus based on conversion
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-08-14
AI Technical Summary
【0019】 本文書によると、全般的な映像/ビデオ圧縮効率を上げることができる。
Smart Images

Figure 2026131834000001_ABST
Abstract
Description
Technical Field
[0001] This document relates to image coding technology, and more particularly, to an image coding method and apparatus based on transform in an image coding system.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from those of real images, such as game images, has been increasing.
[0004] Therefore, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The technical problem of this document is to provide a method and apparatus for increasing the coding efficiency of images.
[0006] Another technical problem of this document is to provide a method and apparatus for increasing the efficiency of transform index coding.
[0007] Another technical objective of this document is to provide a video coding method and apparatus utilizing LFNST and MTS.
[0008] Another technical issue addressed in this paper is to provide a video coding method and apparatus for LFNST index and MTS index signaling. [Means for solving the problem]
[0009] According to one embodiment of this document, a video decoding method is provided that is performed by a decoding device. The method includes a residual coding step of parsing residual information received at a residual coding level and arranging conversion coefficients for the current block in a predetermined scanning order; a step of applying at least one of LFNST or MTS to the conversion coefficients to derive a residual sample; and a step of generating a restored picture based on the residual sample, wherein the LFNST is performed based on an LFNST index that points to an LFNST kernel, and the MTS is performed based on an MTS index that points to an MTS kernel, and the LFNST index and the MTS index are signaled at the coding unit level, and the MTS index is signaled immediately after the LFNST index is signaled.
[0010] If the tree type of the current block is a single-tree type, the LFNST index is parsed after the residual coding is performed on the luma block and chroma block of the current block.
[0011] If the current block's tree type is a dual-tree type and chroma components are coded, the LFNST index is parsed after the residual coding for the Cb and Cr components of the chroma block is performed.
[0012] If the current block is divided into multiple subpartition blocks, the LFNST index is parsed after the residual coding is performed on the multiple subpartition blocks.
[0013] If the current block is divided into the plurality of subpartition blocks, the LFNST index is parsed regardless of whether the conversion coefficient exists in the region excluding the DC position of each of the plurality of subpartition blocks.
[0014] The residual coding step includes deriving a first variable indicating whether the conversion coefficient exists in a region excluding the DC position of the current block; and deriving a second variable indicating whether the conversion coefficient exists in a second region excluding the upper left region of the current block or the subpartition block divided from the current block; wherein if the conversion coefficient exists in the region excluding the DC position and does not exist in the second region, the LFNST index is parsed.
[0015] The residual coding step includes deriving a third variable indicating whether the conversion coefficient exists in the region excluding the upper left 16x16 region of the current block, and if the conversion coefficient does not exist in the region excluding the 16x16 region, the MTS index is parsed.
[0016] According to one embodiment of this document, a video encoding method is provided that is performed by an encoding device. The method includes a conversion coefficient derivation step of applying at least one of LFNST or MTS to a residual sample to derive conversion coefficients for the current block and arranging the conversion coefficients in a predetermined scanning order; a step of encoding at least one of LFNST indices that point to an LFNST kernel or MTS indices that point to an MTS kernel; and a step of structuring and outputting video information such that the LFNST index and the MTS index are signaled at the coding unit level and the MTS index is signaled immediately after the LFNST index is signaled.
[0017] According to another embodiment of this document, a digital storage medium is provided which stores video data containing encoded video information and a bitstream generated by a video encoding method performed by an encoding device.
[0018] According to another embodiment of this document, a digital storage medium is provided which stores video data containing encoded video information and a bitstream, causing a decoding device to perform the video decoding method. [Effects of the Invention]
[0019] According to this document, it is possible to improve the overall video compression efficiency.
[0020] According to this document, the efficiency of conversion index coding can be improved.
[0021] According to this document, a video coding method and apparatus utilizing LFNST and MTS can be provided.
[0022] According to this document, a video coding method and apparatus for LFNST index and MTS index signaling can be provided.
[0023] The effects obtained through a specific example in this specification are not limited to the effects listed above. For example, there may be various technical effects that can be understood or induced by a person having ordinary skill in the related art from this specification. Accordingly, the specific effects of this specification are not limited to those explicitly described in this specification, and may include various effects that can be understood or induced from the technical features of this specification.
Brief Description of Drawings
[0024] [Figure 1] This is a diagram schematically illustrating the configuration of a video / image encoding apparatus to which this document can be applied.
[0025] [Figure 2] This is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which this document can be applied.
[0026] [Figure 3] This schematically shows a multiple conversion technique according to an embodiment of this document.
[0027] [Figure 4] Exemplarily shows an intra prediction mode with 65 prediction directions.
[0028] [Figure 5] This is a diagram for explaining RST according to an embodiment of this document.
[0029] [Figure 6] This is a diagram showing the order of arranging the output data of forward first-order conversion as a one-dimensional vector by way of an example.
[0030] [Figure 7] This diagram illustrates the order in which the output data of a forward quadratic transformation is arranged in a two-dimensional block, as an example.
[0031] [Figure 8] This diagram shows the block shape to which LFNST is applied.
[0032] [Figure 9] This diagram shows an example of the arrangement of output data in a forward LFNST.
[0033] [Figure 10] This figure illustrates an example where the number of output data points for a forward LFNST was limited to a maximum of 16.
[0034] [Figure 11] This diagram shows a zero-out in a block to which a 4x4 LFNST is applied, as an example.
[0035] [Figure 12] This diagram shows a zero-out in a block to which an 8x8 LFNST is applied, as an example.
[0036] [Figure 13] Another example shows a zero-out in a block where 8x8LFNST is applied.
[0037] [Figure 14] This diagram shows an example of subblocks into which a single coding block is divided.
[0038] [Figure 15] This figure shows another example of subblocks, where a single coding block is divided.
[0039] [Figure 16] This diagram illustrates the symmetry between an M×2 (M×1) block and a 2×M (1×M) block as an example.
[0040] [Figure 17] This figure shows an example of a 2×M block being transposed.
[0041] [Figure 18] This shows an example of the scanning sequence for an 8x2 or 2x8 area.
[0042] [Figure 19] This is a diagram illustrating a video decoding method for an example.
[0043] [Figure 20] This is a diagram illustrating an example of a video encoding method.
[0044] [Figure 21] This document outlines some examples of video / image coding systems to which it can be applied.
[0045] [Figure 22] An illustrative diagram of a content streaming system structure to which this document applies is shown below. [Modes for carrying out the invention]
[0046] This document may be modified in various ways and may have various embodiments, but it attempts to illustrate and describe in detail specific embodiments with the drawings. However, this does not mean that this document is intended to limit itself to any particular embodiment. The terms used herein are used solely to describe specific embodiments and are not intended to limit the technical ideas contained herein. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "includes" or "has" are intended to indicate the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood not to preemptively exclude the possibility of the existence or addition of one or more different features, figures, steps, actions, components, parts, or combinations thereof.
[0047] On the other hand, each configuration shown in the diagrams described in this document is shown independently for the convenience of explaining its distinct characteristic functions, and does not mean that each configuration is implemented with separate hardware or separate software. For example, two or more configurations may be combined to form one configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the rights of this document, as long as they do not deviate from the essence of this document.
[0048] The following describes preferred embodiments of this document in more detail with reference to the attached figures. The same reference numerals are used for the same components in the drawings, and redundant descriptions of the same components are omitted.
[0049] This document relates to video / image coding. For example, the methods / examples disclosed in this document may be related to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), next-generation video / image coding standards after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), EVC (essential video coding) standard, AVS2 standard, etc.).
[0050] This document presents various embodiments relating to video / image coding, and unless otherwise noted, these embodiments may be performed in combination with each other.
[0051] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a single image representing a specific time period, while "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more CTUs (coding tree units). A single picture can consist of one or more slices or tiles. A single picture can consist of one or more tile groups. A tile group can contain one or more tiles.
[0052] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" can be used as a counterpart to pixel. A sample generally refers to a pixel or a pixel value, sometimes only the luma component pixel / pixel value, or sometimes only the chroma component pixel / pixel value. Alternatively, a sample may refer to a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it may refer to the conversion coefficient in the frequency domain.
[0053] A unit can represent a basic unit of image processing. A unit can include at least one specific region of a picture and information about that region. A unit can include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area, as appropriate. In general, an M×N block can include a sample (or sample array) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0054] In this document, the terms " / " and "," should be interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Furthermore, "A / B / C" means "at least one of A, B, and / or C." Similarly, "A, B, C" also means "at least one of A, B, and / or C."
[0055] Furthermore, in this document, "or" is interpreted as "and / or." For example, "A or B" may mean 1) only "A," 2) only "B," or 3) both "A and B." In other words, "or" in this document may mean "additionally or alternatively."
[0056] In this specification, "at least one of A and B" may mean "just A," "just B," or "both A and B." Furthermore, in this specification, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."
[0057] Furthermore, in this specification, "at least one of A, B and C" may mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."
[0058] Furthermore, parentheses used herein may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" as used herein is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."
[0059] Technical features described individually in each drawing in this specification may be implemented individually or simultaneously.
[0060] Figure 1 is a schematic diagram illustrating the configuration of a video / image encoding device to which this document applies. Hereinafter, the term "video encoding device" may include an image encoding device.
[0061] Referring to Figure 1, the encoding device 100 may be configured to include an image partitioner 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 may include an inter-predictor 121 and an intra-predictor 122. The residual processor 130 may include a transformer 132, a quantizer 133, a dequantizer 134, and an inverse transformer 135. The residual processor 130 may further include a subtractor 131. The adder 150 may be called a reconstructor or a reconstructed block generator. The aforementioned image segmentation unit 110, prediction unit 120, residual processing unit 130, entropy encoding unit 140, addition unit 150, and filtering unit 160 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 170 as an internal / external component.
[0062] The image splitting unit 110 can split an input image (or picture, frame) input to the encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary structure. Alternatively, the binary-tree structure may be applied first. Based on the final coding unit that cannot be further split, the coding procedure described in this document may be performed. In this case, based on coding efficiency due to image characteristics, the largest coding unit can be immediately used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0063] The term "unit" can be used interchangeably with terms such as "block" or "area," depending on the context. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.
[0064] The encoding device 100 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the inter-prediction unit 121 or intra-prediction unit 122 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 132. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoding device 100 can be called the subtraction unit 131. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes the predicted sample for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit 140, as will be described later in the explanation of each prediction mode. Information regarding the prediction can be encoded by the entropy encoding unit 140 and output in bitstream format.
[0065] The intra-prediction unit 122 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) of the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 122 can also determine the prediction mode to apply to the current block using the prediction modes applied to the surrounding blocks.
[0066] The interprediction unit 121 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between the surrounding block and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding block may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporal neighboring block may also be called a collocated picture (colPic). For example, the interpretation unit 121 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 121 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0067] The prediction unit 120 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values within the picture can be signaled based on information about the palette table and palette index.
[0068] The prediction signal generated via the prediction unit (including the inter-prediction unit 121 and / or the intra-prediction unit 122) can be used to generate a reconstructed signal or a residual signal. The transformation unit 132 can generate transformation coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and obtaining a transformation based on it. The transformation process can be applied to pixel blocks of the same size that are square, or to blocks of variable size that are not square.
[0069] The quantization unit 133 quantizes the conversion coefficients and transmits them to the entropy encoding unit 140, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 133 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 140 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 140 can also encode information necessary for video / image restoration (e.g., the values of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of network abstraction layer (NAL). The video / image information may further include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device in this document may be included in the video / image information. The video / image information may be encoded via the encoding procedure described above and included in the bitstream.The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 140 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 100, or the transmitting unit may be included in the entropy encoding unit 140.
[0070] The quantized conversion coefficients output from the quantization unit 133 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 134 and the inverse transformation unit 135. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 121 or the intra-prediction unit 122. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 150 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.
[0071] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.
[0072] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. As will be described later in the explanation of each filtering method, the filtering unit 160 can generate various filtering-related information and transmit it to the entropy encoding unit 140. The filtering-related information can be encoded by the entropy encoding unit 140 and output in the form of a bitstream.
[0073] The corrected restored picture sent to memory 170 can be used as a reference picture in the interpretation unit 121. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.
[0074] The DPB in memory 170 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 121. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the already restored picture. The stored motion information can be transmitted to the inter-prediction unit 121 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 122.
[0075] Figure 2 is a schematic diagram illustrating the configuration of a video / image decoding device to which this document applies.
[0076] Referring to Figure 2, the decoding device 200 can be configured to include an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The predictor 230 may include an inter-predictor 231 and an intra-predictor 232. The residual processor 220 may include a dequantizer 221 and an inverse transformer 222. The aforementioned entropy decoder 210, residual processor 220, predictor 230, adder 240, and filtering device 250 can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 260 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware component may further include memory 260 as an internal / external component.
[0077] When a bitstream containing video / image information is input, the decoding device 200 can reconstruct the image corresponding to the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 200 can derive units / blocks based on block division information obtained from the bitstream. The decoding device 200 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing units may be, for example, coding units, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed image signal decoded and output via the decoding device 200 can then be reproduced via a playback device.
[0078] The decoding device 200 can receive the signal output from the encoding device shown in Figure 1 in bitstream form, and the received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image reconstruction, quantized values of conversion coefficients related to residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the information of the syntax element to be decoded, the decoded information of the surrounding and decoded blocks, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction unit (inter-prediction unit 232 and intra-prediction unit 231), and the residual values that have been entropy decoded by the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 220. The residual processing unit 220 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 210, information related to filtering can be provided to the filtering unit 250. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 200, or the receiving unit can be a component of the entropy decoding unit 210. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 210, and the sample decoder may include at least one of the inverse quantization unit 221, inverse transformation unit 222, addition unit 240, filtering unit 250, memory 260, inter-prediction unit 232, and intra-prediction unit 231.
[0079] The inverse quantization unit 221 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 221 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 221 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) to obtain the transformation coefficients.
[0080] In the inverse conversion unit 222, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0081] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.
[0082] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0083] The intra-prediction unit 231 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) of the current block or at a distance from it, depending on the prediction mode. The prediction mode in intra-prediction can include a plurality of non-directional modes and a plurality of directional modes. The intra-prediction unit 231 can also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0084] The interprediction unit 232 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. To reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In interprediction, neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the interprediction unit 232 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.
[0085] The summing unit 240 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 232 and / or intra-prediction unit 231). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.
[0086] The summing unit 240 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and may be output after filtering, as described later, or may be used for intra-prediction of the next picture.
[0087] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0088] The filtering unit 250 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 250 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 260, specifically to the DPB of the memory 260. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0089] The (modified) restored picture stored in the DPB of memory 260 can be used as a reference picture by the inter-prediction unit 232. Memory 260 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 232 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 260 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 231.
[0090] In this document, the embodiments described for the filtering unit 160, inter-prediction unit 121, and intra-prediction unit 122 of the encoding device 100 can also be applied identically or in a corresponding manner to the filtering unit 250, inter-prediction unit 232, and intra-prediction unit 231 of the decoding device 200, respectively.
[0091] As described above, prediction is performed during video coding to improve compression efficiency. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived identically by the encoding and decoding devices, and the encoding device can improve image coding efficiency by signaling the decoding device information about the residual between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can generate a restored block containing restored samples by combining the residual block and the predicted block, and can generate a restored picture containing the restored block.
[0092] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, perform a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, perform a quantization procedure on the transformation coefficients to derive quantized transformation coefficients, and signal the associated residual information (via a bitstream) to a decoding device. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can perform an inverse quantization / inverse transformation procedure based on the residual information to derive a residual sample (or residual block). The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can further derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.
[0093] Figure 3 schematically illustrates the multiplexing technique used in this document.
[0094] Referring to Figure 3, the conversion unit may correspond to the conversion unit in the encoding device shown in Figure 1, and the inverse conversion unit may correspond to the inverse conversion unit in the encoding device shown in Figure 1 or the inverse conversion unit in the decoding device shown in Figure 3.
[0095] The transformation unit can perform a primary transformation based on the residual sample (residual sample array) in the residual block to derive (primary) transformation coefficients (S310). Such a primary transformation may be referred to as a core transformation. Here, the primary transformation is obtained based on Multiple Transform Selection (MTS), and when a multiple transformation is applied as the primary transformation, it may be referred to as a multiple core transformation.
[0096] Multiple core transforms can represent a method of transformation that further uses a Discrete Cosine Transform (DCT) type 2, a Discrete Sine Transform (DST) type 7, a Discrete Cosine Transform (DST) type 8, and / or a DST type 1. That is, the multiple core transforms can represent a method of transformation that converts a spatial domain residual signal (or residual block) into frequency domain transformation coefficients (or first-order transformation coefficients) based on a plurality of transformation kernels selected from the DCT type 2, DST type 7, DCT type 8, and DST type 1. Here, the first-order transformation coefficients may be called provisional transformation coefficients from the perspective of the transformer.
[0097] In other words, when an existing transformation method is applied, a spatial-domain to frequency-domain transformation of a residual signal (or residual block) can be applied based on DCT type 2 to generate transformation coefficients. In contrast, when the multiple core transformation is applied, a spatial-domain to frequency-domain transformation of a residual signal (or residual block) can be applied based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate transformation coefficients (or first-order transformation coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc. may be called transformation types, transformation kernels, or transformation cores. Such DCT / DST transformation types can be defined based on basis functions.
[0098] When the multi-core transformation is performed, a vertical transformation kernel and a horizontal transformation kernel can be selected from the transformation kernels for the target block, and a vertical transformation can be performed on the target block based on the vertical transformation kernel, and a horizontal transformation can be performed on the target block based on the horizontal transformation kernel. Here, the horizontal transformation may represent a transformation to the horizontal component of the target block, and the vertical transformation may represent a transformation to the vertical component of the target block. The vertical transformation kernel / horizontal transformation kernel can be adaptively determined based on the prediction mode and / or transformation index of the target block (CU or subblock) including the residual block.
[0099] Furthermore, for example, when applying MTS to perform a linear transformation, a mapping relationship to the transformation kernel can be established by setting a specific basis function to a predetermined value and combining it with whether or not a particular basis function is applied when it is a vertical or horizontal transformation. For example, if the horizontal transformation kernel is represented by trTypeHor and the vertical transformation kernel is represented by trTypeVer, then a value of 0 for trTypeHor or trTypeVer can be set to DCT2, a value of 1 for trTypeHor or trTypeVer can be set to DST7, and a value of 2 for trTypeHor or trTypeVer can be set to DCT8.
[0100] In this case, MTS index information can be encoded and signaled to a decoder to indicate one of a number of conversion kernel sets. For example, an MTS index of 0 can indicate that the values of trTypeHor and trTypeVer are all 0; an MTS index of 1 can indicate that the values of trTypeHor and trTypeVer are all 1; an MTS index of 2 can indicate that the value of trTypeHor is 2 and the value of trTypeVer is 1; an MTS index of 3 can indicate that the value of trTypeHor is 1 and the value of trTypeVer is 2; and an MTS index of 4 can indicate that the values of trTypeHor and trTypeVer are all 2.
[0101] As an example, the conversion kernel sets based on MTS index information are shown in the table below.
[0102] [Table 1]
[0103] The conversion unit performs a quadratic transformation based on the (primary) transformation coefficients to derive modified (secondary) transformation coefficients (S320). The primary transformation is a transformation from the spatial domain to the frequency domain, and the secondary transformation means transforming into a more compressed representation by utilizing the correlations that exist between the (primary) transformation coefficients. The secondary transformation includes a non-separable transform. In this case, the secondary transformation may be called a non-separable secondary transform (NSST) or MDNSST (mode-dependent non-separable secondary transform). The non-separable secondary transform represents a transformation that generates modified transformation coefficients (or secondary transformation coefficients) for the residual signal by performing a quadratic transformation on the (primary) transformation coefficients derived by the primary transformation based on a non-separable transform matrix. Here, the transformation can be applied at once to the (primary) transformation coefficients based on the non-separable transform matrix without separating the vertical and horizontal transformations (or applying the horizontal and vertical transformations independently). In other words, the non-separable quadratic transformation is not applied separately to the (primary) transformation coefficients in the vertical and horizontal directions, but rather represents a transformation method that, for example, rearranges a two-dimensional signal (transformation coefficient) into a one-dimensional signal in a specific fixed direction (e.g., row-first or column-first direction), and then generates a modified transformation coefficient (or quadratic transformation coefficient) based on the non-separable transformation matrix. For example, row-first ordering means arranging the 1st row, 2nd row, ..., Nth row in a column for an M×N block, and column-first ordering means arranging the 1st column, 2nd column, ..., Mth column in a column for an M×N block. The non-separable quadratic transformation can be applied to the top-left region of a block composed of (primary) transformation coefficients (hereinafter referred to as a transformation coefficient block). For example, if both the width (W) and height (H) of the conversion coefficient block are 8 or greater, an 8x8 non-separable quadratic transformation can be applied to the upper left 8x8 region of the conversion coefficient block.Furthermore, if both the width (W) and height (H) of the conversion coefficient block are 4 or greater, but either the width (W) or height (H) of the conversion coefficient block is less than 8, the 4×4 unseparable quadratic transformation can be applied to the upper left min(8,W)×min(8,H) region of the conversion coefficient block. However, the embodiment is not limited to this, and for example, even if only the condition that both the width (W) or height (H) of the conversion coefficient block are 4 or greater is satisfied, the 4×4 unseparable quadratic transformation can also be applied to the upper left min(8,W)×min(8,H) region of the conversion coefficient block.
[0104] Specifically, for example, if a 4x4 input block is used, the unseparated quadratic transform can be performed as follows:
[0105] The aforementioned 4x4 input block X can be represented as follows:
[0106]
number
[0107] When X is shown in the form of a vector, JPEG2026131834000004.jpg12150 can be represented as follows:
[0108]
number
[0109] As shown in equation 2, vector JPEG2026131834000006.jpg15147 rearranges the 2D block of X in equation 1 into a 1D vector in row-first order.
[0110] In this case, the quadratic inseparable transform can be calculated as follows:
[0111]
number
[0112] Here, JPEG2026131834000008.jpg13139 shows the transformation coefficient vector, and T shows a 16x16 (non-separable) transformation matrix.
[0113] Through the aforementioned equation 3, a 16 × 1 transformation coefficient vector JPEG2026131834000009.jpg12155 can be derived, and the above JPEG2026131834000010.jpg12156 can be reorganized into 4x4 blocks via a scan order (horizontal, vertical, diagonal, etc.). However, the calculation described above is illustrative, and to reduce the computational complexity of the inseparable quadratic transform, HyGT (Hypercube-Givens Transform), etc., can also be used for the calculation of the inseparable quadratic transform.
[0114] On the other hand, the unseparated quadratic transform allows the transform kernel (or transform core, transform type) to be selected as mode-dependent. Here, the mode may include intra-predictive mode and / or inter-predictive mode.
[0115] As described above, the non-separable quadratic transformation can be performed based on an 8x8 transformation or a 4x4 transformation determined based on the width (W) and height (H) of the transformation coefficient block. An 8x8 transformation refers to a transformation that can be applied to an 8x8 region contained within the transformation coefficient block when W and H are all equal to or greater than 8, and this 8x8 region may be the upper left 8x8 region inside the transformation coefficient block. Similarly, a 4x4 transformation refers to a transformation that can be applied to a 4x4 region contained within the transformation coefficient block when W and H are all equal to or greater than 4, and this 4x4 region may be the upper left 4x4 region inside the transformation coefficient block. For example, an 8x8 transformation kernel matrix may be a 64x64 / 16x64 matrix, and a 4x4 transformation kernel matrix may be a 16x16 / 8x16 matrix.
[0116] In this case, for the selection of mode-based conversion kernels, two non-separable quadratic conversion kernels may be configured for each conversion set for non-separable quadratic conversions for both 8x8 and 4x4 conversions, and there may be four conversion sets. That is, four conversion sets may be configured for 8x8 conversions and four conversion sets may be configured for 4x4 conversions. In this case, the four conversion sets for 8x8 conversions may each contain two 8x8 conversion kernels, and in this case, the four conversion sets for 4x4 conversions may each contain two 4x4 conversion kernels.
[0117] However, the size of the transformation, i.e., the size of the region to which the transformation is applied, may be a size other than 8x8 or 4x4 as an example, the number of sets may be n, and the number of transformation kernels in each set may be k.
[0118] The aforementioned transformation set may be called an NSST set or an LFNST set. The selection of a particular set from among the transformation sets can be performed, for example, based on the intra-prediction mode of the current block (CU or sub-block). LFNST (Low-Frequency Non-Separable Transform) may be an example of a reduced non-separable transform described later, and represents a non-separable transform for low-frequency components.
[0119] For reference, for example, an intra-prediction mode may include two non-directinoal (or non-angular) intra-prediction modes and 65 directional (or angular) intra-prediction modes. The non-directinoal intra-prediction mode may include a planar intra-prediction mode (number 0) and a DC intra-prediction mode (number 1), and the directional intra-prediction mode may include 65 intra-prediction modes (numbers 2 through 66). However, this is illustrative, and this document can also be applied when the number of intra-prediction modes differs. On the other hand, a 67th intra-prediction mode may be used as needed, and this 67th intra-prediction mode may represent a linear model (LM) mode.
[0120] Figure 4 illustrates 65 intradirectional modes for predicting directions.
[0121] Referring to Figure 4, we can distinguish between intra-prediction modes with horizontal directionality and intra-prediction modes with vertical directionality, centered around intra-prediction mode 34, which has a prediction direction on the lower right diagonal. In Figure 4, H and V represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate a displacement of 1 / 32 units on the sample grid position. This can be used to indicate an offset relative to the mode index value. Intra-prediction modes 2 through 33 are horizontally oriented, while intra-prediction modes 34 through 66 are vertically oriented. On the other hand, intra-prediction mode 34 can be seen as neither strictly horizontal nor vertical, but from the perspective of determining the transformation set of the quadratic transformation, it can be classified as belonging to the horizontal directionality. This is because the input data is transposed for the vertical modes which are symmetrical with respect to intra-prediction mode 34, and the input data alignment method for the horizontal modes is used for intra-prediction mode 34. Transposing input data means that for a 2D block of data MxN, rows become columns and columns become rows, resulting in NxM data. Intra-prediction modes 18 and 50 represent the horizontal intra-prediction mode and the vertical intra-prediction mode, respectively. Intra-prediction mode 2 predicts the upper right direction using the left reference pixel, so it can be called the upper right diagonal intra-prediction mode. In the same context, intra-prediction mode 34 can be called the lower right diagonal intra-prediction mode, and intra-prediction mode 66 can be called the lower left diagonal intra-prediction mode.
[0122] For example, the mapping of four transformation sets in intra-predictive mode may be shown, for instance, as in the following table.
[0123] [Table 2]
[0124] As shown in Table 2, the intra prediction mode allows mapping to one of four transformation sets, i.e., lfnstTrSetIdx to any of the four values from 0 to 3.
[0125] On the other hand, once it is determined that a specific set is to be used for an inseparable transformation, one of the k transformation kernels within that specific set can be selected via the inseparable quadratic transformation index. The encoding device can derive an inseparable quadratic transformation index that points to a specific transformation kernel based on an RD (rate-distortion) check, and can signal the decoding device to the inseparable quadratic transformation index. The decoding device can select one of the k transformation kernels within the specific set based on the inseparable quadratic transformation index. For example, index value 0 of lfnst can point to the first inseparable quadratic transformation kernel, index value 1 of lfnst can point to the second inseparable quadratic transformation kernel, index value 2 of lfnst can point to the third inseparable quadratic transformation kernel. Alternatively, index value 0 of lfnst can indicate that the first inseparable quadratic transformation is not applied to the target block, and index values 1 to 3 of lfnst can point to the three transformation kernels.
[0126] The transformation unit can perform the non-separable quadratic transformation based on the selected transformation kernel to obtain the modified (quadratic) transformation coefficients. The modified transformation coefficients can be derived from the quantized transformation coefficients via the quantization unit as described above, encoded, and transmitted to the decoder for signaling and to the inverse quantization / inverse transformation unit within the encoding unit.
[0127] On the other hand, as mentioned above, if the quadratic transformation is omitted, the (primary) transformation coefficients, which are the output of the primary (separated) transformation, can be derived as quantized transformation coefficients via the quantization unit as described above, encoded, and transmitted to the decoder for signaling and to the inverse quantization / inverse transformation unit within the encoding unit.
[0128] The inverse transform unit can perform a series of steps in the reverse order of the steps performed by the transform unit described above. The inverse transform unit can receive the (inversely quantized) transform coefficients, perform a quadratic (inverse) transform to derive the (primary) transform coefficients (S350), and perform a primary (inverse) transform on the (primary) transform coefficients to obtain a residual block (residual sample) (S360). Here, the primary transform coefficients may be called modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding and decoding devices can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.
[0129] On the other hand, the decoding device may further include a quadratic inverse transform applicability determination unit (or an element that determines whether a quadratic inverse transform is applicable) and a quadratic inverse transform determination unit (or an element that determines a quadratic inverse transform). The quadratic inverse transform applicability determination unit can determine whether a quadratic inverse transform is applicable. For example, the quadratic inverse transform may be NSST, RST, or LFNST, and the quadratic inverse transform applicability determination unit can determine whether a quadratic inverse transform is applicable based on a quadratic transform flag parsed from the bitstream. As another example, the quadratic inverse transform applicability determination unit can also determine whether a quadratic inverse transform is applicable based on the transformation coefficients of the residual block.
[0130] The quadratic inverse transform determination unit can determine the quadratic inverse transform. At that time, the quadratic inverse transform determination unit can determine the quadratic inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified by the intra-prediction mode. Furthermore, as one embodiment, the quadratic transform determination method can be determined dependent on the linear transform determination method. Various combinations of linear and quadratic transforms can be determined by the intra-prediction mode. Also, as an example, the quadratic inverse transform determination unit can determine the region to which the quadratic inverse transform is applied based on the size of the current block.
[0131] On the other hand, as mentioned above, if the quadratic (inverse) transformation is omitted, the (inversely quantized) transformation coefficients can be received, and the primary (separated) inverse transformation can be performed to obtain a residual block (residual sample). As mentioned above, the encoding and decoding devices can generate a reconstructed block based on the residual block and the predicted block, and generate a reconstructed picture based on this.
[0132] On the other hand, in this document, in order to reduce the computational complexity and memory requirements associated with unseparable quadratic transforms, we can apply RST (reduced secondary transform), which is a reduced version of the NSST concept with a smaller transform matrix (kernel).
[0133] On the other hand, the transformation kernel, transformation matrix, and coefficients constituting the transformation kernel matrix described in this document—that is, kernel coefficients or matrix coefficients—can be represented in 8 bits. This may be a requirement for implementation in decoding and encoding devices, and it can reduce the memory requirements for storing the transformation kernel while resulting in a reasonably acceptable performance degradation compared to existing 9-bit or 10-bit representations. Furthermore, representing the kernel matrix in 8 bits allows for the use of smaller multipliers and may be more suitable for SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.
[0134] In this specification, RST can mean a transformation performed on a residual sample for a target block based on a transform matrix whose size has been reduced by a simplification factor. When a simplification transformation is performed, the amount of computation required during the transformation can be reduced due to the reduction in the size of the transform matrix. In other words, RST can be used to resolve the complexity issue that arises when transforming large blocks or during non-separable transformations.
[0135] RST can be referred to by a variety of terms, including reduced transform, reduced secondary transform, reduction transform, simplified transform, and simple transform, and the names to which RST can be referred are not limited to those given. Alternatively, since RST mainly occurs in the low-frequency region with non-zero coefficients in the transform block, it may also be referred to as LFNST (Low-Frequency Non-Separable Transform). The aforementioned transform index may be named the LFNST index.
[0136] On the other hand, when the quadratic inverse transform is performed based on RST, the inverse transform unit 135 of the encoding device 100 and the inverse transform unit 222 of the decoding device 200 may each include an inverse RST unit that derives corrected transform coefficients based on the inverse RST of the transform coefficients, and an inverse linear transform unit that derives the residual sample for the target block based on the inverse linear transform of the corrected transform coefficients. The inverse linear transform means the inverse transform of the linear transform that was applied to the residual. In this document, deriving transform coefficients based on a transform means deriving transform coefficients by applying the transform.
[0137] Figure 5 is a diagram illustrating an RST according to one embodiment of this document.
[0138] In this specification, “target block” may mean the current block, residual block, or transformed block on which coding is performed.
[0139] In one embodiment of the RST, an N-dimensional vector is mapped to an R-dimensional vector located in a different space, and a reduced transformation matrix can be determined, where R is less than N. N can represent the square of the side length of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can represent the R / N value. The simplification factor can be referred to by various terms such as reduced factor, reduction factor, reduced factor, simplified factor, and simple factor. On the other hand, R can be referred to as the reduced coefficient, but in some cases the simplification factor may mean R. Also, in some cases the simplification factor may mean the N / R value.
[0140] In one embodiment, the simplification factor or simplification coefficient can be signaled via a bitstream, but the embodiment is not limited to this. For example, predefined values for the simplification factor or simplification coefficient may be stored in each encoding device 100 and decoding device 200, in which case the simplification factor or simplification coefficient may not be signaled separately.
[0141] The size of the simplified transformation matrix in one embodiment is RxN, which is smaller than the size NxN of a normal transformation matrix, and can be defined as shown in Equation 4 below.
[0142]
number
[0143] The matrix T in the Reduced Transform block shown in Figure 5(a) can be interpreted as the matrix TRxN in Equation 4. As shown in Figure 5(a), when the simplified transformation matrix TRxN is multiplied by the residual sample for the target block, the transformation coefficients for the target block can be derived.
[0144] In one embodiment, if the size of the block to which the transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), then the RST shown in Figure 5(a) can be expressed by a matrix operation as shown in Equation 5 below. In this case, memory and multiplication operations can be reduced to approximately 1 / 4 by the simplification factor.
[0145] In this text, matrix operations can be understood as operations that involve placing a matrix to the left of a column vector and multiplying the matrix by the column vector to obtain a new column vector.
[0146]
number
[0147] In Equation 5, r1 to r64 can represent residual samples for the target block, and more specifically, they can be transformation coefficients generated by applying a linear transformation. From the calculation result of Equation 5, the transformation coefficient ci for the target block can be derived, and the derivation process of ci is as shown in Equation 6.
[0148]
number
[0149] The result of the calculation in equation 6 allows us to derive the conversion coefficients c1 to cR for the target block. That is, if R=16, we can derive the conversion coefficients c1 to c16 for the target block. If a regular conversion were applied instead of RST, and a conversion matrix of size 64x64 (NxN) was multiplied by a residual sample of size 64x1 (Nx1), 64 (N) conversion coefficients for the target block might be derived. However, because RST was applied, only 16 (R) conversion coefficients for the target block are derived. The total number of conversion coefficients for the target block decreases from N to R, and the amount of data that the encoding device 100 sends to the decoding device 200 decreases, so the transmission efficiency between the encoding device 100 and the decoding device 200 may increase.
[0150] From the perspective of the size of the transformation matrix, the size of a normal transformation matrix is 64x64 (NxN), but the size of a simplified transformation matrix is reduced to 16x64 (RxN). Therefore, compared to performing a normal transformation, the memory usage when executing RST can be reduced by a ratio of R / N. In addition, compared to the number of multiplication operations NxN when using a normal transformation matrix, the number of multiplication operations can be reduced by a ratio of R / N (RxN) when using a simplified transformation matrix.
[0151] In one embodiment, the conversion unit 132 of the encoding device 100 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the target block's residual sample. These conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 200, and the inverse conversion unit 222 of the decoding device 200 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) on the conversion coefficients, and derive residual samples for the target block based on an inverse primary conversion on the modified conversion coefficients.
[0152] The size of the inverse RST matrix TNxR in one embodiment is NxR, which is smaller than the size NxN of a normal inverse transform matrix, and it is in a transpose relationship with the simplified transform matrix TRxN shown in Equation 4.
[0153] The matrix Tt in the Reduced Inv. Transform block shown in Figure 5(b) can represent the inverse RST matrix TRxNT (the superscript T indicates transpose). As shown in Figure 5(b), when the inverse RST matrix TRxNT is multiplied by the transformation coefficients for the target block, the modified transformation coefficients or residual samples for the target block can be derived. The inverse RST matrix TRxNT can also be expressed as (TRxN)TNxR.
[0154] More specifically, when the inverse RST is applied to a quadratic inverse transform, multiplying the transformation coefficients for the target block by the inverse RST matrix TRxNT yields the modified transformation coefficients for the target block. On the other hand, when the inverse RST is applied to a linear inverse transform, multiplying the transformation coefficients for the target block by the inverse RST matrix TRxNT yields the residual sample for the target block.
[0155] In one embodiment, when the size of the block to which the inverse transform is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the RST shown in Figure 5(b) can be expressed by a matrix operation as shown in Equation 7 below.
[0156]
number
[0157] In equation 7, c1 to c16 can represent the conversion coefficients for the target block. From the calculation result of equation 7, rj, which represents the modified conversion coefficients for the target block or the residual sample for the target block, can be derived, and the derivation process of rj is as shown in equation 8.
[0158]
number
[0159] The result of the calculation in Equation 8 allows us to derive r1 to rN, which represent the modified transformation coefficients for the target block or the residual samples for the target block. Considering this from the perspective of the size of the inverse transformation matrix, the size of a normal inverse transformation matrix is 64x64 (NxN), but the size of the simplified inverse transformation matrix is reduced to 64x16 (NxR), so compared to performing a normal inverse transformation, the memory usage when performing the inverse RST can be reduced by a ratio of R / N. Also, compared to the number of multiplication operations NxN when using a normal inverse transformation matrix, the number of multiplication operations can be reduced by a ratio of R / N (NxR) when using a simplified inverse transformation matrix.
[0160] On the other hand, the transformation set configuration shown in Table 2 can also be applied to an 8x8 RST. That is, the 8x8 RST can be applied using the transformation sets in Table 2. Since one transformation set consists of two or three transformations (kernels) depending on the prediction mode on the screen, it can be configured to select one of up to four transformations, including the case where a quadratic transformation is not applied. The transformation when a quadratic transformation is not applied can be considered as the application of the identity matrix. If we assign indices 0, 1, 2, and 3 to the four transformations respectively (for example, index 0 can be assigned to the identity matrix, i.e., when a quadratic transformation is not applied), then the transformation to be applied can be specified by signaling a transformation index or lfnst index, which is a syntax element, for each block of transformation coefficients. That is, via the transformation index, an 8x8 RST can be specified for the upper left block of the 8x8 in the RST configuration, or an 8x8 lfnst can be specified if LFNST is applied. 8x8 lfnst and 8x8 RST refer to transformations that can be applied to an 8x8 region contained within the block of the transformation coefficient when the W and H of the target block to be transformed are all equal to or greater than 8, and this 8x8 region may be the upper left 8x8 region inside the block of the transformation coefficient. Similarly, 4x4 lfnst and 4x4 RST refer to transformations that can be applied to a 4x4 region contained within the block of the transformation coefficient when the W and H of the target block are all equal to or greater than 4, and this 4x4 region may be the upper left 4x4 region inside the block of the transformation coefficient.
[0161] On the other hand, according to one embodiment of this document, in the encoding process, instead of a 16x64 transformation kernel matrix, it is possible to select only 48 data points from the 64 data points constituting an 8x8 region and apply a transformation kernel matrix of up to 16x48. Here, "up to" means that for an mx48 transformation kernel matrix that can generate m coefficients, the maximum value of m is 16. That is, when RST is executed by applying an mx48 transformation kernel matrix (m ≤ 16) to an 8x8 region, it is possible to receive 48 data inputs and generate m coefficients. When m is 16, it receives 48 data inputs and generates 16 coefficients. That is, if the 48 data points form a 48x1 vector, a 16x48 matrix and a 48x1 vector can be multiplied in order to generate a 16x1 vector. At that time, the 48 data points constituting the 8x8 region can be appropriately arranged to construct a 48x1 vector. At that time, applying a transformation kernel matrix of up to 16x48 and performing matrix operations will generate 16 modified transformation coefficients, which can be placed in the upper left 4x4 region according to the scanning order, while the upper right 4x4 region and the lower left 4x4 region may be filled with 0.
[0162] For the inverse transformation of the decoding process, the transposed matrix of the transformation kernel matrix described above can be used. That is, when inverse RST or LFNST is performed in the inverse transformation process executed by the decoding device, the input coefficient data to which the inverse RST is applied consists of one-dimensional vectors according to a predetermined arrangement order, and the modified coefficient vectors obtained by multiplying the one-dimensional vectors by the matrix of the inverse RST on the left side can be arranged in two-dimensional blocks according to a predetermined arrangement order.
[0163] To summarize, when RST or LFNST is applied to an 8x8 region during the transformation process, a matrix operation is performed between the 48 transformation coefficients in the upper left, upper right, and lower left regions of the 8x8 region (excluding the lower right region) and a 16x48 transformation kernel matrix. For this matrix operation, the 48 transformation coefficients are input into a one-dimensional array. After this matrix operation, 16 modified transformation coefficients are derived, and these modified transformation coefficients can be arranged in the upper left region of the 8x8 region.
[0164] Conversely, when inverse RST or LFNST is applied to an 8x8 region during the inverse transformation process, the 16 transformation coefficients corresponding to the upper left side of the 8x8 region are input in a one-dimensional array form according to the scanning order and can be used in matrix operations with a 48x16 transformation kernel matrix. That is, the matrix operation in such a case can be expressed as (48x16 matrix) * (16x1 transformation coefficient vector) = (48x1 modified transformation coefficient vector). Here, the nx1 vector can be interpreted as an nx1 matrix, so it may also be expressed as an nx1 column vector. Also, * means matrix multiplication. When such a matrix operation is performed, 48 modified transformation coefficients can be derived, and these 48 modified transformation coefficients can be arranged in the upper left, upper right, and lower left regions of the 8x8 region, excluding the lower right region.
[0165] On the other hand, when the quadratic inverse transform is performed based on RST, the inverse transform unit 135 of the encoding device 100 and the inverse transform unit 222 of the decoding device 200 may each include an inverse RST unit that derives corrected transform coefficients based on the inverse RST of the transform coefficients, and an inverse linear transform unit that derives the residual sample for the target block based on the inverse linear transform of the corrected transform coefficients. The inverse linear transform means the inverse transform of the linear transform that was applied to the residual. In this document, deriving transform coefficients based on a transform means deriving transform coefficients by applying the transform.
[0166] Looking specifically at the previously mentioned non-separated transform, LFNST, it is as follows: LFNST can include both a forward transform by the encoding device and an inverse transform by the decoding device.
[0167] The encoding device applies a primary (core) transform, and then applies a secondary transform to the derived result (or part of the result) as input.
[0168]
number
[0169] In equation 9 above, x and y are the input and output of the quadratic transformation, respectively, G is the matrix representing the quadratic transformation, and the transform basis vectors are composed of column vectors. In the case of inverse LFNST, when the dimension of the transformation matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the dimension of GT is obtained by transposing matrix G.
[0170] In the case of inverse LFNST, the dimensions of matrix G are [48x16], [48x8], [16x16], and [16x8], where the [48x8] matrix and the [16x8] matrix are submatrices obtained by sampling eight transformation basis vectors from the left side of the [48x16] matrix and the [16x16] matrix, respectively.
[0171] On the other hand, in the case of forward LFNST, the dimensions of the matrix GT are [16x48], [8x48], [16x16], and [8x16], and the [8x48] matrix and the [8x16] matrix are submatrices obtained by sampling eight transformation basis vectors from above the [16x48] matrix and the [16x16] matrix, respectively.
[0172] Therefore, in the case of forward LFNST, the input x can be a [48x1] vector or a [16x1] vector, and the output y can be a [16x1] vector or an [8x1] vector. Since the output of the forward linear transformation in video coding and decoding is 2D data, in order to construct a [48x1] vector or a [16x1] vector as input x, the 2D data that is the output of the forward transformation must be appropriately arranged to construct a 1D vector.
[0173] Figure 6 shows an example of the sequence for arranging the output data of a forward linear transformation into a one-dimensional vector. The left-hand diagrams of Figure 6(a) and (b) show the sequence for creating a [48x1] vector, and the right-hand diagrams of Figure 6(a) and (b) show the sequence for creating a [16x1] vector. In the case of LFNST, the 2D data is sequentially arranged in the order shown in Figure 6(a) and (b) to obtain a one-dimensional vector x.
[0174] The orientation of the output data of such a forward linear transformation can be determined by the intra-prediction mode of the current block. For example, if the intra-prediction mode of the current block is horizontal with respect to the diagonal direction, the output data of the forward linear transformation can be arranged in the order shown in Figure 6(a), and if the intra-prediction mode of the current block is vertical with respect to the diagonal direction, the output data of the forward linear transformation can be arranged in the order shown in Figure 6(b).
[0175] For example, a different ordering can be applied than the orderings in Figures 6(a) and (b). To derive the same result (y vector) as when applying the orderings in Figures 6(a) and (b), the column vectors of matrix G can be rearranged to match the chosen ordering. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0176] Since the output y derived via equation 9 is a one-dimensional vector, if a configuration that processes the result of a forward quadratic transformation as input, such as a configuration that performs quantization or residual coding, requires two-dimensional data as input data, then the output y vector from equation 9 must again be appropriately positioned in 2D data.
[0177] Figure 7 shows an example of the order in which the output data of a forward quadratic transformation is arranged in a two-dimensional block.
[0178] In the case of LFNST, the output values can be placed in a 2D block according to a predetermined scan order. Figure 7(a) shows that when the output y is a [16x1] vector, the output values are placed in 16 positions in the 2D block according to the diagonal scan order. Figure 7(b) shows that when the output y is an [8x1] vector, the output values are placed in 8 positions in the 2D block according to the diagonal scan order, and the remaining 8 positions are filled with 0. In Figure 7(b), X is shown to be filled with 0.
[0179] In another example, depending on the configuration performing quantization or residual coding, the order in which the output vector y is processed can be according to a pre-set order, so the output vector y may not be placed in a 2D block, as shown in Figure 7. However, in the case of residual coding, data coding can be performed in units of 2D blocks (e.g., 4x4) such as CG (Coefficient Group), in which case the data can be arranged according to a specific order, as in the diagonal scan order of Figure 7.
[0180] On the other hand, the decoding device can construct a one-dimensional input vector y by arranging the two-dimensional data output through an inverse quantization process, etc., according to a pre-set scan order for the reverse transformation. The input vector y can be output as the input vector x by the following formula.
[0181]
number
[0182] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16x1] vector or an [8x1] vector, by the G matrix. In the case of inverse LFNST, the output vector x can be a [48x1] vector or a [16x1] vector.
[0183] The output vector x is arranged in a 2D block according to the order shown in Figure 6, and this 2D data becomes the input data (or part of the input data) for the inverse linear transformation.
[0184] Therefore, the inverse quadratic transformation is generally the opposite of the forward quadratic transformation process. In the case of the inverse transformation, unlike the forward transformation, the inverse quadratic transformation is applied first, followed by the inverse linear transformation.
[0185] In reverse LFNST, one of eight [48x16] matrices or eight [16x16] matrices can be selected as the transformation matrix G. Which of the [48x16] and [16x16] matrices to apply is determined by the size and shape of the block.
[0186] Furthermore, the eight matrices can be derived from four transformation sets, as shown in Table 2 above, and each transformation set can consist of two matrices. Which of the four transformation sets to use is determined by the intra-prediction mode, and more specifically, the transformation set is determined based on the extended intra-prediction mode value, taking into account the wide-angle intra-prediction mode (WAIP). Which of the two matrices that make up the selected transformation set is chosen is derived via index signaling. More specifically, the index value that is transmitted can be 0, 1, or 2, where 0 indicates that LFNST should not be applied, and 1 and 2 can indicate either of the two transformation matrices that make up the transformation set selected based on the intra-prediction mode value.
[0187] On the other hand, as mentioned above, whether or not to apply the transformation matrix (48x16 or 16x16) to LFNST depends on the size and shape of the block being transformed.
[0188] Figure 8 shows the shapes of blocks to which LFNST is applied. Figure 8(a) shows a 4x4 block, (b) shows 4x8 and 8x4 blocks, (c) shows a 4xN or Nx4 block where N is 16 or greater, (d) shows an 8x8 block, and (e) shows an MxN block where M≧8, N≧8, and N>8 or M>8.
[0189] In Figure 8, blocks with thick borders indicate the regions to which LFNST is applied. For blocks (a) and (b) in Figure 8, LFNST is applied to the top-left 4x4 region, and for block (c) in Figure 8, LFNST is applied to two consecutively placed top-left 4x4 regions. In Figures (a), (b), and (c), LFNST is applied in units of 4x4 regions, so such LFNST will be hereafter referred to as "4x4 LFNST," and the transformation matrix can be a [16x16] or [16x8] matrix with a matrix dimension of G in equations 9 and 10.
[0190] More specifically, a [16x8] matrix is applied to the 4x4 block (4x4TU or 4x4CU) in Figure 8(a), and a [16x16] matrix is applied to the blocks in Figures 8(b) and (c). This is to match the computational complexity for the worst case to 8 multiplications per sample.
[0191] In Figures 8(d) and (e), LFNST is applied to the upper left 8x8 region, and such LFNST will be hereafter referred to as "8x8 LFNST". The transformation matrix can be either a [48x16] or [48x8] matrix. In the case of forward LFNST, a [48x1] vector (the x vector in equation 9) is input as input data, so not all sample values in the upper left 8x8 region are used as input values for forward LFNST. That is, as seen in the left-to-right order of Figure 6(a) or Figure 6(b), the bottom-right 4x4 block is left as is, and the [48x1] vector can be constructed based on the samples belonging to the remaining three 4x4 blocks.
[0192] A [48x8] matrix can be applied to the 8x8 block (8x8TU or 8x8CU) in Figure 8(d), and a [48x16] matrix can be applied to the 8x8 block in Figure 8(e). This is also to match the computational complexity for the worst case to 8 multiplications per sample.
[0193] Depending on the shape of the block, applying the corresponding forward LFNST (4x4LFNST or 8x8LFNST) generates 8 or 16 output data (y vectors in Equation 9, [8x1] or [16x1] vectors). In forward LFNST, due to the properties of the matrix GT, the number of output data is equal to or less than the number of input data.
[0194] Figure 9 is a diagram illustrating the arrangement of output data for a forward LFNST as an example, showing blocks in which the output data for the forward LFNST is arranged according to the block shape.
[0195] In Figure 9, the shaded area in the upper left of the block corresponds to the region where the output data of the forward LFNST is located. Positions marked with 0 indicate samples filled with a value of 0, while the remaining area represents the region that is not modified by the forward LFNST. In the region not modified by the LFNST, the output data of the forward linear transformation remains unchanged.
[0196] As mentioned above, the dimension of the transformation matrix applied changes depending on the shape of the block, so the number of output data also changes. As shown in Figure 9, the output data of the forward LFNST may not fill the entire upper left 4x4 block. In the cases of Figure 11(a) and (d), the blocks or parts of the inside of the blocks shown by the thick lines are to which the [16x8] matrix and [48x8] matrix are applied, respectively, and an [8x1] vector is generated in the output of the forward LFNST. That is, according to the scan order shown in Figure 7(b), only 8 output data are filled as in Figure 9(a) and (d), and the remaining 8 positions can be filled with 0. In the case of the LFNST applied block in Figure 8(d), the two 4x4 blocks on the upper right and lower left sides adjacent to the upper left 4x4 block are also filled with 0 values as in Figure 9(d).
[0197] As described above, the basic approach is to signal the LFNST index and specify whether LFNST is applicable and which transformation matrix to apply. As shown in Figure 9, when LFNST is applied, the number of output data for forward LFNST may be equal to or less than the number of input data, resulting in a region filled with zero values as follows.
[0198] 1) As shown in Figure 9(a), within the 4x4 block on the upper left, the 8th and subsequent positions in the scan order, i.e., samples 9 through 16.
[0199] 2) As shown in Figures 9(d) and (e), a [16×48] matrix or an [8×48] matrix is applied to two 4×4 blocks adjacent to the upper left 4×4 block, or the second and third 4×4 blocks in the scan order.
[0200] Therefore, by checking the areas in 1) and 2) above, if non-zero data exists, it is certain that LFNST is not applied, and thus the signaling of the LFNST index can be omitted.
[0201] For example, in the case of LFNST adopted in the VVC standard, signaling of the LFNST index is performed after residual coding, so the encoding device can determine whether or not there is non-zero data (effectiveness factor) at all positions within the TU or CU block via residual coding. Therefore, the encoding device can determine whether or not to perform signaling for the LFNST index based on the presence or absence of non-zero data, and the decoding device can determine whether or not the LFNST index can be parsed. If there is no non-zero data in the area specified in 1) and 2) above, then signaling for the LFNST index will be performed.
[0202] Since truncated unary code is applied to the LFNST index using a binary evolution method, the LFNST index consists of a maximum of two bins, and the binary codes assigned to the possible LFNST index values 0, 1, and 2 are 0, 10, and 11, respectively. In the case of LFNST adopted in the current VVC, context-based CABAC coding (regular coding) is applied to the first bin, and bypass coding is applied to the second bin. The total number of contexts for the first bin is two, and (DCT-2, DCT-2) is applied as the primary transform pair for the horizontal and vertical directions. If the lumen and chroma components are coded in a dual-tree type, one context is assigned, and the other context is applied for the remaining cases. The coding of such an LFNST index is shown in the table below.
[0203] [Table 3]
[0204] On the other hand, the following simplification method can be applied to the adopted LFNST.
[0205] (i) For example, the number of output data for a forward LFNST can be limited to a maximum of 16.
[0206] In the case of Figure 8(c), a 4x4 LFNST can be applied to each of the two adjacent 4x4 regions in the upper left, generating a maximum of 32 LFNST output data. If the number of output data for forward LFNST is limited to a maximum of 16, then for a 4xN / Nx4 (N≧16) block (TU or CU), a 4x4 LFNST can be applied to only one 4x4 region in the upper left, allowing LFNST to be applied only once to all blocks in Figure 8. This simplifies the implementation for image coding.
[0207] Figure 10 shows an example where the number of output data for a forward LFNST is limited to a maximum of 16. As shown in Figure 10, when an LFNST is applied to the upper leftmost 4x4 region of a 4xN or Nx4 block where N is 16 or greater, the output data for the forward LFNST will be 16.
[0208] (ii) For example, zero-out can be applied to regions to which LFNST is not applied. In this document, zero-out can mean that the values of all positions belonging to a particular region are filled with 0. That is, zero-out can be applied to regions that maintain the result of the forward linear transformation without being changed by LFNST. As mentioned above, LFNST is divided into 4x4LFNST and 8x8LFNST, so zero-out can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.
[0209] (ii)-(A) When a 4x4 LFNST is applied, areas to which the 4x4 LFNST is not applied can be zeroed out. Figure 11 shows an example of zeroing out in a block to which a 4x4 LFNST is applied.
[0210] As shown in Figure 11, for a block to which a 4x4 LFNST is applied, that is, for blocks (a), (b), and (c) in Figure 9, the entire region up to the area to which LFNST is not applied can be filled with 0.
[0211] On the other hand, Figure 11(d) shows that when the maximum number of output data points for the forward LFNST is limited to 16, as in Figure 12, zero-out is performed on the remaining blocks to which the 4x4 LFNST is not applied.
[0212] (ii)-(B) When an 8x8 LFNST is applied, areas to which the 8x8 LFNST is not applied can be zeroed out. Figure 12 shows an example of zeroing out in a block to which an 8x8 LFNST is applied.
[0213] As shown in Figure 12, for a block to which an 8x8 LFNST is applied, that is, for blocks (d) and (e) in Figure 9, it is possible to fill all areas, including those to which LFNST is not applied, with zeros.
[0214] (iii) Due to the zero-out method presented in (ii) above, the area filled with zeros may change when LFNST is applied. Therefore, the zero-out method proposed in (ii) above allows checking for the presence of non-zero data over a wider area than in the case of LFNST in Figure 9.
[0215] For example, when applying (ii)-(B), in addition to the areas filled with zero values in (d) and (e) of Figure 9, it is checked whether there is any non-zero data in the areas further filled with zeros in Figure 12. Signaling for the LFNST index can only be performed if no non-zero data is found.
[0216] Of course, even if the zero-out proposed in (ii) above is applied, it is still possible to check whether or not non-zero data exists, just as with the signaling of existing LFNST indexes. That is, for blocks filled with zeros in Figure 9, it is possible to check whether or not non-zero data exists and apply the signaling of the LFNST index. In such a case, zero-out can be performed only on the encoding device, and the decoding device can not assume that zero-out exists, that is, it can only check whether or not non-zero data exists for areas explicitly indicated as zeros in Figure 9, and then perform parsing of the LFNST index.
[0217] Alternatively, a zero-out can be performed as shown in Figure 13, using another example. Figure 13 shows a zero-out in a block to which an 8x8 LFNST is applied, using another example.
[0218] As shown in Figures 11 and 12, zero-out can be applied to all areas except those to which LFNST is applied, and as shown in Figure 13, zero-out can also be applied to only partial areas. In Figure 13, zero-out can be applied only to areas other than the upper left 8x8 area, and zero-out does not need to be applied to the lower right 4x4 block inside the upper left 8x8 area.
[0219] Various embodiments are derived by applying combinations of the simplification methods ((i), (ii)-(A), (ii)-(B), (iii)) to the aforementioned LFNST. Of course, the combinations for the simplification methods are not limited to the embodiments described below, and any combination can be applied to the LFNST.
[0220] Embodiment
[0221] - Limit the number of output data for forward LFNST to a maximum of 16 → (i)
[0222] When a -4x4 LFNST is applied, all areas where the 4x4 LFNST is not applied are zeroed out → (ii)-(A)
[0223] When an 8x8 LFNST is applied, all areas where the 8x8 LFNST is not applied are zeroed out → (ii)-(B)
[0224] - Check whether there is any non-zero data in the regions filled with existing zero values and in the regions filled with zeros due to additional zero-outs ((ii)-(A), (ii)-(B)), and only if no non-zero data is found, signal LFNST indexing → (iii)
[0225] In the above embodiment, when LFNST is applied, the area in which non-zero output data can exist is limited to the upper left 4x4 area. More specifically, in Figures 11(a) and 12(a), the 8th position in the scan sequence is the last position in which non-zero data can exist, and in Figures 11(b) and (d) and 12(b), the 16th position in the scan sequence (i.e., the lower right position of the upper left 4x4 block) is the last position in which non-zero data can exist.
[0226] Therefore, when LFNST is applied, it is possible to determine whether or not the LFNST index can signal after checking whether or not there is non-zero data at positions where the residual coding process is not permitted (beyond the last position).
[0227] In the zero-out method proposed in (ii), the number of data points generated when both linear transformation and LFNST are applied is reduced, thus reducing the computational load required when performing the overall transformation process. Specifically, when LFNST is applied, zero-out is also applied to the forward linear transformation output data in areas where LFNST is not applied, so there is no need to generate data for areas that will be zero-out from the time of the forward linear transformation. Therefore, the computational load required for generating such data can be saved. The additional effects of the zero-out method proposed in (ii) can be summarized as follows:
[0228] Firstly, as mentioned above, the amount of computation required to execute the entire transformation process is reduced.
[0229] In particular, applying (ii)-(B) reduces the computational complexity in the worst-case scenario, thereby lightening the transformation process. More specifically, while large-scale linear transformations generally require a large amount of computation, applying (ii)-(B) can reduce the number of data points derived as a result of a forward LFNST execution to 16 or less, and the effect of reducing the computational complexity of the transformation increases further as the overall block (TU or CU) size increases.
[0230] Secondly, the amount of computation required for the entire conversion process is reduced, thereby lowering the power consumption required to perform the conversion.
[0231] Thirdly, it reduces the latency associated with the conversion process.
[0232] Quadratic transformations like LFNST add computational complexity to existing linear transformations, thus increasing the overall delay time associated with the transformation. In particular, in the case of intra-prediction, since the reconstruction data of adjacent blocks is used in the prediction process, the increase in delay time due to the quadratic transformation during encoding leads to an increase in the delay time until reconstruction, which can lead to an overall increase in the delay time of intra-prediction encoding.
[0233] However, by applying the zero-out method presented in (ii), the delay time of the primary conversion execution can be significantly reduced when LFNST is applied, so the delay time for the entire conversion execution will remain the same or be reduced, making it easier to implement an encoding device.
[0234] On the other hand, conventional intra-prediction treated the block currently to be encoded as a single encoding unit and performed encoding without division. However, ISP (Intra Sub-Paritions) coding means dividing the block currently to be encoded horizontally or vertically and performing intra-predictive coding. In this case, encoding / decoding is performed on each divided block to generate a reconstructed block, and the reconstructed block is used as a reference block for the next divided block. For example, during ISP coding, one coding block may be divided into two or four sub-blocks and coded, and in ISP, one sub-block performs intra-prediction by referencing the reconstructed pixel value of the adjacent sub-block located to its left or adjacent upper. Hereafter, "coding" is used as a concept that includes both encoding performed in the encoding device and decoding performed in the decoding device.
[0235] Table 4 shows the number of subblocks that are divided according to the block size when an ISP is applied. Subpartitions divided by the ISP may also be called translated blocks (TUs).
[0236] [Table 4]
[0237] The ISP divides the block predicted by Luminetra into two or four subpartitionings vertically or horizontally, depending on the block size. For example, the minimum block size to which the ISP can apply is 4x8 or 8x4. If the block size is larger than 4x8 or 8x4, the block is divided into four subpartitionings.
[0238] Figures 14 and 15 show examples of subblocks into which a single coding block is divided. More specifically, Figure 14 shows an example of division when the coding block (width (W) × height (H)) is a 4 × 8 block or an 8 × 4 block, and Figure 15 shows an example of division when the coding block is not a 4 × 8 block, an 8 × 4 block, or a 4 × 4 block.
[0239] When applying ISP, subblocks are coded sequentially according to the division pattern, for example, horizontally or vertically, from left to right or top to bottom. After the inverse transformation and intra-prediction process for one subblock, followed by the reconstruction process, coding is performed for the next subblock. For the leftmost or topmost subblock, the reconstruction pixels of the already coded block are referenced, as in the normal intra-prediction method. Furthermore, if each edge of a subsequent internal subblock is not adjacent to a previous subblock, the reconstruction pixels of the adjacent coding block that has already been coded are referenced, as in the normal intra-prediction method, to derive the reference pixels adjacent to that edge.
[0240] In ISP coding mode, all subblocks may be coded in the same intra-prediction mode, and a flag indicating whether or not to use ISP coding and a flag indicating the direction (horizontal or vertical) of division are signaled. As shown in Figures 14 and 15, the number of subblocks can be adjusted to two or four depending on the block shape, and if the size (width × height) of one subblock is less than 16, division to that subblock can be prohibited, or ISP coding itself can be restricted from being applied.
[0241] On the other hand, in ISP prediction mode, one coding unit is divided into two or four partition blocks, i.e., subblocks, for prediction, and the same in-screen prediction mode is applied to these two or four partition blocks.
[0242] As mentioned above, the division direction can be either horizontal (when an M×N coding unit with horizontal and vertical lengths M and N respectively is divided horizontally, it is divided into M×(N / 2) blocks if it is divided into two, and into M×(N / 4) blocks if it is divided into four) or vertical (when an M×N coding unit is divided vertically, it is divided into (M / 2)×N blocks if it is divided into two, and into (M / 4)×N blocks if it is divided into four). When divided horizontally, the partition blocks are coded in order from top to bottom, and when divided vertically, the partition blocks are coded in order from left to right. The partition block currently being coded can be predicted by referring to the restored pixel values of the upper (left) partition block when the division is horizontal (vertical).
[0243] A transformation can be applied to the residual signal generated by the ISP prediction method on a partition block basis. Based on the forward direction, a first-order transformation (core transform or primary transform) can be applied, and not only the existing DCT-2 but also the DST-7 / DCT-8 combination-based MTS (Multiple Transform Selection) technology can be applied. The transformation coefficients generated by the first-order transformation can then be subjected to a forward LFNST (Low Frequency Non-Separable Transform) to produce the final modified transformation coefficients.
[0244] In other words, LFNST can be applied to partition blocks that have been divided using the ISP prediction mode, and as mentioned above, the same intra-prediction mode is applied to the divided partition blocks. Therefore, when selecting an LFNST set derived based on the intra-prediction mode, the derived LFNST set can be applied to all partition blocks. That is, since the same intra-prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks.
[0245] On the other hand, LFNST can only be applied to transformation blocks where both the width and height are 4 or greater. Therefore, if the width or height of a partition block divided according to the ISP prediction scheme is less than 4, LFNST will not be applied and the LFNST index will not be signaled. Also, when applying LFNST to each partition block, that partition block can be considered as a single transformation block. Of course, if the ISP prediction scheme is not applied, LFNST will be applied to the coding block.
[0246] The application of LFNST to each partition block is explained in detail as follows:
[0247] For example, after applying forward LFNST to individual partition blocks, zero-out is applied to the upper left 4x4 region, leaving only a maximum of 16 coefficients (8 or 16) according to the conversion coefficient scan order, and filling the remaining positions and regions with zero values.
[0248] Alternatively, for example, if the length of one side of the partition block is 4, LFNST can be applied only to the upper left 4x4 region. If the length of all sides of the partition block, i.e., the width and height, is 8 or greater, LFNST can be applied to the remaining 48 coefficients, excluding the lower right 4x4 region within the upper left 8x8 region.
[0249] Alternatively, as an example, to match the worst-case computational complexity to 8 multiplications per sample, if each partition block is 4x4 or 8x8, only 8 transformation coefficients can be output after applying forward LFNST. That is, if the partition block is 4x4, an 8x16 matrix is applied as the transformation matrix, and if the partition block is 8x8, an 8x48 matrix is applied as the transformation matrix.
[0250] On the other hand, in the current VVC standard, LFNST index signaling is performed on a coding unit basis. Therefore, in ISP prediction mode, when LFNST is applied to all partition blocks, the same LFNST index value can be applied to those partition blocks. That is, once an LFNST index value is sent at the coding unit level, that LFNST index can be applied to all partition blocks within the coding unit. As mentioned above, LFNST index values have values of 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate two transformation matrices that exist within a single LFNST set when LFNST is applied.
[0251] As described above, the LFNST set is determined by the intra-prediction mode, and in the case of ISP prediction mode, all partition blocks within the coding unit are predicted in the same intra-prediction mode, so the partition blocks can refer to the same LFNST set.
[0252] As another example, while LFNST index signaling is still performed on a coding unit basis, in ISP prediction mode, the decision to apply LFNST is not made uniformly for all partition blocks. Instead, a separate condition is used to determine whether to apply the LFNST index value signaled at the coding unit level to each partition block, or not to apply LFNST. Here, the separate condition is signaled in the form of a flag for each partition block via the bitstream. If the flag value is 1, the LFNST index value signaled at the coding unit level is applied; if the flag value is 0, LFNST is not applied.
[0253] On the other hand, in a coding unit to which ISP mode is applied, if the length of one side of the partition block is less than 4, an example of applying LFNST is as follows:
[0254] Firstly, if the partition block size is N×2 (2×N), then LFNST can be applied to the upper left M×2 (2×M) region (where M≦N). For example, if M=8, the upper left region becomes 8×2 (2×8), so the region containing 16 residual signals becomes the input to the forward LFNST, and an R×16 (R≦16) forward transformation matrix can be applied.
[0255] Here, the forward LFNST matrix may be a separate, additional matrix that is not currently included in the VVC standard. Also, for worst-case complexity adjustment, an 8x16 matrix obtained by sampling only the top 8 row vectors of a 16x16 matrix is used in the transformation. The method of complexity adjustment will be described in detail later.
[0256] Secondly, if the partition block size is N×1 (1×N), then LFNST can be applied to the upper left M×1 (1×M) region (where M≦N). For example, if M=16, then the upper left region becomes 16×1 (1×16), so the region containing 16 residual signals becomes the input to the forward LFNST, and an R×16 (R≦16) forward transformation matrix can be applied.
[0257] Here, the forward LFNST matrix in question may be a separate, additional matrix not currently included in the VVC standard. Furthermore, for worst-case complexity adjustment, an 8x16 matrix obtained by sampling only the top eight row vectors of a 16x16 matrix can be used in the transformation. The complexity adjustment method will be described in detail later.
[0258] The first and second embodiments may be applied simultaneously, or only one of the two embodiments may be applied. In particular, in the case of the second embodiment, experiments have shown that the improvement in compression performance obtained with existing LFNST is not relatively large compared to the LFNST index signaling cost, due to the consideration of a one-dimensional transformation in LFNST. However, in the case of the first embodiment, an improvement in compression performance similar to that obtained with existing LFNST was observed; that is, in the case of ISP, experiments have confirmed that the application of LFNST for 2×N and N×2 contributes to the actual compression performance.
[0259] In the current LFNST in VVC, symmetry between intra prediction modes is applied. The same LFNST set is applied to the two-directional modes centered around Mode 34 (prediction in the 45-degree diagonal direction to the lower right), for example, the same LFNST set is applied to Mode 18 (horizontal prediction mode) and Mode 50 (vertical prediction mode). However, when applying the forward LFNST to Modes 35 to 66, after transposing the input data, the LFNST is applied.
[0260] On the other hand, VVC supports the Wide Angle Intra Prediction (WAIP) mode, and the LFNST set is derived based on the intra prediction mode modified considering the WAIP mode. For the modes extended by WAIP, the symmetry is utilized to determine the LFNST set in the same way as general intra prediction direction modes. For example, since Mode -1 is symmetric to Mode 67, the same LFNST set is applied, and since Mode -14 is symmetric to Mode 80, the same LFNST set is applied. For Modes 67 to 80, after transposing the input data before applying the forward LFNST, the LFNST transformation is applied.
[0261] In the case of the LFNST applied to the upper left M×2 (M×1) block, the symmetry for the aforementioned LFNST cannot be applied because the block to which the LFNST is applied is non-square. Therefore, instead of applying the symmetry based on the intra prediction mode like the LFNST in Table 2, the symmetry between the M×2 (M×1) block and the 2×M (1×M) block can be applied.
[0262] Figure 16 is a diagram showing the symmetry between the M×2 (M×1) block and the 2×M (1×M) block by way of an example.
[0263] As shown in Figure 16, the second mode in an M×2 (M×1) block is symmetrical to the 66th mode in a 2×M (1×M) block, so the same set of LFNSTs can be applied to both the 2×M (1×M) block and the M×2 (M×1) block.
[0264] In this case, to apply the LFNST set that was applied to the M×2 (M×1) block to the 2×M (1×M) block, the LFNST set is selected based on mode 2 instead of mode 66. That is, before applying the forward LFNST, the input data of the 2×M (1×M) block can be truncated and then the LFNST can be applied.
[0265] Figure 17 is a diagram showing an example of a 2×M block being transposed.
[0266] Figure 17(a) illustrates how LFNST can be applied to a 2×M block by reading the input data in column-first order, and Figure 17(b) illustrates how LFNST can be applied to an M×2 (M×1) block by reading the input data in row-first order. The methods for applying LFNST to the upper left M×2 (M×1) or 2×M (M×1) block can be summarized as follows.
[0267] 1. First, the input data is arranged as shown in Figures 17(a) and (b) to construct the input vector for the forward LFNST. For example, referring to Figure 16, the input data can be arranged according to the order in Figure 17(b) for an M×2 block predicted in mode 2, and according to the order in Figure 17(a) for a 2×M block predicted in mode 66, after which the LFNST set for mode 2 can be applied.
[0268] 2. For M×2 (M×1) blocks, the LFNST set is determined based on a modified intra-prediction mode that takes WAIP into account. As mentioned above, a pre-configured mapping relationship exists between the intra-prediction mode and the LFNST set, which can be shown as a mapping table as in Table 2.
[0269] For a 2×M (1×M) block, the modes that are symmetric to the prediction mode in the diagonal direction of 45 degrees downward to the right (mode 34 in the case of VVC standard) are determined from the intra-prediction modes modified to take WAIP into account. Then, the LFNST set is determined based on these symmetric modes and the mapping table. The mode (y) that is symmetric to mode 34 is derived by the following number. The mapping table is explained in more detail below.
[0270]
number
[0271] 3. When applying forward LFNST, the input data prepared in step 1 is multiplied by the LFNST kernel to derive the conversion coefficient. The LFNST kernel is selected from the LFNST set determined in step 2 and the pre-specified LFNST index.
[0272] For example, if M=8 and a 16x16 matrix is applied as the LFNST kernel, this matrix is multiplied by 16 input data to generate 16 transformation coefficients. The generated transformation coefficients are then placed in the upper left 8x2 or 2x8 region according to the scan order used in the VVC standard.
[0273] Figure 18 shows an example of the scan sequence for an 8x2 or 2x8 region.
[0274] For areas other than the upper left 8x2 or 2x8 region, all values may be filled with zero (zero-out), or the existing transformation coefficients obtained by applying a linear transformation may be maintained as they are. The aforementioned pre-specified LFNST index may be one of the LFNST index values (0, 1, 2) that are tried when calculating the RD cost while changing the LFNST index value during the encoding process.
[0275] In configurations where the computational complexity for the worst-case scenario is kept below a certain level (e.g., 8 multiplications per sample), for example, one could generate only 8 transformation coefficients by multiplying by an 8x16 matrix obtained by taking only the top 8 rows of the 16x16 matrix, then arrange the 8 transformation coefficients according to the scan order shown in Figure 18, and apply zero-out to the remaining coefficient area. Adjusting the complexity for the worst-case scenario will be discussed later.
[0276] 4. When applying LFNST in the reverse direction, the number of conversion coefficients already set (e.g., 16) is used as the input vector. An LFNST kernel (e.g., a 16x16 matrix) derived from the LFNST set obtained in step 2 and the parsed LFNST index is selected, and then the LFNST kernel is multiplied by the input vector to derive the output vector.
[0277] For M×2 (M×1) blocks, the output vectors are arranged according to row-major order as shown in Figure 17(b), and for 2×M (1×M) blocks, the output vectors are arranged according to column-major order as shown in Figure 17(a).
[0278] The remaining regions, excluding the region where the output vector is located within the upper left M×2 (M×1) or 2×M (M×2) region, and the regions within the partition block other than the upper left M×2 (M×1) or 2×M (M×2) region, are configured to either be filled with zero values (zero-out) or to maintain the transformation coefficients restored during the residual coding and inverse quantization processes.
[0279] Similar to the third case, when constructing the input vector, the input data is arranged according to the scan order in FIG. 18, and the number of input data can be reduced (for example, from 16 to 8) to construct the input vector in order to keep the computational complexity for the worst case below a certain level.
[0280] For example, when M = 8 and using 8 input data, only the left 16×8 matrix can be taken from the 16×16 matrix for multiplication, and then 16 output data can be obtained. The adjustment of the complexity for the worst case will be described later.
[0281] In the above embodiment, when applying LFNST, the case of applying symmetry between the M×2 (M×1) block and the 2×M (1×M) block is presented. However, different LFNST sets can also be applied to the shapes of the two blocks according to other examples.
[0282] Hereinafter, various examples regarding the LFNST set configuration for the ISP mode and the mapping method using the intra prediction mode will be described.
[0283] In the case of the ISP mode, the LFNST set configuration is different from the existing LFNST set. In other words, a kernel different from the existing LFNST kernel may be applied, or a mapping table different from the mapping table between the intra prediction mode index currently applied to the VVC standard and the LFNST set may be applied. The mapping table currently applied to the VVC standard is as shown in Table 2.
[0284] In Table 2, the preModeIntra value means the intra prediction mode value changed considering WAIP, and the lfnstTrSetIdx value is the index value indicating a specific LFNST set. Each LFNST set is composed of two LFNST kernels.
[0285] When ISP prediction mode is applied, if both the width and height of each partition block are greater than or equal to 4, the same kernel as the LFNST kernel currently applied in the VVC standard may be applied, and the mapping table may also be applied as is. Of course, a different LFNST kernel and a different mapping table than those currently in the VVC standard may also be applied.
[0286] When ISP prediction mode is applied, if the width or height of each partition block is less than 4, a different LFNST kernel and mapping table may be applied, which differ from those currently in the VVC standard. Tables 5 through 7 below show the mapping tables between intra-prediction mode values (intra-prediction mode values modified to account for WAIP) and LFNST sets that can be applied to M×2 (M×1) blocks or 2×M (1×M) blocks.
[0287] [Table 5]
[0288] [Table 6]
[0289] [Table 7]
[0290] The first mapping table in Table 5 consists of seven LFNST sets, the mapping table in Table 6 consists of four LFNST sets, and the mapping table in Table 7 consists of two LFNST sets. As another example, when consisting of one LFNST set, the lfnstTrSetIdx value is fixed to 0 for the preModeIntra value.
[0291] The following describes how to maintain worst-case computational complexity when applying LFNST to ISP mode.
[0292] In ISP mode, LFNST application can be restricted to maintain the number of multiplications per sample (or per coefficient, per position) below a certain value when LFNST is applied. Depending on the size of the partition block, LFNST can be applied as follows to maintain the number of multiplications per sample (or per coefficient, per position) at 8 or less.
[0293] 1. If both the horizontal and vertical dimensions of a partition block are 4 or greater, the same calculation complexity adjustment method as the worst-case scenario for LFNST in the current VVC standard can be applied.
[0294] In other words, if the partition block is a 4x4 block, instead of a 16x16 matrix, an 8x16 matrix obtained by sampling the top 8 rows from a 16x16 matrix can be applied in the forward direction, and a 16x8 matrix obtained by sampling the left 8 columns from a 16x16 matrix can be applied in the reverse direction. Also, if the partition block is an 8x8 block, instead of a 16x48 matrix in the forward direction, an 8x48 matrix obtained by sampling the top 8 rows from a 16x48 matrix can be applied, and instead of a 48x16 matrix in the reverse direction, a 48x8 matrix obtained by sampling the left 8 columns from a 48x16 matrix can be applied.
[0295] For 4×N or N×4 (N>4) blocks, when performing a forward transformation, a 16×16 matrix is applied only to the upper-left 4×4 block. The resulting 16 coefficients are then placed in the upper-left 4×4 region, and the remaining region is filled with zeros. When performing a reverse transformation, the 16 coefficients located in the upper-left 4×4 block are arranged according to the scan order to construct an input vector. These vectors are then multiplied by a 16×16 matrix to generate 16 output data. The generated output data is placed in the upper-left 4×4 region, and the remaining region is filled with zeros.
[0296] In the case of 8×N or N×8 (N>8) blocks, when performing a forward transformation, a 16×48 matrix is applied only to the ROI region within the upper-left 8×8 block (the remaining region after excluding the lower-right 4×4 block from the upper-left 8×8 block). The 16 generated coefficients are then placed in the upper-left 4×4 region, and all other regions are filled with zero values. When performing a reverse transformation, the 16 coefficients located in the upper-left 4×4 block are arranged according to the scan order to construct an input vector, which is then multiplied by a 48×16 matrix to generate 48 output data. The generated output data fills the ROI region, and all remaining regions are filled with zero values.
[0297] 2. When the partition block size is N×2 or 2×N, and (M≦N)LFNST is applied to the upper left M×2 or 2×M region, a matrix sampled according to the N value can be applied.
[0298] If M=8, then for partition blocks where N=8, i.e., 8×2 or 2×8 blocks, in the case of a forward transformation, an 8×16 matrix obtained by sampling the top 8 rows from a 16×16 matrix is applied instead of a 16×16 matrix, and in the case of a reverse transformation, a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix is applied instead of a 16×16 matrix.
[0299] If N is greater than 8, in the case of a forward transformation, a 16x16 matrix is applied to the upper left 8x2 or 2x8 block, and the resulting 16 output data are placed in the upper left 8x2 or 2x8 block, with the remaining area filled with zeros. In the case of a reverse transformation, the 16 coefficients located in the upper left 8x2 or 2x8 block are arranged in scan order to form an input vector, and then multiplied by the corresponding 16x16 matrix to generate 16 output data. The resulting output data are placed in the upper left 8x2 or 2x8 block, and the remaining area is filled with zeros.
[0300] 3. When the partition block size is N×1 or 1×N, and (M≦N)LFNST is applied to the upper left M×1 or 1×M region, a matrix sampled according to the N value is applied.
[0301] When M=16, for partition blocks where N=16, i.e., 16×1 or 1×16 blocks, in the case of a forward transformation, an 8×16 matrix obtained by sampling the top 8 rows from a 16×16 matrix is applied instead of a 16×16 matrix, and in the case of a reverse transformation, a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix is applied instead of a 16×16 matrix.
[0302] If N is greater than 16, in the case of a forward transformation, a 16x16 matrix is applied to the upper left 16x1 or 1x16 block, and the resulting 16 output data are placed in the upper left 16x1 or 1x16 block, with the remaining area filled with zeros. In the case of a reverse transformation, the 16 coefficients located in the upper left 16x1 or 1x16 block are arranged in scan order to form an input vector, and then multiplied by the corresponding 16x16 matrix to generate 16 output data. The resulting output data are placed in the upper left 16x1 or 1x16 block, and the remaining area is filled with zeros.
[0303] Another example is to maintain the number of multiplication factors per sample (or per coefficient, per position) below a certain value by keeping it at 8 or less based on the size of the ISP coding unit, not the size of the ISP partition block. If there is only one ISP partition block that satisfies the conditions for LFNST to apply, the LFNST worst-case complexity calculation is applied based on the size of that coding unit, not the size of the partition block. For example, if a Luma coding block for a coding unit is divided into four 4x4 partition blocks and coded by ISP, and there are no non-zero conversion factors for two of these partition blocks, the other two partition blocks can be configured to generate 16 conversion factors each (not 8, based on the encoder).
[0304] The following describes how to signal the LFNST index when using ISP mode.
[0305] As mentioned above, the LFNST index has values of 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate one of the two LFNST kernel matrices included in the selected set of LFNSTs. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The current method of transmitting the LFNST index in the VVC standard is as follows:
[0306] 1. An LFNST index can be sent once per coding unit (CU), and in the case of a dual-tree, separate LFNST indexes are signaled for the luma block and chroma block, respectively.
[0307] 2. If the LFNST index is not signaled, the LFNST index value is determined to be the default value of 0 (infer). The following are cases in which the LFNST index value is inferred to be 0:
[0308] A. When a mode in which no transformation is applied (e.g., transform skip, BDPCM, lossless coding, etc.)
[0309] B. When the primary transformation is not DCT-2 (e.g., DST7 or DCT8), i.e., when the horizontal or vertical transformation is not DCT-2.
[0310] C. If the horizontal or vertical dimensions of the coding unit relative to the Luma block exceed the maximum size of the Luma conversion that can be converted, for example, if the maximum size of the Luma conversion that can be converted is 64, then LFNST cannot be applied if the size of the coding block relative to the Luma block is equivalent to 128 × 16.
[0311] In the case of a dual tree, it is determined whether the coding unit for the luma component and the coding unit for the chroma component will exceed the maximum luma conversion size. That is, it is checked whether the size of the luma block exceeds the maximum luma conversion size that can be converted, and for the chroma block, it is checked whether the length / width of the corresponding luma block for the color format and the size of the maximum luma conversion that can be converted exceed each other. For example, if the color format is 4:2:0, the length / width of the corresponding luma block will be twice that of the chroma block, and the size of the corresponding luma block conversion will be twice that of the chroma block. As another example, if the color format is 4:4:4, the length / width of the corresponding luma block and the size of the conversion will be the same as that of the corresponding chroma block.
[0312] A 64-length conversion or a 32-length conversion refers to a conversion applied to a horizontal or vertical dimension having lengths of 64 or 32, respectively, and "conversion size" refers to the respective lengths of 64 or 32.
[0313] If it is a single tree, check whether the size of the horizontally or vertically convertible luma block exceeds the maximum size of the luma conversion block that can be converted. If it does exceed this limit, LFNST index signaling may be omitted.
[0314] D. The LFNST index can only be sent if both the width and height of the coding unit are 4 or greater.
[0315] In the case of a dual tree, the LFNST index can only be signaled if both the horizontal and vertical dimensions for the relevant component (i.e., the luma or chroma component) are 4 or greater.
[0316] In the case of a single tree, the LFNST index can be signaled when both the horizontal and vertical lengths of the luma components are 4 or greater.
[0317] E. If the position of the last non-zero coefficient is not the DC position (top-left position of the block), in a dual-tree type luma block, if the position of the last non-zero coefficient is not the DC position, an LFNST index is sent. In a dual-tree type chroma block, if either the position of the last non-zero coefficient for Cb or the position of the last non-zero coefficient for Cr is not the DC position, the corresponding LNFST index is sent.
[0318] If it is a single-tree type, the LFNST index is sent if the position of the last non-zero coefficient of any of the Luma, Cb, or Cr components is not at the DC position.
[0319] Here, if the CBF (coded block flag) value, which indicates whether or not a conversion coefficient exists for a single conversion block, is 0, the position of the last non-zero coefficient for that conversion block is not checked in order to determine whether or not to perform LFNST index signaling. In other words, if the CBF value is 0, the conversion is not applied to that block, so when checking the conditions for LFNST index signaling, the position of the last non-zero coefficient does not need to be considered.
[0320] For example, 1) in a dual-tree type, if the component is luma and its CBF value is 0, the LFNST index is not signaled; 2) in a dual-tree type, if the component is chroma and its CBF value for Cb is 0 and its CBF value for Cr is 1, only the position of the last non-zero coefficient for Cr is checked and the corresponding LFNST index is sent; 3) in a single-tree type, only the position of the last non-zero coefficient is checked for components where the CBF value for luma, Cb, and Cr is 1.
[0321] If it is confirmed that a conversion coefficient exists at a location where an F.LFNST conversion coefficient cannot exist, LFNST index signaling can be omitted. For 4x4 and 8x8 conversion blocks, LFNST conversion coefficients exist at 8 locations from the DC position according to the conversion coefficient scan order in the VVC standard, and all remaining positions are filled with 0. For conversion blocks that are not 4x4 or 8x8, LFNST conversion coefficients exist at 16 locations from the DC position according to the conversion coefficient scan order in the VVC standard, and all remaining positions are filled with 0.
[0322] Therefore, if a non-zero conversion coefficient exists in the region that should be filled with zero values after residual coding, LFNST index signaling can be omitted.
[0323] On the other hand, ISP mode may apply only to luma blocks, or it may apply to both luma and chroma blocks. As mentioned above, when ISP prediction is applied, the coding unit is divided into two or four partition blocks for prediction, and the transformation is applied to each of these partition blocks. Therefore, when determining the conditions for signaling the LFNST index on a coding unit basis, the fact that LFNST can be applied to each of the partition blocks must be taken into consideration. Also, if ISP prediction mode is applied only to a specific component (e.g., luma blocks), the fact that the component is divided into partition blocks only must be taken into consideration when signaling the LFNST index. When ISP mode is applied, the possible LFNST index signaling methods can be summarized as follows.
[0324] 1. An LFNST index can be sent once per coding unit (CU), and in the case of a dual-tree, separate LFNST indexes can be signaled for the luma block and chroma block, respectively.
[0325] 2. If the LFNST index is not signaled, the LFNST index value is determined to the default value of 0 (infer). The following are cases in which the LFNST index value is inferred to 0:
[0326] A. When a mode in which no transformation is applied (e.g., transform skip, BDPCM, lossless coding, etc.)
[0327] B. If the horizontal or vertical dimensions of the coding unit relative to the Luma block exceed the maximum size of the Luma conversion that can be converted, for example, if the maximum size of the Luma conversion that can be converted is 64, then LFNST cannot be applied if the size of the coding block relative to the Luma block is the same as 128 × 16.
[0328] Alternatively, the decision to signal the LFNST index can be based on the size of the partition block instead of the coding unit. That is, if the horizontal or vertical length of the partition block relative to the given luma block exceeds the size of the maximum luma conversion that can be converted, the LFNST index signaling can be omitted, and the LFNST index value can be inferred as 0.
[0329] In the case of a dual tree, it is determined whether the coding unit or partition block for the luma component and the coding unit or partition block for the chroma component each exceed the maximum conversion block size. Specifically, the length and width of the coding unit or partition block for luma are compared with the maximum luma conversion size, and if either is greater than the maximum luma conversion size, LFNST is not applied. In the case of a coding unit or partition block for chroma, the width / length of the corresponding luma block for the color format is compared with the maximum possible luma conversion size. For example, if the color format is 4:2:0, the width / length of the corresponding luma block will be twice that of the chroma block, and the conversion size of the corresponding luma block will be twice that of the chroma block. As another example, if the color format is 4:4:4, the width / length and conversion size of the corresponding luma block will be the same as that of the corresponding chroma block.
[0330] In the case of a single tree, after checking whether the horizontal or vertical conversion of a luma block (coding unit or partition block) exceeds the maximum luma conversion block size that can be converted, LFNST index signaling may be omitted if it does.
[0331] C. If LFNST, which is included in the current VVC standard, is applied, the LFNST index can only be sent if both the width and height of the partition block are 4 or greater.
[0332] If we were to apply LFNST to 2×M(1×M) or M×2(M×1) blocks in addition to the LFNST currently included in the VVC standard, then the LFNST index could only be sent if the partition block size is greater than or equal to a 2×M(1×M) or M×2(M×1) block. Here, P×Q block being greater than or equal to an R×S block means that P≧R and Q≧S.
[0333] In summary, an LFNST index can only be sent if the partition block is greater than or equal to the minimum size for which LFNST is applicable. In the case of a dual tree, an LFNST index can only be signaled if the partition block for the luma or chroma component is greater than or equal to the minimum size for which LFNST is applicable. In the case of a single tree, an LFNST index can only be signaled if the partition block for the luma component is greater than or equal to the minimum size for which LFNST is applicable.
[0334] In this document, an M×N block being greater than or equal to a K×L block means that M is greater than or equal to K and N is greater than or equal to L. An M×N block being greater than a K×L block means that M is greater than or equal to K and N is greater than or equal to L, while M is greater than K or N is greater than L. An M×N block being less than or equal to a K×L block means that M is less than or equal to K and N is less than or equal to L, while M is less than or equal to K and N is less than or equal to L.
[0335] D. If the position of the last non-zero coefficient is not the DC position (top-left corner of the block), then in the case of a dual-tree type chroma block, an LFNST can be transmitted if the position of the last non-zero coefficient in at least one of the partition blocks is not the DC position. In the case of a dual-tree type chroma block, the LNFST index can be transmitted if the position of the last non-zero coefficient in all partition blocks for Cb (assuming there is one partition block if the ISP mode is not applied to the chroma component) and the position of the last non-zero coefficient in all partition blocks for Cr (assuming there is one partition block if the ISP mode is not applied to the chroma component) are not the DC position.
[0336] In the case of a single-tree type, if the position of the last non-zero coefficient in any one of the partition blocks for the Luma component, Cb component, and Cr component is not at the DC position, the corresponding LFNST index can be sent.
[0337] Here, if the CBF (coded block flag) value, which indicates whether or not a conversion coefficient exists for each partition block, is 0, the position of the last non-zero coefficient for that partition block is not checked in order to determine whether or not to perform LFNST index signaling. In other words, since no conversion is applied to that block when the CBF value is 0, the position of the last non-zero coefficient for that partition block is not considered when checking the conditions for LFNST index signaling.
[0338] For example, 1) in a dual-tree type with luma components, if the corresponding CBF value is 0 for each partition block, the corresponding partition block is excluded when deciding whether or not to perform LFNST index signaling; 2) in a dual-tree type with chroma components, if the CBF value for Cb is 0 and the CBF value for Cr is 1 for each partition block, only the position of the last non-zero coefficient for Cr is checked to decide whether or not to perform LFNST index signaling; 3) in a single-tree type, the position of the last non-zero coefficient is checked only for blocks where the CBF value is 1 for all partition blocks of luma, Cb, and Cr components to decide whether or not to perform LFNST index signaling.
[0339] In ISP mode, the video information may be configured so as not to check the position of the last non-zero coefficient, and embodiments relating to this are as follows:
[0340] i. In ISP mode, the check for the position of the last non-zero coefficient is omitted for both luma blocks and chroma blocks, and LFNST index signaling is allowed. That is, for all partition blocks, LFNST index signaling is allowed even if the position of the last non-zero coefficient is the DC position or the corresponding CBF value is 0.
[0341] ii. In ISP mode, the check regarding the position of the last non-zero coefficient is omitted only for luma blocks, while in chroma blocks, the check regarding the position of the last non-zero coefficient is performed using the method described above. For example, in the case of a dual-tree type and a luma block, LFNST index signaling is allowed without checking the position of the last non-zero coefficient, while in the case of a dual-tree type and a chroma block, the existence of a DC position for the position of the last non-zero coefficient is checked using the method described above to determine whether or not to signal the corresponding LFNST index.
[0342] iii. If the system is in ISP mode and of a single-tree type, then method i or ii above is applied. That is, when method i is applied to an ISP mode and single-tree type, the check for the position of the last non-zero coefficient is omitted for both the luma block and the chroma block, and LFNST index signaling is permitted. Alternatively, method ii is applied, and the check for the position of the last non-zero coefficient is omitted for partition blocks of the luma component, while for partition blocks of the chroma component (if ISP is not applied to the chroma component, the number of partition blocks is assumed to be 1), the check for the position of the last non-zero coefficient is performed using the method described above to determine whether or not to perform the corresponding LFNST index signaling.
[0343] E. If it is confirmed that a conversion coefficient exists in a location that is not a possible location for any of the partition blocks, LFNST index signaling can be omitted.
[0344] For example, in the case of a 4x4 partition block and an 8x8 partition block, LFNST conversion coefficients exist at 8 positions from the DC position according to the conversion coefficient scan order in the VVC standard, and all remaining positions are filled with 0. Also, if the partition block is larger than or equal to 4x4 but is not a 4x4 or 8x8 partition block, LFNST conversion coefficients exist at 16 positions from the DC position according to the conversion coefficient scan order in the VVC standard, and all remaining positions are filled with 0.
[0345] Therefore, if a non-zero conversion coefficient exists in the region that should be filled with zero values after residual coding, LFNST index signaling can be omitted.
[0346] If LFNST can be applied to partition blocks of 2×M(1×M) or M×2(M×1), the region where LFNST conversion coefficients can be located can be specified as follows. Regions outside the region where conversion coefficients can be located are filled with zeros, and if there are non-zero conversion coefficients in regions that should be filled with zeros when LFNST is assumed to be applied, LFNST index signaling can be omitted.
[0347] i. LFNST can be applied to a 2×M or M×2 block, and if M=8, then only 8 LFNST conversion coefficients are generated for a 2×8 or 8×2 partition block. If the conversion coefficients are arranged in the scan order shown in Figure 18, then 8 conversion coefficients are placed in scan order starting from the DC position, and the remaining 8 positions are filled with 0.
[0348] For 2×N or N×2 (N>8) partition blocks, 16 LFNST conversion coefficients are generated. If the conversion coefficients are arranged in the scan order shown in Figure 18, the 16 conversion coefficients are placed in scan order from the DC position, and the remaining area is filled with 0. That is, in a 2×N or N×2 (N>8) partition block, the area other than the upper left 2×8 or 8×2 block is filled with 0. For 2×8 or 8×2 partition blocks, 16 conversion coefficients are generated instead of 8 LFNST conversion coefficients, and in this case, no area that must be filled with 0 occurs. As mentioned above, when LFNST is applied, if it is detected that a non-zero conversion coefficient exists in an area designated to be filled with 0 in even one partition block, LFNST index signaling can be omitted, and the LFNST index can be inferred as 0.
[0349] ii. LFNST can be applied to a 1×M or M×1 block, and if M=16, only 8 LFNST conversion coefficients are generated for a 1×16 or 16×1 partition block. If the conversion coefficients are placed in a left-to-right or top-to-bottom scan order, 8 conversion coefficients are placed in that scan order starting from the DC position, and the remaining 8 positions are filled with 0.
[0350] For a 1×N or N×1 (N>16) partition block, 16 LFNST conversion coefficients are generated. If the conversion coefficients are placed in a scan order from left to right or top to bottom, the 16 conversion coefficients are placed in that scan order starting from the DC position, and the remaining area is filled with 0. In other words, in a 1×N or N×1 (N>16) partition block, the area other than the upper left 1×16 or 16×1 block is filled with 0.
[0351] For 1x16 or 16x1 partition blocks, 16 conversion coefficients are generated instead of 8 LFNST conversion coefficients, and in this case, no areas that must be filled with zeros occur. As mentioned above, when LFNST is applied, if it is detected that a non-zero conversion coefficient exists in an area designated to be filled with zeros in even one partition block, LFNST index signaling can be omitted and the LFNST index can be inferred as 0.
[0352] On the other hand, in ISP mode, the current VVC standard considers the length conditions independently for the horizontal and vertical directions, and applies DST-7 instead of DCT-2 without signaling to the MTS index. It is determined whether the vertical or horizontal length is greater than or equal to 4 and less than or equal to 16, and the primary conversion kernel is determined according to the result of the determination. Therefore, the following conversion combination configurations are possible when LFNST can be applied while in ISP mode.
[0353] 1. If the LFNST index is 0 (including cases where the LFNST index is inferred to be 0), the determination conditions for the linear transformation when it is an ISP currently included in the VVC standard are followed. That is, the length conditions (greater than 4, equal to 4 and less than 16, or equal to 4) are checked independently for the horizontal and vertical directions, and if they are satisfied, DST-7 is applied instead of DCT-2 for the linear transformation; otherwise, DCT-2 is applied.
[0354] 2. When the LFNST index is greater than 0, the following two configurations are possible with a linear transformation:
[0355] A. DCT-2 can be applied to both horizontal and vertical directions.
[0356] B. The determination conditions for the first-order conversion when it is an ISP currently included in the VVC standard can be followed. That is, the length conditions (greater than 4, equal to 4 and less than 16, or equal to 4) are checked independently for the horizontal and vertical directions, and if they are satisfied, DST-7 is applied instead of DCT-2; otherwise, DCT-2 is applied.
[0357] In ISP mode, the video information can be configured so that the LFNST index is transmitted per partition block rather than per coding unit. In such cases, the LFNST index signaling scheme described above can be used to determine whether or not to perform LFNST index signaling, assuming that there is only one partition block within the unit to which the LFNST index is transmitted.
[0358] On the other hand, the following section examines the signaling order of the LFNST index and the MTS index.
[0359] For example, the LFNST index, signaled by residual coding, can be coded after the coding position for the last non-zero coefficient position, and the MTS index can be coded immediately after the LFNST index. In such a configuration, the LFNST index can be signaled for each conversion unit. Alternatively, even without signaling by residual coding, the LFNST index can be coded after the coding for the last effective coefficient position, and the MTS index can be coded after the LFNST index.
[0360] The syntax for a typical residual coding example is as follows:
[0361] [Table 8-1]
[0362] [Table 8-2]
[0363] The meanings of the major variables shown in Table 8 are as follows:
[0364] 1. cbWidth, cbHeight: The current width and height of the coding block.
[0365] 2. log2TbWidth, log2TbHeight: Base - 2 log values of the current width and height of the Transform Block. Zero-out is reflected, and non-zero coefficients can exist in the upper left region.
[0366] 3. sps_lfnst_enabled_flag: This flag indicates whether LFNST is applicable (enable). A flag value of 0 indicates that LFNST is not applicable, and a flag value of 1 indicates that LFNST is applicable. It is defined in the Sequence Parameter Set (SPS).
[0367] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponds to the variable chType and the (x0, y0) position. chType can have values of 0 or 1, where 0 indicates the luma component and 1 indicates the chroma component. The (x0, y0) position indicates the position on the picture, and the CuPredMode[chType][x0][y0] value allows for MODE_INTRA (intra prediction) and MODE_INTER (inter prediction).
[0368] 5. IntraSubPartitionsSplit[x0][y0]: The content for the (x0, y0) position is the same as in 4 above. It indicates what kind of ISP splitting was applied at the (x0, y0) position, and ISP_NO_SPLIT indicates that the coding unit corresponding to the (x0, y0) position is not split into partition blocks.
[0369] 6. intra_mip_flag[x0][y0]: The content for the (x0, y0) position is the same as in 4 above. intra_mip_flag is a flag that indicates whether the MIP (Matrix-based Intra Prediction) prediction mode is applied. A flag value of 0 indicates that MIP is not applicable, and a flag value of 1 indicates that MIP is applied.
[0370] 7. cIdx: A value of 0 indicates luma, while values of 1 and 2 indicate the chromatic components Cb and Cr, respectively.
[0371] 8. treeType: Refers to single-tree and dual-tree types (SINGLE_TREE: single-tree, DUAL_TREE_LUMA: dual-tree for luma component, DUAL_TREE_CHROMA: dual-tree for chroma component)
[0372] 9. tu_cbf_cb[x0][y0]: The content for the position (x0, y0) is the same as in 4 above. It indicates the CBF (Coded Block Flag) for the Cb component. If its value is 0, it means that there are no non-zero coefficients in the corresponding transformation unit for the Cb component, and if it is 1, it means that there are non-zero coefficients in the corresponding transformation unit for the Cb component.
[0373] 10. lastSubBlock: Indicates the scan order position of the sub-block (Coefficient Group (CG)) where the last effective coefficient (lastnon-zero coefficient) is located. 0 indicates a sub-block containing the DC component, while a value greater than 0 indicates a sub-block that does not contain the DC component.
[0374] 11. lastScanPos: Indicates the position of the last effective coefficient in the scan order within a subblock. If a subblock consists of 16 positions, the possible values are from 0 to 15.
[0375] 12.lfnst_idx[x0][y0]: This is the LFNST index syntax element to be parsed. If it is not parsed, it is inferred to a value of 0. That is, the default value is set to 0, indicating that LFNST will not be applied.
[0376] 13. LastSignificantCoeffX, LastSignificantCoeffY: Indicates the x and y coordinates in which the last significant coefficient is located within the transformation block. The x coordinate increases from left to right, starting from 0, and the y coordinate increases from top to bottom, starting from 0. If the values of both variables are all 0, it means that the last significant coefficient is located at DC.
[0377] 14. cu_sbt_flag: This flag indicates whether the SubBlock Transform (SBT) currently included in the VVC standard is applicable. A flag value of 0 indicates that SBT is not applicable, and a flag value of 1 indicates that SBT is applicable.
[0378] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: These flags indicate whether explicit MTS has been applied to the interCU and intraCU, respectively. A flag value of 0 indicates that MTS cannot be applied to the interCU or intraCU, while a value of 1 indicates that it can be applied.
[0379] 16. tu_mts_idx[x0][y0]: This is the MTS index syntax element to be parsed. If parsed, it is inferred to a value of 0. That is, the default value is set to 0, indicating that DCT-2 is applied to all horizontal and vertical directions.
[0380] As shown in Table 8, in the case of a single tree, the signaling of the LFNST index can be determined by only the condition of the position of the last effective coefficient for Luma. That is, if the position of the last effective coefficient is not DC, and the last effective coefficient is located inside the upper left subblock (CG), for example, a 4x4 block, the LFNST index is signaled. In the case of 4x4 and 8x8 transformation blocks, the LFNST index is signaled only if the last effective coefficient is located within the upper left subblock at a position between 0 and 7.
[0381] In the case of a dual tree, the lumana and chromana are each independently signaled with LFNST indices. In the case of chromana, the LFNST index can be signaled by applying the last effective coefficient position condition only to the Cb component. The condition is not checked for the Cr component, and if the CBF value for Cb is 0, the LFNST index can be signaled by applying the last effective coefficient position condition to the Cr component.
[0382] In Table 8, "Min(log2TbWidth, log2TbHeight)>=2" can be expressed as "Min(tbWidth, tbHeight)>=4", and "Min(log2TbWidth, log2TbHeight)>=4" can be expressed as "Min(tbWidth, tbHeight)>=16".
[0383] In Table 8, log2ZoTbWidth and log2ZoTbHeight represent the base-2 (base-2) logarithmic values of the width and height of the upper-left region where the last effective coefficient can exist due to zeroing out.
[0384] As shown in Table 8, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places: firstly, before the MTS index or LFNST index values are parsed, and secondly, after the MTS index has been parsed.
[0385] The first update occurs before the MTS index (tu_mts_idx[x0][y0]) value is parsed, so log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.
[0386] After the MTS index is parsed, log2ZoTbWidth and log2ZoTbHeight are set if the MTS index value is greater than 0 (i.e., a DST-7 / DCT-8 combination). When applying DST-7 / DCT-8 independently to the horizontal and vertical directions in a linear transformation, there can be up to 16 effective coefficients per row or column for each direction. That is, after applying DST-7 / DCT-8 for a length of 32 or more, up to 16 transformation coefficients can be derived per row or column from the left or top. Therefore, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions of a 2D block, effective coefficients can only exist up to a maximum of 16x16 area at the top left.
[0387] Furthermore, when DCT-2 is applied independently to the horizontal and vertical directions in a linear transformation, there can be up to 32 effective coefficients per row or column in each direction. That is, when applying DCT-2 to a length of 64 or more, up to 32 transformation coefficients can be derived from the left or top side for each row or column. Therefore, when DCT-2 is applied to both the horizontal and vertical directions of a 2D block, effective coefficients can only exist up to a maximum of 32x32 area at the top left.
[0388] Furthermore, when DST-7 / DCT-8 is applied to one direction and DCT-2 is applied to the other direction, 16 effective coefficients can exist in the former direction and 32 effective coefficients can exist in the latter direction. For example, in a 64x8 conversion block, where DCT-2 is applied to the horizontal direction and DST-7 is applied to the vertical direction (which can occur in situations where implicit MTS is applied), effective coefficients can exist in a maximum 32x8 area at the upper left corner.
[0389] If log2ZoTbWidth and log2ZoTbHeight are updated in two places, as shown in Table 8, that is, before MTS index parsing, then the ranges of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight, as shown in the table below.
[0390] [Table 9]
[0391] Furthermore, in such cases, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2TbHeight values in the binary evolution process for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.
[0392] [Table 10]
[0393] On the other hand, for example, if ISP mode is enabled and LFNST is applied, the spec text can be constructed as shown in Table 11 when the signaling in Table 8 is applied. Compared with Table 8, the condition for signaling the LFNST index only when ISP mode is not enabled (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT in Table 8) has been removed.
[0394] In the case of a single tree, if the LFNST index sent when it is a luma (when cIdx=0) is to be reused when it is a chroma, the LFNST index sent to the first ISP partition block where the effective coefficient exists can be applied to the chroma conversion block. Alternatively, even in the case of a single tree, the LFNST index can be signaled separately for the chroma component from the luma component. Explanations for the variables listed in Table 11 are as shown in Table 8.
[0395] [Table 11]
[0396] If, as in another example, in Table 11, when it is an ISP, the last effective coefficient is allowed to be located only at the DC position for all partition blocks, then the parsing conditions for the LFNST index can be modified as follows:
[0397] [Table 12]
[0398] On the other hand, as an example, the LFNST index and / or the MTS index can be signaled at the coding unit level. As mentioned above, the LFNST index can have three values: 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate the first and second candidates, respectively, of the two LFNST kernel candidates included in the selected LFNST set. The LFNST index is coded via truncated unary binarization, and the values 0, 1, and 2 can be coded in bin strings 0, 10, and 11, respectively.
[0399] For example, LFNST can only be applied when DCT-2 is applied to both the horizontal and vertical directions in a linear transformation. Therefore, if the MTS index is signaled after LFNST index signaling, the MTS index can only be signaled if the LFNST index value is 0. If the LFNST index is not 0, the linear transformation can be performed by applying DCT-2 to both the horizontal and vertical directions without signaling the MTS index.
[0400] The MTS index value can have values of 0, 1, 2, 3, and 4, where 0, 1, 2, 3, and 4 indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, and DCT-8 / DCT-8 are applied to the horizontal and vertical directions, respectively. The MTS index can also be coded via truncated unali dimorphism, where the values 0, 1, 2, 3, and 4 can be coded in bin strings 0, 10, 110, 1110, and 1111, respectively.
[0401] The signaling of LFNST indexes at the coding unit level can be shown in the table below. LFNST indexes can be signaled later in the coding unit syntax table.
[0402] [Table 13]
[0403] On the other hand, as an example, the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag in Table 13 can be set as shown in Table 16 below.
[0404] The variable LfnstDcOnly is set to 1 if the last effective coefficients of a transformation block with a corresponding CBF (Coded Block Flag, 1 if at least one effective coefficient exists in the block, 0 otherwise) value are all located at the DC position (top left corner), and 0 otherwise. More specifically, in the case of a dual-tree luma, the position of the last effective coefficient is checked for one luma transformation block, and in the case of a dual-tree chroma, the position of the last effective coefficient is checked for both the transformation block for Cb and the transformation block for Cr. In the case of a single tree, the position of the last effective coefficient can be checked for the transformation blocks for luma, Cb, and Cr.
[0405] The variable LfnstZeroOutSigCoeffFlag is 0 if an effective coefficient exists at a position where LFNST would result in a zero-out when applied, and 1 otherwise.
[0406] In Table 13 and the following tables, lfnst_idx[x0][y0] indicates the LFNST index for the corresponding coding unit, and tu_mts_idx[x0][y0] indicates the MTS index for the corresponding coding unit.
[0407] For example, if you intend to code an MTS index at the coding unit level, following an LFNST index, the coding unit syntax table can be configured as shown in Table 14.
[0408] [Table 14]
[0409] When comparing Table 14 with Table 13, the condition for signaling lfnst_idx[x0][y0] has been changed from checking whether the tu_mts_idx[x0][y0] value is 0 (i.e., checking whether both the horizontal and vertical directions are DCT-2) to checking whether the transform_skip_flag[x0][y0] value is 0 (!transform_skip_flag[x0][y0]). transform_skip_flag[x0][y0] indicates whether the coding unit has been coded into a transformation skip mode where transformations are omitted, and this flag is signaled before the MTS index and LFNST index. That is, because lfnst_idx[x0][y0] is signaled before the tu_mtx_idx[x0][y0] value is signaled, only the condition for the transform_skip_flag[x0][y0] value can be checked.
[0410] As shown in Table 14, when coding tu_mts_idx[x0][y0], various conditions are checked, and as mentioned above, tu_mts_idx[x0][y0] is signaled only when the lfnst_idx[x0][y0] value is 0.
[0411] Furthermore, tu_cbf_luma[x0][y0] is a flag indicating whether an effective coefficient exists for the luma component, and cbWidth and cbHeight represent the width and height of the coding unit for the luma component, respectively.
[0412] Furthermore, in Table 14, (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT) indicates that ISP mode is not being used, and (!cu_sbt_flag) indicates that SBT is not being applied.
[0413] According to Table 14, when both the width and height of the coding unit for the luma component are 32 or less, tu_mts_idx[x0][y0] is signaled, meaning that the applicability of MTS is determined by the width and height of the coding unit for the luma component.
[0414] In other cases where TU tiling occurs (for example, if the maximum transformation size is set to 32, a 64x64 coding unit is divided into four 32x32 transformation blocks for coding), the MTS index can be signaled based on the size of each transformation block. For example, when both the width and height of a transformation block are 32 or less, the same MTS index value can be applied to all transformation blocks within a coding unit, and the same primary transformation can be applied. Also, when TU tiling occurs, the tu_cbf_luma[x0][y0] value in Table 14 is the CBF value for the top-left transformation block, or it can be set to 1 if at least one transformation block has a corresponding CBF value of 1.
[0415] For example, if ISP mode is currently applied to a block, LFNST can be applied, in which case Table 14 can be modified as shown in Table 15.
[0416] [Table 15]
[0417] As shown in Table 15, even in ISP mode, it is possible to configure the system to signal (IntraSubPartitionsSplitType!=ISP_NO_SPLIT)lfnst_idx[x0][y0], and the same LFNST index value can be applied to all ISP partition blocks.
[0418] Furthermore, as shown in Table 15, tu_mts_idx[x0][y0] can only be signaled when not in ISP mode, so the MTS index coding portion is as shown in Table 14.
[0419] As shown in Tables 14 and 15, when the MTS index is signaled immediately after the LFNST index, information for the linear transformation cannot be known when performing residual coding. That is, the MTS index is signaled after residual coding. Therefore, the part of the residual coding section that performs zero-out, leaving only 16 coefficients for a 32-length DST-7 or DCT-8, can be modified as shown in Table 16 below.
[0420] [Table 16-1]
[0421] [Table 16-2]
[0422] As shown in Table 16, the part of the process of determining log2ZoTbWidth and log2ZoTbHeight (where log2ZoTbWidth and log2ZoTbHeight represent the base-2 log values of the width and height of the upper-left region remaining after zeroing out, respectively) in which the tu_mts_idx[x0][y0] value is checked can be omitted.
[0423] The binary representations for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 16 can be determined based on log2ZoTbWidth and log2ZoTbHeight, as shown in Table 10.
[0424] Furthermore, as shown in Table 16, when determining log2ZoTbWidth and log2ZoTbHeight using residual coding, a condition can be added to check sps_mts_enable_flag.
[0425] The TR in Table 10 represents the truncated rice binarization method, and based on the cMax and cRiceParam defined in Table 10, the final effective coefficient information can be binarized using the method described in the following table.
[0426] [Table 17]
[0427] On the other hand, as in other examples, the coding unit syntax table and the residual coding syntax table are as shown in the table below.
[0428] [Table 18]
[0429] [Table 19]
[0430] In Table 18, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed by residual coding in Table 19. The variable MtsZeroOutSigCoeffFlag is changed from 1 to 0 if there are significant coefficients in the region that should be filled with 0 by zero-out (LastSignificantCoeffX>15||LastSignificantCoeffY>15), in which case the MTS index is not signaled, as shown in Table 19.
[0431] On the other hand, as an example, the MTS index coding portion in Table 18 can be modified as shown in the following table.
[0432] [Table 20]
[0433] Unlike Table 18, the variable MtsZeroOutSigCoeffFlag in Table 20 is initialized to 0 instead of 1 (MtsZeroOutSigCoeffFlag=0). The variable MtsZeroOutSigCoeffFlag maintains its value of 0 if there are significant coefficients in the region that should be filled with 0 by zero-out (LastSignificantCoeffX>15||LastSignificantCoeffY>15), and in this case, the MTS index is not signaled.
[0434] The following drawings have been prepared to illustrate a specific example of this specification. The names of specific devices and signals / messages / fields shown in the drawings are presented illustratively, and the technical features of this specification are not limited to the specific names used in the following drawings.
[0435] Figure 19 is a flowchart illustrating the operation of a video decoding device according to one embodiment described in this document.
[0436] Each step disclosed in Figure 19 is based in part on the details described in Figures 2 through 18. Therefore, specific details that overlap with those described in Figures 2 through 18 are omitted or simplified in their explanation.
[0437] A decoding device 200 according to one embodiment can perform residual coding based on residual information received from a bitstream, and specifically can parse the residual information received at the residual coding level and arrange the conversion coefficients for the current block in a predetermined scanning order (S1910).
[0438] The decoding device 200 can decode information about the quantized transformation coefficients for the current block from the bitstream and derive the quantized transformation coefficients for the target block based on the information about the quantized transformation coefficients for the current block. The information about the quantized transformation coefficients for the target block can be contained in an SPS (Sequence Parameter Set) or a slice header and may include at least one of the following: information on whether a simplified transformation (RST) is applied, information on the simplification factor, information on the minimum transformation size to apply the simplified transformation, information on the maximum transformation size to apply the simplified transformation, the inverse simplified transformation size, and information on a transformation index that points to one of the transformation kernel matrices contained in the transformation set.
[0439] The decoding device can also receive further information regarding the intra-predictive mode for the current block and information regarding whether ISP is applied to the current block. By receiving and parsing flag information indicating whether ISP coding or ISP mode is applied, the decoding device can derive whether the current block will be divided into a predetermined number of subpartition conversion blocks. Here, the current block is the coding block. The decoding device can also derive the size and number of subpartition blocks to be divided via flag information indicating the direction in which the current block will be divided.
[0440] The decoding device 200 can perform inverse quantization on the current block's residual information, i.e., the quantized conversion coefficients, to derive the conversion coefficients, and can arrange the derived conversion coefficients in a predetermined scanning order.
[0441] Specifically, the derived transformation coefficients can be arranged in a 4x4 block unit according to the reverse diagonal scan order, and the transformation coefficients within a 4x4 block can also be arranged according to the reverse diagonal scan order. In other words, the transformation coefficients after inverse quantization can be arranged according to the reverse scan order applied in VVC and HEVC video codecs.
[0442] The transformation coefficients derived based on such residual information may be inversely quantized transformation coefficients as described above, or they may be quantized transformation coefficients. In other words, the transformation coefficients should be data that can be checked to see if they are non-zero data in the current block, regardless of whether they can be quantized or not.
[0443] The decoding device can derive residual samples by applying an inverse transform to the quantized transformation coefficients.
[0444] As mentioned above, the decoding device can derive residual samples by applying either the non-separating transform LFNST or the separating transform MTS, respectively, which can be performed based on an LFNST kernel, i.e., an LFNST index that points to the LFNST matrix and an MTS index that points to the MTS kernel.
[0445] The decoding device can, for example, receive and parse at least one of the LFNST index or MTS index at the coding unit level, and can parse the LFNST index indicating the LFNST kernel before, i.e., immediately after, the MTS index indicating the MTS kernel (S1920).
[0446] Based on specific conditions for the LFNST index, for example, when the LFNST index value is 0, the MTS index can be parsed.
[0447] On the other hand, the residual coding level may include syntax for the last effective coefficient position information, and the LFNST index may be parsed after the last effective coefficient position information has been parsed.
[0448] For example, if the current block is a Luma block and the LFNST index is 0, the MTS index can be parsed. That is, if the current block is a Luma block and the LFNST index is greater than 0, the MTS index will not be parsed.
[0449] For example, if the current block tree type is a dual tree, the LFNST index for each of the luma blocks and chroma blocks can be parsed.
[0450] On the other hand, in the step of deriving the conversion coefficients, the width and height of the upper left region where the last effective coefficient can exist can be derived by zeroing out within the current block, and the width and height of the upper left region can be derived before parsing the MTS index.
[0451] On the other hand, the last effective coefficient position can be derived from the width and height of the upper left region, and the last effective coefficient position information can be binary based on the width and height of the upper left region.
[0452] Alternatively, for example, if the current block's tree type is a single-tree type, the decoding device can perform residual coding on the current block's luma and chroma blocks and then parse the LFNST index.
[0453] When an LFNST index is parsed at the coding unit level after residual coding has been performed, not at the conversion block level or the residual coding level, it is possible to receive an LFNST index that reflects the positions of the complete conversion coefficients for both the luma block and the chroma block, as well as zero-out information that occurred during the conversion process, rather than just conversion coefficient information for one of the luma blocks or chroma blocks.
[0454] Furthermore, if the current block tree type is a dual-tree type and chroma components are being coded, the decoding device can perform residual coding for the Cb and Cr components of the chroma block and then parse the LFNST index.
[0455] When an LFNST index is parsed at the coding unit level after residual coding has been performed, not at the conversion block level or the residual coding level, it is possible to receive an LFNST index that reflects the positions of the complete conversion coefficients for both the Cb and Cr components of the chroma block, as well as zero-out information that occurred during the conversion process, rather than just the conversion coefficient information for one of the Cb and Cr components of the chroma block.
[0456] Furthermore, if the current block is divided into multiple sub-partition blocks, the decoding device can perform residual coding on the multiple sub-partition blocks and then parse the LFNST index.
[0457] As described above, if the LFNST index is parsed at the coding unit level after residual coding has been performed, rather than at the conversion block level or the residual coding level, it is possible to receive an LFNST index that reflects the complete conversion coefficient positions for all subpartition blocks and zero-out information that occurred during the conversion process, rather than just conversion coefficient information for some or individual subpartition blocks.
[0458] On the other hand, if, for example, the current block is divided into the multiple subpartition blocks, the LFNST index can be parsed regardless of whether a conversion coefficient exists in the region excluding the DC location of each of the multiple subpartition blocks. That is, if ISP is applied to the current block, signaling of the LFNST index can be permitted if it is allowed that the last effective coefficient is located only at the DC location for all subpartition blocks.
[0459] On the other hand, the decoding device can derive a first variable in the residual coding step that indicates whether a conversion coefficient exists in the region excluding the DC position of the current block, and a second variable that indicates whether a conversion coefficient exists in a second region excluding the upper left corner first region of the current block or a subpartition block separated from the current block.
[0460] The decoding device can parse the LFNST index if a conversion coefficient exists in the region excluding the DC position, and no conversion coefficient exists in the second region.
[0461] Specifically, the decoding device can derive a first variable indicating whether the conversion coefficient, i.e., a valid coefficient, exists in the region excluding the DC position of the current block, in order to determine whether the LFNST index can be parsed.
[0462] The first variable is LfnstDcOnly, which can be derived during the residual coding process. The first variable can be derived to 0 if the index of the subblock containing the last effective coefficient in the current block is 0, and the position of the said last effective coefficient in the subblock is greater than 0. If the first variable is 0, the LFNST index can be parsed. A subblock refers to a 4x4 block used as a coding unit in residual coding, and can also be named CG (Coefficient Group). An index of 0 in a subblock refers to the top-left 4x4 subblock.
[0463] The first variable can initially be set to 1, and depending on whether an effective coefficient exists in the region excluding the DC position, it may remain at 1 or be changed to 0.
[0464] The variable LfnstDcOnly indicates whether a non-zero coefficient exists in a position other than the DC component for at least one transformation block within a coding unit. It can be 0 if a non-zero coefficient exists in a position other than the DC component for at least one transformation block within a coding unit, and 1 if no non-zero coefficients exist in a position other than the DC component for any of the transformation blocks within a coding unit.
[0465] Furthermore, the decoding device can check whether zeroing out has been performed on the second region by deriving a second variable that indicates whether an effective coefficient exists in the second region, which is the first region excluding the upper left corner of the current block.
[0466] The second variable is LfnstZeroOutSigCoeffFlag, which indicates that zero-out has been performed when LFNST is applied. The second variable is initially set to 1, and can also be changed to 0 if an effective coefficient exists in the second region.
[0467] The variable LfnstZeroOutSigCoeffFlag can be derived to 0 if the index of the subblock containing the last non-zero coefficient is greater than 0, the width and height of the transformation block are both equal to or greater than 4, or the last position of the non-zero coefficient within the subblock containing the last non-zero coefficient is greater than 7, or the size of the transformation block is 4x4 or 8x8. A subblock refers to a 4x4 block used as a coding unit in residual coding, and can also be named a CG (Coefficient Group). An index of 0 in a subblock refers to the top-left 4x4 subblock.
[0468] In other words, if a non-zero coefficient is derived in a region other than the upper-leftmost region where an LFNST conversion coefficient can exist in the conversion block, or if a non-zero coefficient exists in a region other than the 8th position in the scan order for 4x4 and 8x8 blocks, the variable LfnstZeroOutSigCoeffFlag is set to 0.
[0469] For example, when ISP is applied to a coding unit, if it is confirmed that a conversion coefficient exists in a location where an LFNST conversion coefficient cannot exist in at least one of the subpartition blocks, LFNST index signaling can be omitted. In other words, if zeroing out is not performed in one subpartition block and an effective coefficient exists in a second area, LFNST index signaling will not be performed.
[0470] On the other hand, the first region can be derived based on the current block size.
[0471] For example, if the current block size is 4x4 or 8x8, the first region extends from the top-left corner of the current block to the 8th sample position in the scan direction. When the current block is divided, if the subpartition block size is 4x4 or 8x8, the first region extends from the top-left corner of the subpartition block to the 8th sample position in the scan direction.
[0472] Currently, if the block size is 4x4 or 8x8, eight data points are output via the forward LFNST, so the eight conversion coefficients received by the decoding device can be arranged from the top left corner of the block to the eighth sample position in the scanning direction, as shown in Figures 11(a) and 12(a).
[0473] Furthermore, if the current block size is not 4x4 or 8x8, the first region is the 4x4 region at the top left of the current block. If the current block size is not 4x4 or 8x8, 16 data points are output via the forward LFNST, and the 16 conversion coefficients received by the decoding device can be arranged in the 4x4 region at the top left of the current block, as shown in Figures 11(b) to (d) and Figure 12(b).
[0474] On the other hand, the conversion coefficients that can be arranged in the first region can be arranged according to the diagonal scan direction, as shown in Figure 7.
[0475] As mentioned above, when the current block is divided into subpartition blocks, the decoding device can parse the LFNST index if no conversion coefficients exist in any of the individual second areas for any of the subpartition blocks. If a conversion coefficient exists in the second area for any one of the subpartition blocks, the LFNST index is not parsed.
[0476] As mentioned above, LFNST can be applied to subpartition blocks with a width and height of 4 or more, and the LFNST index for the current block, which is the coding block, can be applied to multiple subpartition blocks.
[0477] On the other hand, zero-outs reflected by LFNST (including all zero-outs that may occur as a result of applying LFNST) are also applied to the sub-partition block, so they are applied to both the first area and the sub-partition block in the same way. That is, if the divided sub-partition block is a 4x4 block or an 8x8 block, LFNST is applied to the conversion coefficients from the top left corner of the sub-partition block up to the 8th in the scan direction, and if the sub-partition block is not a 4x4 block or an 8x8 block, LFNST can be applied to the conversion coefficients of the top left 4x4 area of the sub-partition block.
[0478] Furthermore, in the residual coding step, the decoding device derives a third variable indicating whether a conversion coefficient exists in the region excluding the top-left 16x16 area of the current block. If no conversion coefficient exists in the region excluding the 16x16 area, the MTS index can be parsed.
[0479] The third variable is MtsZeroOutSigCoeffFlag, which indicates whether zero-out has been performed when MTS is applied. The variable MtsZeroOutSigCoeffFlag indicates whether a conversion coefficient exists in an area other than the upper left corner region where the last effective coefficient can exist due to zero-out after MTS execution, i.e., the upper left corner 16x16 area. It is initially set to 1, and if a conversion coefficient exists in an area other than the 16x16 area, its value can be changed from 1 to 0. If the value of the third variable is 0, the MTS index is not signaled.
[0480] The decoding device can derive a residual sample by applying at least one of either LFNST, which is performed based on the LFNST index, or MTS, which is performed based on the MTS index (S1930).
[0481] Next, the decoding device 200 can generate a reconstructed sample based on the residual sample for the current block and the predicted sample for the current block (S1940).
[0482] The following drawings have been prepared to illustrate a specific example of this specification. The names of specific devices and signals / messages / fields shown in the drawings are presented illustratively, and the technical features of this specification are not limited to the specific names used in the following drawings.
[0483] Figure 20 is a flowchart illustrating the operation of a video encoding device according to one embodiment described in this document.
[0484] Each step disclosed in Figure 20 is based in part on the details described in Figures 3 through 18. Therefore, specific details that overlap with those described in Figures 1 and 3 through 18 are omitted or simplified in their explanation.
[0485] An encoding device 100 according to one embodiment can derive predicted samples for the current block based on an intra-prediction mode applied to the current block (S2010).
[0486] The encoding device can perform predictions on a per-subpartition conversion block basis if an ISP is applied to the current block.
[0487] The encoding device can determine whether to apply ISP coding or ISP mode to the current block, i.e., the coding block. Based on this determination, it can decide in which direction the current block will be divided and derive the size and number of subblocks to be divided.
[0488] The encoding device 100 can derive the residual sample for the current block based on the predicted sample (S2020).
[0489] The encoding device 100 can apply at least one of LFNST or MTS to the residual sample to derive a conversion coefficient for the current block, and arrange the conversion coefficients in a predetermined scanning order (S2030).
[0490] The first-order transformation can be performed via multiple transformation kernels, similar to MTS, in which case the transformation kernel can be selected based on the intra-prediction mode.
[0491] Furthermore, the encoding device 100 can determine whether to perform a quadratic transformation or a non-separable transformation, specifically LFNST, on the transformation coefficients for the current block, and can derive modified transformation coefficients by applying LFNST to the transformation coefficients.
[0492] Unlike linear transformations, which separate the coefficients to be transformed vertically or horizontally, LFNST is a non-separated transformation that applies the transformation without separating the coefficients in a specific direction. Such a non-separated transformation is a low-frequency non-separated transformation that applies the transformation only to the low-frequency region, not to the entire target block being transformed.
[0493] The encoding device can determine whether LFNST can be applied to the height and width of the divided subpartition block if ISP is currently applied to the block.
[0494] The encoding device can determine whether LFNST can be applied to the height and width of the divided subpartition block. In this case, the decoding device can parse the LFNST index if the height and width of the subpartition block are 4 or greater.
[0495] The encoding device can encode at least one of the following: an LFNST index pointing to an LFNST kernel or an MTS index pointing to an MTS kernel (S2040).
[0496] Based on specific conditions for the LFNST index, for example, when the LFNST index value is 0, the MTS index can be encoded.
[0497] For example, the encoding device can encode the MTS index if the current block is a Luma block and the LFNST index indicates 0.
[0498] For example, if the current block tree type is a dual tree, the encoding device can encode LFNST indices for both the luma block and the chroma block.
[0499] For example, once the conversion coefficients are derived, the encoding device can derive the width and height of the upper-left region where the last effective coefficient can exist due to zero-outs in the current block, derive the position of the last effective coefficient based on this width and height of the upper-left region, and binary the information of the last effective coefficient position.
[0500] For example, the width and height of the upper left region can be derived before signaling the MTS index.
[0501] Alternatively, for example, if the current block's tree type is a single-tree type, the encoding device can encode the LFNST index at the coding unit level after deriving all the conversion coefficients for the current block's luma block and chroma block.
[0502] If the LFNST index is encoded at the coding unit level after all conversion coefficients that are not at the conversion block level or residual coding level have been derived, the encoded LFNST index can reflect the positions of the complete conversion coefficients for both the luma block and the chroma block, as well as the zero-out information that occurred during the conversion process, rather than just the conversion coefficient information for one of the luma blocks or chroma blocks.
[0503] Furthermore, if the current block tree type is a dual-tree type and the chroma component is being coded, the encoding device can derive all the conversion coefficients for the Cb and Cr components of the chroma block, and then encode the LFNST index at the coding unit level.
[0504] If the LFNST index is encoded at the coding unit level after all conversion coefficients that are not at the conversion block level or residual coding level have been derived, the encoded LFNST index can reflect the positions of the complete conversion coefficients for both the Cb and Cr components of the chroma block, as well as the zero-out information that occurred during the conversion process, rather than just the conversion coefficient information for one of the Cb and Cr components of the chroma block.
[0505] Furthermore, if a block is currently divided into multiple subpartition blocks, the encoding device can encode the LFNST index at the coding unit level after deriving all the conversion coefficients for the multiple subpartition blocks.
[0506] As described above, if the LFNST index is parsed at the coding unit level after all conversion coefficients that are not at the conversion block level or the residual coding level have been derived, an LFNST index can be encoded that reflects the positions of the complete conversion coefficients for all subpartition blocks, not just the conversion coefficient information for some or individual subpartition blocks, and the zero-out information that occurred during the conversion process.
[0507] On the other hand, if the current block is divided into the plurality of subpartition blocks, the encoding device can encode the LFNST index regardless of whether a conversion coefficient exists in the region excluding the DC position of each of the plurality of subpartition blocks. That is, if ISP is applied to the current block, signaling of the LFNST index can be permitted if it is allowed that the last effective coefficient is located only at the DC position for all subpartition blocks.
[0508] In the process of deriving the conversion coefficients, the encoding device can derive a first variable indicating whether the conversion coefficients exist in a region excluding the DC position of the current block, and a second variable indicating whether the conversion coefficients exist in a second region excluding the upper left corner of the current block or a subpartition block separated from the current block.
[0509] The encoding device can encode the LFNST index if a conversion coefficient exists in the region excluding the DC position, and no conversion coefficient exists in the second region.
[0510] Specifically, the first variable is the variable LfnstDcOnly, and can be derived to 0 if the index of the subblock containing the last effective coefficient in the current block is 0, and the position of the last effective coefficient in the subblock is greater than 0, and if the first variable is 0, the LFNST index can be encoded.
[0511] The first variable can initially be set to 1, and depending on whether an effective coefficient exists in the region excluding the DC position, it may remain at 1 or be changed to 0.
[0512] The variable LfnstDcOnly indicates whether a non-zero coefficient exists in a position other than the DC component for at least one transformation block within a coding unit. It can be 0 if a non-zero coefficient exists in a position other than the DC component for at least one transformation block within a coding unit, and 1 if no non-zero coefficients exist in a position other than the DC component for any of the transformation blocks within a coding unit.
[0513] Furthermore, after LFNST is performed, the encoding device can zero out a second region of the current block where no modified conversion coefficients exist, and derive a second variable indicating whether the conversion coefficients exist in the second region.
[0514] As shown in Figures 11 and 12, the remaining area of the current block where no corrected conversion coefficients exist can all be processed as 0. This zeroing out reduces the computational load required to execute the entire conversion process, thereby reducing the power consumption required for the conversion. In addition, the latency associated with the conversion process can be reduced, increasing the efficiency of video coding.
[0515] The second variable is LfnstZeroOutSigCoeffFlag, which indicates that zero-out has been performed when LFNST is applied. The second variable is initially set to 1, and can also be changed to 0 if an effective coefficient exists in the second region.
[0516] The variable LfnstZeroOutSigCoeffFlag can be derived to 0 if the index of the subblock containing the last non-zero coefficient is greater than 0, the width and height of the transformation block are both 4 or greater, or the last position of the non-zero coefficient within the subblock containing the last non-zero coefficient is greater than 7, or the size of the transformation block is 4x4 or 8x8.
[0517] In other words, if a non-zero coefficient is derived in a region other than the upper-leftmost region where an LFNST conversion coefficient can exist in the conversion block, or if a non-zero coefficient exists in a region other than the 8th position in the scan order for 4x4 and 8x8 blocks, the variable LfnstZeroOutSigCoeffFlag is set to 0.
[0518] The explanation for the first area and the zero-out when ISP applies are the same as the explanation for the decoding method, so redundant explanations are omitted.
[0519] For example, an encoding device can perform a zero-out when applying MTS to the primary transformation of the current block. The encoding device can perform a zero-out by filling the area excluding the top-left 16x16 region of the current block or subpartition block with zeros, and can encode the MTS index with a third variable indicating whether a transformation coefficient exists in the zero-out region.
[0520] The third variable is MtsZeroOutSigCoeffFlag, which indicates whether zero-out has been performed when MTS is applied. The variable MtsZeroOutSigCoeffFlag indicates whether a conversion coefficient exists in an area other than the upper left corner region where the last effective coefficient can exist due to zero-out after MTS execution, i.e., the upper left corner 16x16 area. It is initially set to 1, and if a conversion coefficient exists in an area other than the 16x16 area, its value can be changed from 1 to 0. If the value of the third variable is 0, the MTS index is not encoded or signaled.
[0521] The encoding device can configure and output video information such that at least one of the LFNST index and the MTS index is signaled at the coding unit level, and the MTS index is signaled immediately after the LFNST index is signaled (S2050).
[0522] Furthermore, the encoding device can perform quantization based on the conversion coefficients for the current block or modified conversion coefficients to derive quantized conversion coefficients, and encode and output video information including information about the quantized conversion coefficients.
[0523] The encoding device can generate residual information containing information about the quantized conversion coefficients. This residual information may include the aforementioned conversion-related information / syntax elements. The encoding device can encode the video information containing this residual information and output it in bitstream format.
[0524] More specifically, the encoding device can generate information about the quantized transformation coefficients and encode the generated information about the quantized transformation coefficients.
[0525] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficient may be called a transformation coefficient. If transformation / inverse transformation is omitted, the transformation coefficient may also be called a coefficient or residual coefficient, or may still be called a transformation coefficient for consistency of expression.
[0526] Furthermore, in this document, quantized transformation coefficients and transformation coefficients may be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients, and such information about the transformation coefficients may be signaled via residual coding syntax. Based on the residual information (or information about the transformation coefficients), transformation coefficients can be derived, and scaled transformation coefficients can be derived via inverse transformation (scaling) of the transformation coefficients. Based on inverse transformation (transformation) of the scaled transformation coefficients, residual samples can be derived. This can be similarly applied / expressed in other parts of this document.
[0527] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks; however, this document is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of this document.
[0528] The method described in this document above can be implemented in software form, and the encoding and / or decoding device described in this document may be included in, for example, image processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.
[0529] In this document, when embodiments are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0530] Furthermore, decoding and encoding devices to which this document applies may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction equipment, real-time communication equipment such as video communications, mobile streaming equipment, storage media, camcorders, customized video (VoD) service providers, OTT video (Over the Top Video) equipment, internet streaming service providers, 3D video equipment, image telephone video equipment, and medical video equipment, and may be used to process video signals or data signals. For example, OTT video (Over the Top Video) equipment may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.
[0531] Furthermore, the processing methods to which this document applies can be produced in the form of programs executed by a computer and stored on a computer-readable recording medium. Multimedia data having the data structure relating to this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), general-purpose serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored on computer-readable recording media or transmitted over a wireless network. Furthermore, embodiments of this document can be embodied in computer program products in the form of program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.
[0532] Figure 21 schematically shows an example of a video / image coding system to which this document can be applied.
[0533] Referring to Figure 21, a video / image coding system may include a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device in the form of a file or streaming via a digital storage medium or network.
[0534] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may consist of a separate device or external component.
[0535] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.
[0536] An encoding device can encode input video / images. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0537] The transmitting unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0538] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of the encoding device.
[0539] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0540] Figure 22 illustrates the structure of a content streaming system to which this document applies.
[0541] Furthermore, the content streaming system to which this document applies may broadly include encoding servers, streaming servers, web servers, media storage, user devices, and multimedia input devices.
[0542] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting it to the streaming server. As an alternative example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted. The bitstream can be generated by the encoding method or bitstream generation method to which this document applies, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving it.
[0543] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, which then transmits multimedia data to the user. The content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.
[0544] The streaming server can receive content from media storage and / or encoding servers. For example, if it receives content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0545] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage. Each server in the content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.
[0546] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined to embody an apparatus, and the technical features of the apparatus claims herein can be combined to embody a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody a method.
Claims
1. In a decoding device related to video decoding, Memory and The system comprises at least one processor connected to the memory, The aforementioned at least one processor is Obtain residual information at the residual coding level, Based on the aforementioned residual information, a conversion coefficient for the current block is derived. Residual samples are derived by applying at least one of the LFNST kernel derived from the LFNST (low frequency non-separable transform) index or the linear transformation kernel derived from the MTS (multiple transform selection) index to the transformation coefficients. It is configured to generate a restored picture based on the said residual sample, The LFNST index and the MTS index are obtained at the coding unit level. The MTS index is obtained immediately after the LFNST index is obtained. A decoding device in which the MTS index is obtained based on the fact that at least the value of the LFNST index is equal to 0.
2. The decoding apparatus according to claim 1, wherein, if the tree type of the current block is a single-tree type, the LFNST index is parsed after residual coding has been performed on the luma block and chroma block of the current block.
3. The decoding apparatus according to claim 1, wherein, if the tree type of the current block is a dual tree type and a chroma component is coded, the LFNST index is parsed after residual coding is performed on the Cb and Cr components of the chroma block.
4. The decoding apparatus according to claim 1, wherein, if the current block is divided into a plurality of subpartition blocks, the LFNST index is parsed after residual coding is performed on the plurality of subpartition blocks.
5. The decoding apparatus according to claim 4, wherein, when the current block is divided into a plurality of subpartition blocks, the LFNST index is parsed regardless of whether the conversion coefficient exists in the region excluding the DC position of each of the plurality of subpartition blocks.
6. The execution of the aforementioned residual coding is To derive a first variable related to whether the conversion coefficient exists in the region excluding the DC position of the current block, This includes deriving a second variable related to whether the conversion coefficient exists in a second region excluding the first region at the upper left of the current block or the subpartition block from which the current block is divided, The decoding device according to claim 4, wherein the LFNST index is parsed if the conversion coefficient exists in the region excluding the DC position and the conversion coefficient does not exist in the second region.
7. The execution of residual coding includes deriving a third variable related to whether the conversion coefficient exists in the region excluding the upper left 16x16 region of the current block, The decoding device according to claim 1, wherein if the conversion coefficient does not exist in the region excluding the 16x16 region, the MTS index is parsed.
8. In an encoding device related to video encoding, Memory and The system comprises at least one processor connected to the memory, The aforementioned at least one processor is Currently, we derive the predicted sample for the block. Based on the predicted sample, a residual sample is derived for the current block. A conversion coefficient derivation operation is performed to derive conversion coefficients for the current block by applying at least one of the LFNST (low frequency non-separable transform) kernel or a linear transformation kernel to the residual sample, Encode at least one of the LFNST index associated with the LFNST kernel or the MTS index associated with the primary transformation kernel, Video information, The LFNST index and the MTS index are signaled at the coding unit level. The MTS index is signaled immediately after the LFNST index is signaled, and An encoding device configured to signal and output the MTS index based on at least the value of the LFNST index being equal to 0.
9. The encoding device according to claim 8, wherein, if the tree type of the current block is a single-tree type, the LFNST index is encoded after the conversion coefficient derivation operation for the luma block and chroma block of the current block has been performed.
10. The encoding apparatus according to claim 8, wherein, when the tree type of the current block is a dual tree type and a chroma component is coded, the LFNST index is encoded after the conversion coefficient derivation operation for the Cb and Cr components of the chroma block is performed.
11. The encoding device according to claim 8, wherein, if the current block is divided into a plurality of subpartition blocks, the LFNST index is encoded after the conversion coefficient derivation operation is performed for the plurality of subpartition blocks.
12. The encoding device according to claim 11, wherein, when the current block is divided into a plurality of subpartition blocks, the LFNST index is encoded regardless of whether the conversion coefficient exists in the region excluding each DC position of the subpartition block.
13. The execution of the conversion coefficient derivation operation is as follows: To derive a first variable related to whether the conversion coefficient exists in the region excluding the DC position of the current block, This includes deriving a second variable related to whether the conversion coefficient exists in a second region excluding the first region at the upper left of the current block or the subpartition block from which the current block is divided, The encoding device according to claim 11, wherein the conversion coefficient exists in the region excluding the DC position and the conversion coefficient does not exist in the second region, and the LFNST index is encoded.
14. The execution of the conversion coefficient derivation operation includes deriving a third variable related to whether the conversion coefficient exists in the region excluding the upper left 16x16 region of the current block, The encoding device according to claim 8, wherein if the conversion coefficient does not exist in the region excluding the 16x16 region, the MTS index is encoded.
15. In a device for transmitting video data, At least one processor configured to acquire a bitstream, wherein the bitstream is The current step is to derive a predicted sample for the block, The steps include: deriving a residual sample for the current block based on the predicted sample; The steps include performing a transformation coefficient derivation operation to derive transformation coefficients for the current block by applying at least one of an LFNST (low frequency non-separable transform) kernel or a linear transformation kernel to the residual sample, A step of encoding at least one of the LFNST index associated with the LFNST kernel or the MTS index associated with the primary transformation kernel, Video information, The LFNST index and the MTS index are signaled at the coding unit level. The MTS index is signaled immediately after the LFNST index is signaled, and A processor generated by the steps of: configuring and outputting such that the MTS index is signaled based on the value of the LFNST index being equal to at least 0; A device comprising: a transmitting unit configured to transmit the data including the bitstream of the video information.