Image coding method and apparatus based on conversion
The image coding method enhances compression efficiency by optimizing quantization through LFNST-based scaling list application, addressing the need for efficient compression of high-resolution images and VR/AR content.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-02-18
- Publication Date
- 2026-06-02
AI Technical Summary
The increasing demand for high-resolution and high-quality images/videos, including VR and AR content, necessitates a highly efficient image/video compression technology to reduce transmission and storage costs.
An image coding method and apparatus that employs LFNST (Lifted Fractional Transform) to determine whether a scaling list is applied to a current block based on flag information and LFNST index, optimizing quantization efficiency for luma and chroma elements.
Improves overall image/video compression efficiency and quantization efficiency, particularly for chroma components in single-tree types.
Smart Images

Figure 0007869355000032 
Figure 0007869355000033 
Figure 0007869355000034
Abstract
Description
Technical Field
[0001] This document relates to image coding technology, and more particularly, to an image coding method and apparatus based on transform in an image coding system.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) contents, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from those of real images, such as game images, has been increasing.
[0004] Therefore, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The technical problem of this document is to provide a method and apparatus for increasing image coding efficiency.
[0006] Another technical problem of this document is to provide a method and apparatus for increasing quantization efficiency.
[0007] Another technical challenge of this paper is to provide a method and apparatus for improving the quantization efficiency of chroma elements in a single-tree type. [Means for solving the problem]
[0008] According to one embodiment of this document, an image decoding method is provided that is performed by a decoding device. The method includes the steps of: receiving flag information indicating the availability of a scaling list, an LFNST index for the current block, and residual information when LFNST is performed; determining whether the scaling list is applied to the current block based on the flag information, the LFNST index, and the tree type of the current block; deriving a conversion coefficient for the current block from the residual information based on the determination result; and deriving a modified conversion coefficient by applying LFNST to the conversion coefficient, wherein if the tree type of the current block is a single tree and it is a luma element, the scaling list is not applied, and if the tree type of the current block is a single tree and it is a chroma element, the scaling list can be applied.
[0009] The LFNST may not be applied to the chroma element of the current block.
[0010] If the aforementioned flag information indicates that the scaling list is not available and the LFNST index is greater than 0, the scaling list may not be applied to the Luma element.
[0011] If the flag information indicates that the scaling list is not available and the LFNST index is greater than 0, the scaling list may not be applied to the chroma elements if the tree type of the current block is a dual-tree chroma.
[0012] If the flag information indicates that the scaling list is not available and the LFNST index is greater than 0, the scaling list may not be applied to the luma elements if the tree type of the current block is a dual-tree luma.
[0013] The aforementioned current block may contain a transformation block.
[0014] According to one embodiment of this document, an image encoding method is provided that is performed by an encoding device. The method includes the steps of deriving conversion coefficients for the current block from the residual samples based on a conversion process, and quantizing the conversion coefficients based on a scaling list, wherein whether the scaling list is applied to the current block is determined based on whether LFNST is performed on the conversion process and the tree type of the current block, wherein if the tree type of the current block is a single tree and contains chroma elements, the scaling list is not applied, and if the tree type of the current block is a single tree and contains chroma elements, the scaling list can be applied.
[0015] According to another embodiment of this document, a digital storage medium is provided which stores image data containing encoded image information and a bitstream generated by an image encoding method performed by an encoding device.
[0016] According to another embodiment of this document, a digital storage medium is provided which stores image data containing encoded image information and a bitstream for a decoding device to perform the image decoding method. [Effects of the Invention]
[0017] According to this document, it is possible to improve the overall image / video compression efficiency.
[0018] According to this document, the quantization efficiency can be improved.
[0019] According to this document, the quantization efficiency for chroma components in single-tree type can be improved.
[0020] The effects obtained through a specific example in this specification are not limited to the effects listed above. For example, there may be various technical effects that a person having ordinary skill in the related art can understand or derive from this specification. Accordingly, the specific effects of this specification are not limited to those explicitly described in this specification, and may include various effects that can be understood or derived from the technical features of this specification.
Brief Description of Drawings
[0021] [Figure 1] It is a diagram schematically explaining the configuration of a video / image encoding device to which this document can be applied. [Figure 2] It is a diagram schematically explaining the configuration of a video / image decoding device to which this document can be applied. [Figure 3] It schematically shows a multiple conversion technique according to an embodiment of this document. [Figure 4] Exemplarily shows the intra prediction direction mode of 65 prediction directions. [Figure 5] It is a diagram for explaining the RST according to an embodiment of this document. [Figure 6] It is a diagram showing the order of arranging the output data of forward first-order conversion in a one-dimensional vector by way of an example. [Figure 7] It is a diagram showing the order of arranging the output data of forward second-order conversion in a two-dimensional block by way of an example. [Figure 8] It is a diagram showing the block shape to which LFNST is applied. [Figure 9] It is a diagram showing the arrangement of the output data of forward LFNST by way of an example. [Figure 10]A diagram showing zeroing out in a block to which 4x4 LFNST is applied according to one example. [Figure 11] A diagram showing zeroing out in a block to which 8x8 LFNST is applied according to one example. [Figure 12] A diagram for explaining a method of decoding an image according to one example. [Figure 13] A diagram for explaining a method of encoding an image according to one example. [Figure 14] An example of a video / image coding system to which this document can be applied is schematically shown. [Figure 15] A structural diagram of a content streaming system to which this document is applied is exemplarily shown.
Embodiments for Carrying Out the Invention
[0022] This document can be modified in various ways and can have various embodiments, but specific embodiments are illustrated in the drawings and will be described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not intended to limit the technical idea in this document. Singular expressions include plural expressions unless the context clearly indicates a different meaning. In this specification, terms such as "including" or "having" are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude in advance the possibility of the existence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.
[0023] On the other hand, each configuration shown in the diagrams described in this document is shown independently for the convenience of explaining its distinct characteristic functions, and does not mean that each configuration is implemented with separate hardware or separate software. For example, two or more configurations may be combined to form one configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the rights of this document, as long as they do not deviate from the essence of this document.
[0024] The following describes preferred embodiments of this document in more detail with reference to the attached figures. The same reference numerals are used for the same components in the drawings, and redundant descriptions of the same components are omitted.
[0025] This document relates to video / image coding. For example, the methods / examples disclosed in this document may be related to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), next-generation video / image coding standards after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), EVC (essential video coding) standard, AVS2 standard, etc.).
[0026] This document presents various embodiments relating to video / image coding, and unless otherwise noted, these embodiments may be performed in combination with each other.
[0027] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a single image representing a specific time period, while "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more CTUs (coding tree units). A single picture can consist of one or more slices or tiles. A single picture can consist of one or more tile groups. A tile group can contain one or more tiles.
[0028] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" can be used as a counterpart to pixel. A sample generally refers to a pixel or a pixel value, sometimes only the luma component pixel / pixel value, or sometimes only the chroma component pixel / pixel value. Alternatively, a sample may refer to a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it may refer to the conversion coefficient in the frequency domain.
[0029] A unit can represent a basic unit of image processing. A unit can include at least one specific region of a picture and information about that region. A unit can include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area, as appropriate. In general, an M×N block can include a sample (or sample array) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0030] In this document, the terms " / " and "," should be interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Furthermore, "A / B / C" means "at least one of A, B, and / or C." Similarly, "A, B, C" also means "at least one of A, B, and / or C."
[0031] Furthermore, in this document, "or" is interpreted as "and / or." For example, "A or B" may mean 1) only "A," 2) only "B," or 3) both "A and B." In other words, "or" in this document may mean "additionally or alternatively."
[0032] In this specification, "at least one of A and B" may mean "just A," "just B," or "both A and B." Furthermore, in this specification, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."
[0033] Furthermore, in this specification, "at least one of A, B and C" may mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."
[0034] Furthermore, parentheses used herein may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" as used herein is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."
[0035] Technical features described individually in each drawing in this specification may be implemented individually or simultaneously.
[0036] Figure 1 is a schematic diagram illustrating the configuration of a video / image encoding device to which this document applies. Hereinafter, the term "video encoding device" may include an image encoding device.
[0037] Referring to Figure 1, the encoding device 100 may be configured to include an image partitioner 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 may include an inter-predictor 121 and an intra-predictor 122. The residual processor 130 may include a transformer 132, a quantizer 133, a dequantizer 134, and an inverse transformer 135. The residual processor 130 may further include a subtractor 131. The adder 150 may be called a reconstructor or a reconstructed block generator. The aforementioned image segmentation unit 110, prediction unit 120, residual processing unit 130, entropy encoding unit 140, addition unit 150, and filtering unit 160 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 170 as an internal / external component.
[0038] The image splitting unit 110 can split an input image (or picture, frame) input to the encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary structure. Alternatively, the binary-tree structure may be applied first. Based on the final coding unit that cannot be further split, the coding procedure described in this document may be performed. In this case, based on coding efficiency due to image characteristics, the largest coding unit can be immediately used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0039] The term "unit" can be used interchangeably with terms such as "block" or "area," depending on the context. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.
[0040] The encoding device 100 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the inter-prediction unit 121 or intra-prediction unit 122 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 132. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoding device 100 can be called the subtraction unit 131. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes the predicted sample for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit 140, as will be described later in the explanation of each prediction mode. Information regarding the prediction can be encoded by the entropy encoding unit 140 and output in bitstream format.
[0041] The intra-prediction unit 122 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) of the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 122 can also determine the prediction mode to apply to the current block using the prediction modes applied to the surrounding blocks.
[0042] The interprediction unit 121 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between the surrounding block and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding block may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporal neighboring block may also be called a collocated picture (colPic). For example, the interpretation unit 121 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 121 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0043] The prediction unit 120 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values within the picture can be signaled based on information about the palette table and palette index.
[0044] The prediction signal generated via the prediction unit (including the inter-prediction unit 121 and / or the intra-prediction unit 122) can be used to generate a reconstructed signal or a residual signal. The transformation unit 132 can generate transformation coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and obtaining a transformation based on it. The transformation process can be applied to pixel blocks of the same size that are square, or to blocks of variable size that are not square.
[0045] The quantization unit 133 quantizes the conversion coefficients and transmits them to the entropy encoding unit 140, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 133 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 140 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 140 can also encode information necessary for video / image restoration (e.g., the values of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of network abstraction layer (NAL). The video / image information may further include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device in this document may be included in the video / image information. The video / image information may be encoded via the encoding procedure described above and included in the bitstream.The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 140 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 100, or the transmitting unit may be included in the entropy encoding unit 140.
[0046] The quantized conversion coefficients output from the quantization unit 133 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 134 and the inverse transformation unit 135. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 121 or the intra-prediction unit 122. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 150 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.
[0047] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.
[0048] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. As will be described later in the explanation of each filtering method, the filtering unit 160 can generate various filtering-related information and transmit it to the entropy encoding unit 140. The filtering-related information can be encoded by the entropy encoding unit 140 and output in the form of a bitstream.
[0049] The corrected restored picture sent to memory 170 can be used as a reference picture in the interpretation unit 121. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.
[0050] The DPB in memory 170 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 121. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the already restored picture. The stored motion information can be transmitted to the inter-prediction unit 121 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 122.
[0051] Figure 2 is a schematic diagram illustrating the configuration of a video / image decoding device to which this document applies.
[0052] Referring to Figure 2, the decoding device 200 can be configured to include an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The predictor 230 may include an inter-predictor 231 and an intra-predictor 232. The residual processor 220 may include a dequantizer 221 and an inverse transformer 222. The aforementioned entropy decoder 210, residual processor 220, predictor 230, adder 240, and filtering device 250 can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 260 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware component may further include memory 260 as an internal / external component.
[0053] When a bitstream containing video / image information is input, the decoding device 200 can reconstruct the image corresponding to the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 200 can derive units / blocks based on block division information obtained from the bitstream. The decoding device 200 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing units may be, for example, coding units, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed image signal decoded and output via the decoding device 200 can then be reproduced via a playback device.
[0054] The decoding device 200 can receive the signal output from the encoding device shown in Figure 1 in bitstream form, and the received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image reconstruction, quantized values of conversion coefficients related to residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the information of the syntax element to be decoded, the decoded information of the surrounding and decoded blocks, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction unit (inter-prediction unit 232 and intra-prediction unit 231), and the residual values that have been entropy decoded by the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 220. The residual processing unit 220 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 210, information related to filtering can be provided to the filtering unit 250. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 200, or the receiving unit can be a component of the entropy decoding unit 210. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 210, and the sample decoder may include at least one of the inverse quantization unit 221, inverse transformation unit 222, addition unit 240, filtering unit 250, memory 260, inter-prediction unit 232, and intra-prediction unit 231.
[0055] The inverse quantization unit 221 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 221 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 221 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) to obtain the transformation coefficients.
[0056] In the inverse conversion unit 222, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0057] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.
[0058] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0059] The intra-prediction unit 231 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) of the current block or at a distance from it, depending on the prediction mode. The prediction mode in intra-prediction can include a plurality of non-directional modes and a plurality of directional modes. The intra-prediction unit 231 can also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0060] The interprediction unit 232 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. To reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In interprediction, neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the interprediction unit 232 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.
[0061] The summing unit 240 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 232 and / or intra-prediction unit 231). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.
[0062] The summing unit 240 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and may be output after filtering, as described later, or may be used for intra-prediction of the next picture.
[0063] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0064] The filtering unit 250 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 250 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 260, specifically to the DPB of the memory 260. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0065] The (modified) restored picture stored in the DPB of memory 260 can be used as a reference picture by the inter-prediction unit 232. Memory 260 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 232 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 260 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 231.
[0066] In this document, the embodiments described for the filtering unit 160, inter-prediction unit 121, and intra-prediction unit 122 of the encoding device 100 can also be applied identically or in a corresponding manner to the filtering unit 250, inter-prediction unit 232, and intra-prediction unit 231 of the decoding device 200, respectively.
[0067] As described above, prediction is performed during video coding to improve compression efficiency. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived identically by the encoding and decoding devices, and the encoding device can improve image coding efficiency by signaling the decoding device information about the residual between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can generate a restored block containing restored samples by combining the residual block and the predicted block, and can generate a restored picture containing the restored block.
[0068] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, perform a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, perform a quantization procedure on the transformation coefficients to derive quantized transformation coefficients, and signal the associated residual information (via a bitstream) to a decoding device. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can perform an inverse quantization / inverse transformation procedure based on the residual information to derive a residual sample (or residual block). The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can further derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.
[0069] Figure 3 schematically illustrates the multiplexing technique used in this document.
[0070] Referring to Figure 3, the conversion unit may correspond to the conversion unit in the encoding device shown in Figure 1, and the inverse conversion unit may correspond to the inverse conversion unit in the encoding device shown in Figure 1 or the inverse conversion unit in the decoding device shown in Figure 3.
[0071] The transformation unit can perform a primary transformation based on the residual sample (residual sample array) in the residual block to derive (primary) transformation coefficients (S310). Such a primary transformation may be referred to as a core transformation. Here, the primary transformation is obtained based on Multiple Transform Selection (MTS), and when a multiple transformation is applied as the primary transformation, it may be referred to as a multiple core transformation.
[0072] Multiple core transforms can represent a method of transformation that further uses a Discrete Cosine Transform (DCT) type 2, a Discrete Sine Transform (DST) type 7, a Discrete Cosine Transform (DST) type 8, and / or a DST type 1. That is, the multiple core transforms can represent a method of transformation that converts a spatial domain residual signal (or residual block) into frequency domain transformation coefficients (or first-order transformation coefficients) based on a plurality of transformation kernels selected from the DCT type 2, DST type 7, DCT type 8, and DST type 1. Here, the first-order transformation coefficients may be called provisional transformation coefficients from the perspective of the transformer.
[0073] In other words, when an existing transformation method is applied, a spatial-domain to frequency-domain transformation of a residual signal (or residual block) can be applied based on DCT type 2 to generate transformation coefficients. In contrast, when the multiple core transformation is applied, a spatial-domain to frequency-domain transformation of a residual signal (or residual block) can be applied based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate transformation coefficients (or first-order transformation coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc. may be called transformation types, transformation kernels, or transformation cores. Such DCT / DST transformation types can be defined based on basis functions.
[0074] When the multi-core transformation is performed, a vertical transformation kernel and a horizontal transformation kernel can be selected from the transformation kernels for the target block, and a vertical transformation can be performed on the target block based on the vertical transformation kernel, and a horizontal transformation can be performed on the target block based on the horizontal transformation kernel. Here, the horizontal transformation may represent a transformation to the horizontal component of the target block, and the vertical transformation may represent a transformation to the vertical component of the target block. The vertical transformation kernel / horizontal transformation kernel can be adaptively determined based on the prediction mode and / or transformation index of the target block (CU or subblock) including the residual block.
[0075] Furthermore, for example, when applying MTS to perform a linear transformation, a mapping relationship to the transformation kernel can be established by setting a specific basis function to a predetermined value and combining it with whether or not a particular basis function is applied when it is a vertical or horizontal transformation. For example, if the horizontal transformation kernel is represented by trTypeHor and the vertical transformation kernel is represented by trTypeVer, then a value of 0 for trTypeHor or trTypeVer can be set to DCT2, a value of 1 for trTypeHor or trTypeVer can be set to DST7, and a value of 2 for trTypeHor or trTypeVer can be set to DCT8.
[0076] In this case, MTS index information can be encoded and signaled to a decoder to indicate one of a number of conversion kernel sets. For example, an MTS index of 0 can indicate that the values of trTypeHor and trTypeVer are all 0; an MTS index of 1 can indicate that the values of trTypeHor and trTypeVer are all 1; an MTS index of 2 can indicate that the value of trTypeHor is 2 and the value of trTypeVer is 1; an MTS index of 3 can indicate that the value of trTypeHor is 1 and the value of trTypeVer is 2; and an MTS index of 4 can indicate that the values of trTypeHor and trTypeVer are all 2.
[0077] As an example, the conversion kernel sets based on MTS index information are shown in the table below.
[0078] [Table 1]
[0079] The conversion unit performs a quadratic transformation based on the (primary) transformation coefficients to derive modified (secondary) transformation coefficients (S320). The primary transformation is a transformation from the spatial domain to the frequency domain, and the secondary transformation means transforming into a more compressed representation by utilizing the correlations that exist between the (primary) transformation coefficients. The secondary transformation includes a non-separable transform. In this case, the secondary transformation may be called a non-separable secondary transform (NSST) or MDNSST (mode-dependent non-separable secondary transform). The non-separable secondary transform represents a transformation that generates modified transformation coefficients (or secondary transformation coefficients) for the residual signal by performing a quadratic transformation on the (primary) transformation coefficients derived by the primary transformation based on a non-separable transform matrix. Here, the transformation can be applied at once to the (primary) transformation coefficients based on the non-separable transform matrix without separating the vertical and horizontal transformations (or applying the horizontal and vertical transformations independently). In other words, the non-separable quadratic transformation is not applied separately to the (primary) transformation coefficients in the vertical and horizontal directions, but rather represents a transformation method that, for example, rearranges a two-dimensional signal (transformation coefficient) into a one-dimensional signal in a specific fixed direction (e.g., row-first or column-first direction), and then generates a modified transformation coefficient (or quadratic transformation coefficient) based on the non-separable transformation matrix. For example, row-first ordering means arranging the 1st row, 2nd row, ..., Nth row in a column for an M×N block, and column-first ordering means arranging the 1st column, 2nd column, ..., Mth column in a column for an M×N block. The non-separable quadratic transformation can be applied to the top-left region of a block composed of (primary) transformation coefficients (hereinafter referred to as a transformation coefficient block). For example, if both the width (W) and height (H) of the conversion coefficient block are 8 or greater, an 8x8 non-separable quadratic transformation can be applied to the upper left 8x8 region of the conversion coefficient block.Furthermore, if both the width (W) and height (H) of the conversion coefficient block are 4 or greater, but either the width (W) or height (H) of the conversion coefficient block is less than 8, the 4×4 unseparable quadratic transformation can be applied to the upper left min(8,W)×min(8,H) region of the conversion coefficient block. However, the embodiment is not limited to this, and for example, even if only the condition that both the width (W) or height (H) of the conversion coefficient block are 4 or greater is satisfied, the 4×4 unseparable quadratic transformation can also be applied to the upper left min(8,W)×min(8,H) region of the conversion coefficient block.
[0080] Specifically, for example, if a 4x4 input block is used, the unseparated quadratic transform can be performed as follows:
[0081] The aforementioned 4x4 input block X can be represented as follows:
[0082]
number
[0083] When X is shown in the form of a vector, JPEG0007869355000003.jpg84 can be represented as follows:
[0084]
number
[0085] As shown in equation 2, vector JPEG0007869355000005.jpg84 rearranges the 2D block of X in equation 1 into a 1D vector in row-first order.
[0086] In this case, the quadratic inseparable transform can be calculated as follows:
[0087]
number
[0088] Here, JPEG0007869355000007.jpg75 shows the transformation coefficient vector, and T shows a 16x16 (non-separable) transformation matrix.
[0089] Through the above equation 3, a 16 × 1 transformation coefficient vector JPEG0007869355000008.jpg75 can be derived, and the above JPEG0007869355000009.jpg75 can be reorganized into 4x4 blocks via the scan order (horizontal, vertical, diagonal, etc.). However, the calculation described above is illustrative, and to reduce the computational complexity of the non-separable quadratic transform, HyGT (Hypercube-Givens Transform), etc., can also be used for the calculation of the non-separable quadratic transform.
[0090] On the other hand, the unseparated quadratic transform allows the transform kernel (or transform core, transform type) to be selected as mode-dependent. Here, the mode may include intra-predictive mode and / or inter-predictive mode.
[0091] As described above, the non-separable quadratic transformation can be performed based on an 8x8 transformation or a 4x4 transformation determined based on the width (W) and height (H) of the transformation coefficient block. An 8x8 transformation refers to a transformation that can be applied to an 8x8 region contained within the transformation coefficient block when W and H are all equal to or greater than 8, and this 8x8 region may be the upper left 8x8 region inside the transformation coefficient block. Similarly, a 4x4 transformation refers to a transformation that can be applied to a 4x4 region contained within the transformation coefficient block when W and H are all equal to or greater than 4, and this 4x4 region may be the upper left 4x4 region inside the transformation coefficient block. For example, an 8x8 transformation kernel matrix may be a 64x64 / 16x64 matrix, and a 4x4 transformation kernel matrix may be a 16x16 / 8x16 matrix.
[0092] In this case, for the selection of mode-based conversion kernels, two non-separable quadratic conversion kernels may be configured for each conversion set for non-separable quadratic conversions for both 8x8 and 4x4 conversions, and there may be four conversion sets. That is, four conversion sets may be configured for 8x8 conversions and four conversion sets may be configured for 4x4 conversions. In this case, the four conversion sets for 8x8 conversions may each contain two 8x8 conversion kernels, and in this case, the four conversion sets for 4x4 conversions may each contain two 4x4 conversion kernels.
[0093] However, the size of the transformation, i.e., the size of the region to which the transformation is applied, may be a size other than 8x8 or 4x4 as an example, the number of sets may be n, and the number of transformation kernels in each set may be k.
[0094] The aforementioned transformation set may be called an NSST set or an LFNST set. The selection of a particular set from among the transformation sets can be performed, for example, based on the intra-prediction mode of the current block (CU or sub-block). LFNST (Low-Frequency Non-Separable Transform) may be an example of a reduced non-separable transform described later, and represents a non-separable transform for low-frequency components.
[0095] For reference, for example, an intra-prediction mode may include two non-directinoal (or non-angular) intra-prediction modes and 65 directional (or angular) intra-prediction modes. The non-directinoal intra-prediction mode may include a planar intra-prediction mode (number 0) and a DC intra-prediction mode (number 1), and the directional intra-prediction mode may include 65 intra-prediction modes (numbers 2 through 66). However, this is illustrative, and this document can also be applied when the number of intra-prediction modes differs. On the other hand, a 67th intra-prediction mode may be used as needed, and this 67th intra-prediction mode may represent a linear model (LM) mode.
[0096] Figure 4 illustrates 65 intradirectional modes for predicting directions.
[0097] Referring to Figure 4, we can distinguish between intra-prediction modes with horizontal directionality and intra-prediction modes with vertical directionality, centered around intra-prediction mode 34, which has a prediction direction on the lower right diagonal. In Figure 4, H and V represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate a displacement of 1 / 32 units on the sample grid position. This can be used to indicate an offset relative to the mode index value. Intra-prediction modes 2 through 33 are horizontally oriented, while intra-prediction modes 34 through 66 are vertically oriented. On the other hand, intra-prediction mode 34 can be seen as neither strictly horizontal nor vertical, but from the perspective of determining the transformation set of the quadratic transformation, it can be classified as belonging to the horizontal directionality. This is because the input data is transposed for the vertical modes which are symmetrical to intra-prediction mode 34, and the input data alignment method for the horizontal modes is used for intra-prediction mode 34. Transposing input data means that for a 2D block of data MxN, rows become columns and columns become rows, resulting in NxM data. Intra-prediction modes 18 and 50 represent the horizontal intra-prediction mode and the vertical intra-prediction mode, respectively. Intra-prediction mode 2 predicts the upper right direction using the left reference pixel, so it can be called the upper right diagonal intra-prediction mode. In the same context, intra-prediction mode 34 can be called the lower right diagonal intra-prediction mode, and intra-prediction mode 66 can be called the lower left diagonal intra-prediction mode.
[0098] For example, the mapping of four transformation sets in intra-predictive mode may be shown, for instance, as in the following table.
[0099] [Table 2]
[0100] As shown in Table 2, the intra prediction mode allows mapping to one of four transformation sets, i.e., lfnstTrSetIdx to any of the four values from 0 to 3.
[0101] On the other hand, once it is determined that a specific set is to be used for an inseparable transformation, one of the k transformation kernels within that specific set can be selected via the inseparable quadratic transformation index. The encoding device can derive an inseparable quadratic transformation index that points to a specific transformation kernel based on an RD (rate-distortion) check, and can signal the decoding device to the inseparable quadratic transformation index. The decoding device can select one of the k transformation kernels within the specific set based on the inseparable quadratic transformation index. For example, index value 0 of lfnst can point to the first inseparable quadratic transformation kernel, index value 1 of lfnst can point to the second inseparable quadratic transformation kernel, index value 2 of lfnst can point to the third inseparable quadratic transformation kernel. Alternatively, index value 0 of lfnst can indicate that the first inseparable quadratic transformation is not applied to the target block, and index values 1 to 3 of lfnst can point to the three transformation kernels.
[0102] The transformation unit can perform the non-separable quadratic transformation based on the selected transformation kernel to obtain the modified (quadratic) transformation coefficients. The modified transformation coefficients can be derived from the quantized transformation coefficients via the quantization unit as described above, encoded, and transmitted to the decoder for signaling and to the inverse quantization / inverse transformation unit within the encoding unit.
[0103] On the other hand, as mentioned above, if the quadratic transformation is omitted, the (primary) transformation coefficients, which are the output of the primary (separated) transformation, can be derived as quantized transformation coefficients via the quantization unit as described above, encoded, and transmitted to the decoder for signaling and to the inverse quantization / inverse transformation unit within the encoding unit.
[0104] The inverse transform unit can perform a series of steps in the reverse order of the steps performed by the transform unit described above. The inverse transform unit can receive the (inversely quantized) transform coefficients, perform a quadratic (inverse) transform to derive the (primary) transform coefficients (S350), and perform a primary (inverse) transform on the (primary) transform coefficients to obtain a residual block (residual sample) (S360). Here, the primary transform coefficients may be called modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding and decoding devices can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.
[0105] On the other hand, the decoding device may further include a quadratic inverse transform applicability determination unit (or an element that determines whether a quadratic inverse transform is applicable) and a quadratic inverse transform determination unit (or an element that determines a quadratic inverse transform). The quadratic inverse transform applicability determination unit can determine whether a quadratic inverse transform is applicable. For example, the quadratic inverse transform may be NSST, RST, or LFNST, and the quadratic inverse transform applicability determination unit can determine whether a quadratic inverse transform is applicable based on a quadratic transform flag parsed from the bitstream. As another example, the quadratic inverse transform applicability determination unit can also determine whether a quadratic inverse transform is applicable based on the transformation coefficients of the residual block.
[0106] The quadratic inverse transform determination unit can determine the quadratic inverse transform. At that time, the quadratic inverse transform determination unit can determine the quadratic inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified by the intra-prediction mode. Furthermore, as one embodiment, the quadratic transform determination method can be determined dependent on the linear transform determination method. Various combinations of linear and quadratic transforms can be determined by the intra-prediction mode. Also, as an example, the quadratic inverse transform determination unit can determine the region to which the quadratic inverse transform is applied based on the size of the current block.
[0107] On the other hand, as mentioned above, if the quadratic (inverse) transformation is omitted, the (inversely quantized) transformation coefficients can be received, and the primary (separated) inverse transformation can be performed to obtain a residual block (residual sample). As mentioned above, the encoding and decoding devices can generate a reconstructed block based on the residual block and the predicted block, and generate a reconstructed picture based on this.
[0108] On the other hand, in this document, in order to reduce the computational complexity and memory requirements associated with unseparable quadratic transforms, we can apply RST (reduced secondary transform), which is a reduced version of the NSST concept with a smaller transform matrix (kernel).
[0109] On the other hand, the transformation kernel, transformation matrix, and coefficients constituting the transformation kernel matrix described in this document—that is, kernel coefficients or matrix coefficients—can be represented in 8 bits. This may be a requirement for implementation in decoding and encoding devices, and it can reduce the memory requirements for storing the transformation kernel while resulting in a reasonably acceptable performance degradation compared to existing 9-bit or 10-bit representations. Furthermore, representing the kernel matrix in 8 bits allows for the use of smaller multipliers and may be more suitable for SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.
[0110] In this specification, RST can mean a transformation performed on a residual sample for a target block based on a transform matrix whose size has been reduced by a simplification factor. When a simplification transformation is performed, the amount of computation required during the transformation can be reduced due to the reduction in the size of the transform matrix. In other words, RST can be used to resolve the complexity issue that arises when transforming large blocks or during non-separable transformations.
[0111] RST can be referred to by a variety of terms, including reduced transform, reduced secondary transform, reduction transform, simplified transform, and simple transform, and the names to which RST can be referred are not limited to those given. Alternatively, since RST mainly occurs in the low-frequency region with non-zero coefficients in the transform block, it may also be referred to as LFNST (Low-Frequency Non-Separable Transform). The aforementioned transform index may be named the LFNST index.
[0112] On the other hand, when the quadratic inverse transform is performed based on RST, the inverse transform unit 135 of the encoding device 100 and the inverse transform unit 222 of the decoding device 200 may each include an inverse RST unit that derives corrected transform coefficients based on the inverse RST of the transform coefficients, and an inverse linear transform unit that derives the residual sample for the target block based on the inverse linear transform of the corrected transform coefficients. The inverse linear transform means the inverse transform of the linear transform that was applied to the residual. In this document, deriving transform coefficients based on a transform means deriving transform coefficients by applying the transform.
[0113] Figure 5 is a diagram illustrating an RST according to one embodiment of this document.
[0114] In this specification, “target block” may mean the current block, residual block, or transformed block on which coding is performed.
[0115] In one embodiment of the RST, an N-dimensional vector is mapped to an R-dimensional vector located in a different space, and a reduced transformation matrix can be determined, where R is less than N. N can represent the square of the side length of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can represent the R / N value. The simplification factor can be referred to by various terms such as reduced factor, reduction factor, reduced factor, simplified factor, and simple factor. On the other hand, R can be referred to as the reduced coefficient, but in some cases the simplification factor may mean R. Also, in some cases the simplification factor may mean the N / R value.
[0116] In one embodiment, the simplification factor or simplification coefficient can be signaled via a bitstream, but the embodiment is not limited to this. For example, predefined values for the simplification factor or simplification coefficient may be stored in each encoding device 100 and decoding device 200, in which case the simplification factor or simplification coefficient may not be signaled separately.
[0117] The size of the simplified transformation matrix in one embodiment is RxN, which is smaller than the size NxN of a normal transformation matrix, and can be defined as shown in Equation 4 below.
[0118]
number
[0119] The matrix T in the Reduced Transform block shown in Figure 5(a) can be interpreted as the matrix TRxN in Equation 4. As shown in Figure 5(a), when the simplified transformation matrix TRxN is multiplied by the residual sample for the target block, the transformation coefficients for the target block can be derived.
[0120] In one embodiment, if the size of the block to which the transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), then the RST shown in Figure 5(a) can be expressed by a matrix operation as shown in Equation 5 below. In this case, memory and multiplication operations can be reduced to approximately 1 / 4 by the simplification factor.
[0121] In this text, matrix operations can be understood as operations that involve placing a matrix to the left of a column vector and multiplying the matrix by the column vector to obtain a new column vector.
[0122]
number
[0123] In Equation 5, r1 to r64 can represent residual samples for the target block, and more specifically, they can be transformation coefficients generated by applying a linear transformation. From the calculation result of Equation 5, the transformation coefficient ci for the target block can be derived, and the derivation process of ci is as shown in Equation 6.
[0124]
number
[0125] The result of the calculation in equation 6 allows us to derive the conversion coefficients c1 to cR for the target block. That is, if R=16, we can derive the conversion coefficients c1 to c16 for the target block. If a regular conversion were applied instead of RST, and a conversion matrix of size 64x64 (NxN) was multiplied by a residual sample of size 64x1 (Nx1), 64 (N) conversion coefficients for the target block might be derived. However, because RST was applied, only 16 (R) conversion coefficients for the target block are derived. The total number of conversion coefficients for the target block decreases from N to R, and the amount of data that the encoding device 100 sends to the decoding device 200 decreases, so the transmission efficiency between the encoding device 100 and the decoding device 200 may increase.
[0126] From the perspective of the size of the transformation matrix, the size of a normal transformation matrix is 64x64 (NxN), but the size of a simplified transformation matrix is reduced to 16x64 (RxN). Therefore, compared to performing a normal transformation, the memory usage when executing RST can be reduced by a ratio of R / N. In addition, compared to the number of multiplication operations NxN when using a normal transformation matrix, the number of multiplication operations can be reduced by a ratio of R / N (RxN) when using a simplified transformation matrix.
[0127] In one embodiment, the conversion unit 132 of the encoding device 100 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the target block's residual sample. These conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 200, and the inverse conversion unit 222 of the decoding device 200 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) on the conversion coefficients, and derive residual samples for the target block based on an inverse primary conversion on the modified conversion coefficients.
[0128] The size of the inverse RST matrix TNxR in one embodiment is NxR, which is smaller than the size NxN of a normal inverse transform matrix, and it is in a transpose relationship with the simplified transform matrix TRxN shown in Equation 4.
[0129] The matrix Tt in the Reduced Inv. Transform block shown in Figure 5(b) can represent the inverse RST matrix TRxNT (the superscript T indicates transpose). As shown in Figure 5(b), when the inverse RST matrix TRxNT is multiplied by the transformation coefficients for the target block, the modified transformation coefficients or residual samples for the target block can be derived. The inverse RST matrix TRxNT can also be expressed as (TRxN)TNxR.
[0130] More specifically, when the inverse RST is applied to a quadratic inverse transform, multiplying the transformation coefficients for the target block by the inverse RST matrix TRxNT yields the modified transformation coefficients for the target block. On the other hand, when the inverse RST is applied to a linear inverse transform, multiplying the transformation coefficients for the target block by the inverse RST matrix TRxNT yields the residual sample for the target block.
[0131] In one embodiment, when the size of the block to which the inverse transform is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the RST shown in Figure 5(b) can be expressed by a matrix operation as shown in Equation 7 below.
[0132]
number
[0133] In equation 7, c1 to c16 can represent the conversion coefficients for the target block. From the calculation result of equation 7, rj, which represents the modified conversion coefficients for the target block or the residual sample for the target block, can be derived, and the derivation process of rj is as shown in equation 8.
[0134]
number
[0135] The result of the calculation in Equation 8 allows us to derive r1 to rN, which represent the modified transformation coefficients for the target block or the residual samples for the target block. Considering this from the perspective of the size of the inverse transformation matrix, the size of a normal inverse transformation matrix is 64x64 (NxN), but the size of the simplified inverse transformation matrix is reduced to 64x16 (NxR), so compared to performing a normal inverse transformation, the memory usage when performing the inverse RST can be reduced by a ratio of R / N. Also, compared to the number of multiplication operations NxN when using a normal inverse transformation matrix, the number of multiplication operations can be reduced by a ratio of R / N (NxR) when using a simplified inverse transformation matrix.
[0136] On the other hand, the transformation set configuration shown in Table 2 can also be applied to an 8x8 RST. That is, the 8x8 RST can be applied using the transformation sets in Table 2. Since one transformation set consists of two or three transformations (kernels) depending on the prediction mode on the screen, it can be configured to select one of up to four transformations, including the case where a quadratic transformation is not applied. The transformation when a quadratic transformation is not applied can be considered as the application of the identity matrix. If we assign indices 0, 1, 2, and 3 to the four transformations respectively (for example, index 0 can be assigned to the identity matrix, i.e., when a quadratic transformation is not applied), then the transformation to be applied can be specified by signaling a transformation index or lfnst index, which is a syntax element, for each block of transformation coefficients. That is, via the transformation index, an 8x8 RST can be specified for the upper left block of the 8x8 in the RST configuration, or an 8x8 lfnst can be specified if LFNST is applied. 8x8 lfnst and 8x8 RST refer to transformations that can be applied to an 8x8 region contained within the block of the transformation coefficient when the W and H of the target block to be transformed are all equal to or greater than 8, and this 8x8 region may be the upper left 8x8 region inside the block of the transformation coefficient. Similarly, 4x4 lfnst and 4x4 RST refer to transformations that can be applied to a 4x4 region contained within the block of the transformation coefficient when the W and H of the target block are all equal to or greater than 4, and this 4x4 region may be the upper left 4x4 region inside the block of the transformation coefficient.
[0137] On the other hand, according to one embodiment of this document, in the encoding process, instead of a 16x64 transformation kernel matrix, it is possible to select only 48 data points from the 64 data points constituting an 8x8 region and apply a transformation kernel matrix of up to 16x48. Here, "up to" means that for an mx48 transformation kernel matrix that can generate m coefficients, the maximum value of m is 16. That is, when RST is executed by applying an mx48 transformation kernel matrix (m ≤ 16) to an 8x8 region, it is possible to receive 48 data inputs and generate m coefficients. When m is 16, it receives 48 data inputs and generates 16 coefficients. That is, if the 48 data points form a 48x1 vector, a 16x48 matrix and a 48x1 vector can be multiplied in order to generate a 16x1 vector. At that time, the 48 data points constituting the 8x8 region can be appropriately arranged to construct a 48x1 vector. At that time, applying a transformation kernel matrix of up to 16x48 and performing matrix operations will generate 16 modified transformation coefficients, which can be placed in the upper left 4x4 region according to the scanning order, while the upper right 4x4 region and the lower left 4x4 region can be filled with 0.
[0138] For the inverse transformation of the decoding process, the transposed matrix of the transformation kernel matrix described above can be used. That is, when inverse RST or LFNST is performed in the inverse transformation process executed by the decoding device, the input coefficient data to which the inverse RST is applied consists of one-dimensional vectors according to a predetermined arrangement order, and the modified coefficient vectors obtained by multiplying the one-dimensional vectors by the matrix of the inverse RST on the left side can be arranged in two-dimensional blocks according to a predetermined arrangement order.
[0139] To summarize, when RST or LFNST is applied to an 8x8 region during the transformation process, a matrix operation is performed between the 48 transformation coefficients in the upper left, upper right, and lower left regions of the 8x8 region (excluding the lower right region) and a 16x48 transformation kernel matrix. For this matrix operation, the 48 transformation coefficients are input into a one-dimensional array. After this matrix operation, 16 modified transformation coefficients are derived, and these modified transformation coefficients can be arranged in the upper left region of the 8x8 region.
[0140] Conversely, when inverse RST or LFNST is applied to an 8x8 region during the inverse transformation process, the 16 transformation coefficients corresponding to the upper left side of the 8x8 region are input in a one-dimensional array form according to the scanning order and can be used in matrix operations with a 48x16 transformation kernel matrix. That is, the matrix operation in such a case can be expressed as (48x16 matrix) * (16x1 transformation coefficient vector) = (48x1 modified transformation coefficient vector). Here, the nx1 vector can be interpreted as an nx1 matrix, so it may also be expressed as an nx1 column vector. Also, * means matrix multiplication. When such a matrix operation is performed, 48 modified transformation coefficients can be derived, and these 48 modified transformation coefficients can be arranged in the upper left, upper right, and lower left regions of the 8x8 region, excluding the lower right region.
[0141] On the other hand, when the quadratic inverse transform is performed based on RST, the inverse transform unit 135 of the encoding device 100 and the inverse transform unit 222 of the decoding device 200 may each include an inverse RST unit that derives corrected transform coefficients based on the inverse RST of the transform coefficients, and an inverse linear transform unit that derives the residual sample for the target block based on the inverse linear transform of the corrected transform coefficients. The inverse linear transform means the inverse transform of the linear transform that was applied to the residual. In this document, deriving transform coefficients based on a transform means deriving transform coefficients by applying the transform.
[0142] Looking specifically at the previously mentioned non-separated transform, LFNST, it is as follows: LFNST can include both a forward transform by the encoding device and an inverse transform by the decoding device.
[0143] The encoding device applies a primary (core) transform, and then applies a secondary transform to the derived result (or part of the result) as input.
[0144]
number
[0145] In equation 9 above, x and y are the input and output of the quadratic transformation, respectively, G is the matrix representing the quadratic transformation, and the transform basis vectors are composed of column vectors. In the case of inverse LFNST, when the dimension of the transformation matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the dimension of GT is obtained by transposing matrix G.
[0146] In the case of inverse LFNST, the dimensions of matrix G are [48x16], [48x8], [16x16], and [16x8], where the [48x8] matrix and the [16x8] matrix are submatrices obtained by sampling eight transformation basis vectors from the left side of the [48x16] matrix and the [16x16] matrix, respectively.
[0147] On the other hand, in the case of forward LFNST, the dimensions of the matrix GT are [16x48], [8x48], [16x16], and [8x16], and the [8x48] matrix and the [8x16] matrix are submatrices obtained by sampling eight transformation basis vectors from above the [16x48] matrix and the [16x16] matrix, respectively.
[0148] Therefore, in the case of forward LFNST, the input x can be a [48x1] vector or a [16x1] vector, and the output y can be a [16x1] vector or an [8x1] vector. Since the output of the forward linear transformation in video coding and decoding is 2D data, in order to construct a [48x1] vector or a [16x1] vector as input x, the 2D data that is the output of the forward transformation must be appropriately arranged to construct a 1D vector.
[0149] Figure 6 shows an example of the sequence for arranging the output data of a forward linear transformation into a one-dimensional vector. The left-hand diagrams of Figure 6(a) and (b) show the sequence for creating a [48x1] vector, and the right-hand diagrams of Figure 6(a) and (b) show the sequence for creating a [16x1] vector. In the case of LFNST, the 2D data is sequentially arranged in the order shown in Figure 6(a) and (b) to obtain a one-dimensional vector x.
[0150] The orientation of the output data of such a forward linear transformation can be determined by the intra-prediction mode of the current block. For example, if the intra-prediction mode of the current block is horizontal with respect to the diagonal direction, the output data of the forward linear transformation can be arranged in the order shown in Figure 6(a), and if the intra-prediction mode of the current block is vertical with respect to the diagonal direction, the output data of the forward linear transformation can be arranged in the order shown in Figure 6(b).
[0151] For example, a different ordering can be applied than the orderings in Figures 6(a) and (b). To derive the same result (y vector) as when applying the orderings in Figures 6(a) and (b), the column vectors of matrix G can be rearranged to match the chosen ordering. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0152] Since the output y derived via equation 9 is a one-dimensional vector, if a configuration that processes the result of a forward quadratic transformation as input, such as a configuration that performs quantization or residual coding, requires two-dimensional data as input data, then the output y vector from equation 9 must again be appropriately positioned in 2D data.
[0153] Figure 7 shows an example of the order in which the output data of a forward quadratic transformation is arranged in a two-dimensional block.
[0154] In the case of LFNST, the output values can be placed in a 2D block according to a predetermined scan order. Figure 7(a) shows that when the output y is a [16x1] vector, the output values are placed in 16 positions in the 2D block according to the diagonal scan order. Figure 7(b) shows that when the output y is an [8x1] vector, the output values are placed in 8 positions in the 2D block according to the diagonal scan order, and the remaining 8 positions are filled with 0. In Figure 7(b), X is shown to be filled with 0.
[0155] In another example, depending on the configuration performing quantization or residual coding, the order in which the output vector y is processed can be according to a pre-set order, so the output vector y may not be placed in a 2D block, as shown in Figure 7. However, in the case of residual coding, data coding can be performed in units of 2D blocks (e.g., 4x4) such as CG (Coefficient Group), in which case the data can be arranged according to a specific order, as in the diagonal scan order of Figure 7.
[0156] On the other hand, the decoding device can construct a one-dimensional input vector y by arranging the two-dimensional data output through an inverse quantization process, etc., according to a pre-set scan order for the reverse transformation. The input vector y can be output as the input vector x by the following formula.
[0157]
number
[0158] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16x1] vector or an [8x1] vector, by the G matrix. In the case of inverse LFNST, the output vector x can be a [48x1] vector or a [16x1] vector.
[0159] The output vector x is arranged in a 2D block according to the order shown in Figure 6, and this 2D data becomes the input data (or part of the input data) for the inverse linear transformation.
[0160] Therefore, the inverse quadratic transformation is generally the opposite of the forward quadratic transformation process. In the case of the inverse transformation, unlike the forward transformation, the inverse quadratic transformation is applied first, followed by the inverse linear transformation.
[0161] In reverse LFNST, one of eight [48x16] matrices or eight [16x16] matrices can be selected as the transformation matrix G. Which of the [48x16] and [16x16] matrices to apply is determined by the size and shape of the block.
[0162] Furthermore, the eight matrices can be derived from four transformation sets, as shown in Table 2 above, and each transformation set can consist of two matrices. Which of the four transformation sets to use is determined by the intra-prediction mode, and more specifically, the transformation set is determined based on the extended intra-prediction mode value, taking into account the wide-angle intra-prediction mode (WAIP). Which of the two matrices that make up the selected transformation set is chosen is derived via index signaling. More specifically, the index value that is transmitted can be 0, 1, or 2, where 0 indicates that LFNST should not be applied, and 1 and 2 can indicate either of the two transformation matrices that make up the transformation set selected based on the intra-prediction mode value.
[0163] On the other hand, as mentioned above, whether or not to apply the transformation matrix (48x16 or 16x16) to LFNST depends on the size and shape of the block being transformed.
[0164] Figure 8 shows the shapes of blocks to which LFNST is applied. Figure 8(a) shows a 4x4 block, (b) shows 4x8 and 8x4 blocks, (c) shows a 4xN or Nx4 block where N is 16 or greater, (d) shows an 8x8 block, and (e) shows an MxN block where M≧8, N≧8, and N>8 or M>8.
[0165] In Figure 8, blocks with thick borders indicate the regions to which LFNST is applied. For blocks (a) and (b) in Figure 8, LFNST is applied to the top-left 4x4 region, and for block (c) in Figure 8, LFNST is applied to two consecutively placed top-left 4x4 regions. In Figures (a), (b), and (c), LFNST is applied in units of 4x4 regions, so such LFNST will be hereafter referred to as "4x4 LFNST," and the transformation matrix can be a [16x16] or [16x8] matrix with a matrix dimension of G in equations 9 and 10.
[0166] More specifically, a [16x8] matrix is applied to the 4x4 block (4x4TU or 4x4CU) in Figure 8(a), and a [16x16] matrix is applied to the blocks in Figures 8(b) and (c). This is to match the computational complexity for the worst case to 8 multiplications per sample.
[0167] In Figures 8(d) and (e), LFNST is applied to the upper left 8x8 region, and such LFNST will be hereafter referred to as "8x8 LFNST". The transformation matrix can be either a [48x16] or [48x8] matrix. In the case of forward LFNST, a [48x1] vector (the x vector in equation 9) is input as input data, so not all sample values in the upper left 8x8 region are used as input values for forward LFNST. That is, as seen in the left-to-right order of Figure 6(a) or Figure 6(b), the bottom-right 4x4 block is left as is, and the [48x1] vector can be constructed based on the samples belonging to the remaining three 4x4 blocks.
[0168] A [48x8] matrix can be applied to the 8x8 block (8x8TU or 8x8CU) in Figure 8(d), and a [48x16] matrix can be applied to the 8x8 block in Figure 8(e). This is also to match the computational complexity for the worst case to 8 multiplications per sample.
[0169] Depending on the shape of the block, applying the corresponding forward LFNST (4x4LFNST or 8x8LFNST) generates 8 or 16 output data (y vectors in Equation 9, [8x1] or [16x1] vectors). In forward LFNST, due to the properties of the matrix GT, the number of output data is equal to or less than the number of input data.
[0170] Figure 9 is a diagram illustrating the arrangement of output data for a forward LFNST as an example, showing blocks in which the output data for the forward LFNST is arranged according to the block shape.
[0171] In Figure 9, the shaded area in the upper left of the block corresponds to the region where the output data of the forward LFNST is located. Positions marked with 0 indicate samples filled with a value of 0, while the remaining area represents the region that is not modified by the forward LFNST. In the region not modified by the LFNST, the output data of the forward linear transformation remains unchanged.
[0172] As mentioned above, the dimension of the transformation matrix applied changes depending on the shape of the block, so the number of output data also changes. As shown in Figure 9, the output data of the forward LFNST may not fill the entire upper left 4x4 block. In the cases of Figure 11(a) and (d), the blocks or parts of the inside of the blocks shown by the thick lines are to which the [16x8] matrix and [48x8] matrix are applied, respectively, and an [8x1] vector is generated in the output of the forward LFNST. That is, according to the scan order shown in Figure 7(b), only 8 output data are filled as in Figure 9(a) and (d), and the remaining 8 positions can be filled with 0. In the case of the LFNST applied block in Figure 8(d), the two 4x4 blocks on the upper right and lower left sides adjacent to the upper left 4x4 block are also filled with 0 values as in Figure 9(d).
[0173] As described above, the basic approach is to signal the LFNST index and specify whether LFNST is applicable and which transformation matrix to apply. As shown in Figure 9, when LFNST is applied, the number of output data for forward LFNST may be equal to or less than the number of input data, resulting in a region filled with zero values as follows.
[0174] 1) As shown in Figure 9(a), within the 4x4 block on the upper left, the 8th and subsequent positions in the scan order, i.e., samples 9 through 16.
[0175] 2) As shown in Figures 9(d) and (e), a [16×48] matrix or an [8×48] matrix is applied to two 4×4 blocks adjacent to the upper left 4×4 block, or the second and third 4×4 blocks in the scan order.
[0176] Therefore, by checking the areas in 1) and 2) above, if non-zero data exists, it is certain that LFNST is not applied, and thus the signaling of the LFNST index can be omitted.
[0177] For example, in the case of LFNST adopted in the VVC standard, signaling of the LFNST index is performed after residual coding, so the encoding device can determine whether or not there is non-zero data (effectiveness factor) at all positions within the TU or CU block via residual coding. Therefore, the encoding device can determine whether or not to perform signaling for the LFNST index based on the presence or absence of non-zero data, and the decoding device can determine whether or not the LFNST index can be parsed. If there is no non-zero data in the area specified in 1) and 2) above, then signaling for the LFNST index will be performed.
[0178] On the other hand, the following simplification method can be applied to the adopted LFNST.
[0179] (i) For example, the number of output data for a forward LFNST can be limited to a maximum of 16.
[0180] In the case of Figure 8(c), a 4x4 LFNST is applied to each of the two adjacent 4x4 regions in the upper left, and in this case, a maximum of 32 LFNST output data can be generated. If the number of output data for forward LFNST is limited to a maximum of 16, then for a 4xN / Nx4 (N≧16) block (TU or CU), a 4x4 LFNST is applied only to one 4x4 region in the upper left, and LFNST can be applied only once to all blocks in Figure 8. This simplifies the implementation for image coding.
[0181] (ii) For example, zero-out can be applied to regions to which LFNST is not applied. In this document, zero-out can mean filling the values of all positions belonging to a particular region with the value 0. That is, zero-out can be applied to regions that are not modified by LFNST and maintain the result of the forward linear transformation. As mentioned above, LFNST is divided into 4x4 LFNST and 8x8 LFNST, so zero-out can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.
[0182] (ii)-(A) When a 4x4 LFNST is applied, areas to which the 4x4 LFNST is not applied can be zeroed out. Figure 10 shows an example of zeroing out in a block to which a 4x4 LFNST is applied.
[0183] As shown in Figure 10, for a block to which a 4x4 LFNST is applied, that is, for blocks (a), (b), and (c) in Figure 9, the region to which LFNST is not applied can all be filled with 0.
[0184] On the other hand, Figure 10(d) shows, as an example, that when the maximum number of output data for the forward LFNST is limited to 16, zero-out is performed on the remaining blocks to which the 4x4 LFNST is not applied.
[0185] (ii)-(B) When an 8x8 LFNST is applied, areas to which the 8x8 LFNST is not applied can be zeroed out. Figure 11 shows an example of zeroing out in a block to which an 8x8 LFNST is applied.
[0186] As shown in Figure 11, for a block to which an 8x8 LFNST is applied, that is, for blocks (d) and (e) in Figure 9, the region to which LFNST is not applied can be completely filled with zeros.
[0187] (iii) When LFNST is applied by the zero-out method presented in (ii) above, the area filled with zeros may change. Therefore, the zero-out method proposed in (ii) above allows checking for the presence of non-zero data over a wider area than in the case of LFNST in Figure 9.
[0188] For example, when applying (ii)-(B), after checking whether non-zero data exists up to the regions filled with zero values in (d) and (e) of Figure 9, and further up to the regions filled with zeros in Figure 11, signaling to the LFNST index can be performed only if no non-zero data exists.
[0189] Of course, even if the zero-out proposed in (ii) above is applied, it is still possible to check whether non-zero data exists, similar to existing LFNST index signaling. That is, it is possible to check whether non-zero data exists for blocks filled with zeros in Figure 9 and apply LFNST index signaling. In such a case, zero-out can be performed only on the encoding device, and the decoding device can not assume that zero-out exists, that is, it can only check whether non-zero data exists for areas explicitly represented by zeros in Figure 9 and perform LFNST index parsing.
[0190] Various embodiments can be derived by applying combinations of the simplification methods ((i), (ii)-(A), (ii)-(B), (iii)) to the aforementioned LFNST. Of course, the combinations for the simplification methods are not limited to the embodiments described below, and any combination can be applied to the LFNST.
[0191] Embodiment
[0192] - Limit the number of output data for forward LFNST to a maximum of 16 → (i)
[0193] - When 4x4 LFNST is applied, all areas where 4x4 LFNST is not applied are zeroed out → (ii)-(A)
[0194] - When an 8x8 LFNST is applied, all areas where the 8x8 LFNST is not applied are zeroed out → (ii)-(B)
[0195] - After checking whether non-zero data exists in the regions filled with existing zero values and in the regions filled with zeros due to additional zero-outs ((ii)-(A), (ii)-(B)), LFNST indexing signaling is performed only if no non-zero data exists → (iii)
[0196] In the above embodiment, when LFNST is applied, the area where non-zero output data may exist is limited to the upper left 4x4 area. More specifically, in Figures 10(a) and 11(a), the 8th position in the scan sequence is the last position where non-zero data may exist, and in Figures 10(b) and (d) and 11(b), the 16th position in the scan sequence (i.e., the position at the lower right edge of the upper left 4x4 block) is the last position where non-zero data may exist.
[0197] Therefore, when LFNST is applied, it is possible to determine whether or not the LFNST index can signal after checking whether or not there is non-zero data at positions where the residual coding process is not permitted (beyond the last position).
[0198] In the zero-out method proposed in (ii), the number of data points generated when both linear transformation and LFNST are applied is reduced, thus reducing the computational load required when performing the overall transformation process. Specifically, when LFNST is applied, zero-out is also applied to the forward linear transformation output data in areas where LFNST is not applied, so there is no need to generate data for areas that will be zero-out from the time of the forward linear transformation. Therefore, the computational load required for generating such data can be saved. The additional effects of the zero-out method proposed in (ii) can be summarized as follows:
[0199] Firstly, as mentioned above, the amount of computation required to execute the entire transformation process is reduced.
[0200] In particular, applying (ii)-(B) reduces the computational complexity in the worst-case scenario, thereby lightening the transformation process. More specifically, while large-scale linear transformations generally require a large amount of computation, applying (ii)-(B) can reduce the number of data points derived as a result of a forward LFNST execution to 16 or less, and the effect of reducing the computational complexity of the transformation increases further as the overall block (TU or CU) size increases.
[0201] Secondly, the amount of computation required for the entire conversion process is reduced, thereby lowering the power consumption required to perform the conversion.
[0202] Thirdly, it reduces the latency associated with the conversion process.
[0203] Quadratic transformations like LFNST add computational complexity to existing linear transformations, thus increasing the overall delay time associated with the transformation. In particular, in the case of intra-prediction, since the reconstruction data of adjacent blocks is used in the prediction process, the increase in delay time due to the quadratic transformation during encoding leads to an increase in the delay time until reconstruction, which can lead to an overall increase in the delay time of intra-prediction encoding.
[0204] However, by applying the zero-out method presented in (ii), the delay time of the primary conversion execution can be significantly reduced when LFNST is applied, so the delay time for the entire conversion execution will remain the same or be reduced, making it easier to realize an encoding device.
[0205] On the other hand, conventional intra-prediction treated the block currently to be encoded as a single encoding unit and performed encoding without division. However, ISP (Intra Sub-Paritions) coding means dividing the block currently to be encoded horizontally or vertically and performing intra-predictive coding. In this case, encoding / decoding is performed on each divided block to generate a reconstructed block, and the reconstructed block is used as a reference block for the next divided block. For example, during ISP coding, one coding block may be divided into two or four sub-blocks and coded, and in ISP, one sub-block performs intra-prediction by referencing the reconstructed pixel value of the adjacent sub-block located to its left or adjacent upper. Hereafter, "coding" is used as a concept that includes both encoding performed in the encoding device and decoding performed in the decoding device.
[0206] ISP divides a block into two or four subpartitionings vertically or horizontally, depending on its size. For example, the smallest block size to which ISP can be applied is 4x8 or 8x4. If the block size is larger than 4x8 or 8x4, the block will be divided into four subpartitionings.
[0207] When applying ISP, subblocks are coded sequentially depending on the division configuration, for example, horizontally or vertically, from left to right or top to bottom. After the inverse transformation and intra-prediction process for one subblock, followed by the reconstruction process, coding is performed for the next subblock. For the leftmost or topmost subblock, the reconstruction pixels of the already coded block are referenced, as in the usual intra-prediction method. Furthermore, if each edge of a subsequent internal subblock is not adjacent to a previous subblock, the reconstruction pixels of the adjacent coding block that has already been coded are referenced, as in the usual intra-prediction method, to derive the reference pixels adjacent to that edge.
[0208] In ISP coding mode, all subblocks are coded in the same intra-predictive mode, and flags indicating whether to use ISP coding and in which direction (horizontal or vertical) to divide are signaled. At this time, the number of subblocks can be adjusted to 2 or 4 depending on the block shape, and if the size (width x height) of one subblock is less than 16, it is possible to restrict the division to that subblock or to not apply ISP coding at all.
[0209] On the other hand, in ISP prediction mode, one coding unit is divided into two or four partition blocks, i.e., subblocks, for prediction, and the same in-screen prediction mode is applied to these two or four partition blocks.
[0210] As mentioned above, the division direction can be either horizontal (when an M×N coding unit with horizontal and vertical lengths M and N respectively is divided horizontally, it is divided into M×(N / 2) blocks if it is divided into two, and into M×(N / 4) blocks if it is divided into four) or vertical (when an M×N coding unit is divided vertically, it is divided into (M / 2)×N blocks if it is divided into two, and into (M / 4)×N blocks if it is divided into four). When divided horizontally, the partition blocks are coded in order from top to bottom, and when divided vertically, the partition blocks are coded in order from left to right. The partition block currently being coded can be predicted by referring to the restored pixel values of the upper (left) partition block when the division is horizontal (vertical).
[0211] A transformation can be applied to the residual signal generated by the ISP prediction method on a partition block basis. Based on the forward direction, a first-order transformation (core transform or primary transform) can be applied, and not only the existing DCT-2 but also the DST-7 / DCT-8 combination-based MTS (Multiple Transform Selection) technology can be applied. The transformation coefficients generated by the first-order transformation can then be subjected to a forward LFNST (Low Frequency Non-Separable Transform) to produce the final modified transformation coefficients.
[0212] In other words, LFNST can be applied to partition blocks that have been divided using the ISP prediction mode, and as mentioned above, the same intra-prediction mode is applied to the divided partition blocks. Therefore, when selecting an LFNST set derived based on the intra-prediction mode, the derived LFNST set can be applied to all partition blocks. That is, since the same intra-prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks.
[0213] On the other hand, LFNST can only be applied to transformation blocks where both the width and height are 4 or greater. Therefore, if the width or height of a partition block divided according to the ISP prediction scheme is less than 4, LFNST will not be applied and the LFNST index will not be signaled. Also, when applying LFNST to each partition block, that partition block can be considered as a single transformation block. Of course, if the ISP prediction scheme is not applied, LFNST will be applied to the coding block.
[0214] The application of LFNST to each partition block is explained in detail as follows:
[0215] For example, after applying forward LFNST to individual partition blocks, zero-out is applied to the upper left 4x4 region, leaving only a maximum of 16 coefficients (8 or 16) according to the conversion coefficient scan order, and filling all remaining positions and regions with zero values.
[0216] Alternatively, for example, if the length of one side of the partition block is 4, LFNST can be applied only to the upper left 4x4 region. If the length of all sides of the partition block, i.e., the width and height, is 8 or greater, LFNST can be applied to the remaining 48 coefficients, excluding the lower right 4x4 region within the upper left 8x8 region.
[0217] Alternatively, as an example, to match the worst-case computational complexity to 8 multiplications per sample, if each partition block is 4x4 or 8x8, only 8 transformation coefficients can be output after applying forward LFNST. That is, if the partition block is 4x4, an 8x16 matrix is applied as the transformation matrix, and if the partition block is 8x8, an 8x48 matrix is applied as the transformation matrix.
[0218] On the other hand, in the current VVC standard, LFNST index signaling is performed on a coding unit basis. Therefore, in ISP prediction mode, when LFNST is applied to all partition blocks, the same LFNST index value can be applied to those partition blocks. That is, once an LFNST index value is sent at the coding unit level, that LFNST index can be applied to all partition blocks within the coding unit. As mentioned above, LFNST index values have values of 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate two transformation matrices that exist within a single LFNST set when LFNST is applied.
[0219] As described above, the LFNST set is determined by the intra-prediction mode, and in the case of ISP prediction mode, all partition blocks within the coding unit are predicted in the same intra-prediction mode, so the partition blocks can refer to the same LFNST set.
[0220] As another example, while LFNST index signaling is still performed on a coding unit basis, in ISP prediction mode, the decision to apply LFNST is not made uniformly for all partition blocks. Instead, a separate condition is used to determine whether to apply the LFNST index value signaled at the coding unit level to each partition block, or not to apply LFNST. Here, the separate condition is signaled in the form of a flag for each partition block via the bitstream. If the flag value is 1, the LFNST index value signaled at the coding unit level is applied; if the flag value is 0, LFNST is not applied.
[0221] The following describes how to maintain worst-case computational complexity when applying LFNST to ISP mode.
[0222] In ISP mode, LFNST application can be restricted to maintain the number of multiplications per sample (or per coefficient, per position) below a certain value when LFNST is applied. Depending on the size of the partition block, LFNST can be applied as follows to maintain the number of multiplications per sample (or per coefficient, per position) at 8 or less.
[0223] 1. If both the horizontal and vertical dimensions of a partition block are 4 or greater, the same calculation complexity adjustment method as the worst-case scenario for LFNST in the current VVC standard can be applied.
[0224] In other words, if the partition block is a 4x4 block, instead of a 16x16 matrix, an 8x16 matrix obtained by sampling the top 8 rows from a 16x16 matrix can be applied in the forward direction, and a 16x8 matrix obtained by sampling the left 8 columns from a 16x16 matrix can be applied in the reverse direction. Also, if the partition block is an 8x8 block, instead of a 16x48 matrix in the forward direction, an 8x48 matrix obtained by sampling the top 8 rows from a 16x48 matrix can be applied, and instead of a 48x16 matrix in the reverse direction, a 48x8 matrix obtained by sampling the left 8 columns from a 48x16 matrix can be applied.
[0225] For 4×N or N×4 (N>4) blocks, when performing a forward transformation, a 16×16 matrix is applied only to the upper-left 4×4 block. The resulting 16 coefficients are then placed in the upper-left 4×4 region, and the remaining region is filled with zeros. When performing a reverse transformation, the 16 coefficients located in the upper-left 4×4 block are arranged according to the scan order to construct an input vector. These vectors are then multiplied by a 16×16 matrix to generate 16 output data. The generated output data is placed in the upper-left 4×4 region, and the remaining region is filled with zeros.
[0226] In the case of 8×N or N×8 (N>8) blocks, when performing a forward transformation, a 16×48 matrix is applied only to the ROI region within the upper-left 8×8 block (the remaining region after excluding the lower-right 4×4 block from the upper-left 8×8 block). The 16 generated coefficients are then placed in the upper-left 4×4 region, and all other regions are filled with zero values. When performing a reverse transformation, the 16 coefficients located in the upper-left 4×4 block are arranged according to the scan order to construct an input vector, which is then multiplied by a 48×16 matrix to generate 48 output data. The generated output data fills the ROI region, and all remaining regions are filled with zero values.
[0227] Another example is to maintain the number of multiplication factors per sample (or per coefficient, per position) below a certain value by keeping it at 8 or less based on the size of the ISP coding unit, not the size of the ISP partition block. If there is only one ISP partition block that satisfies the conditions for LFNST to apply, the LFNST worst-case complexity calculation is applied based on the size of that coding unit, not the size of the partition block. For example, if a Luma coding block for a coding unit is divided into four 4x4 partition blocks and coded by ISP, and there are no non-zero conversion factors for two of these partition blocks, the other two partition blocks can be configured to generate 16 conversion factors each (not 8, based on the encoder).
[0228] The following describes how to signal the LFNST index when using ISP mode.
[0229] As mentioned above, the LFNST index has values of 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate one of the two LFNST kernel matrices included in the selected set of LFNSTs. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The current method of transmitting the LFNST index in the VVC standard is as follows:
[0230] 1. An LFNST index can be sent once per coding unit (CU), and in the case of a dual-tree, separate LFNST indexes are signaled for the luma block and chroma block, respectively.
[0231] 2. If the LFNST index is not signaled, the LFNST index value is determined to be the default value of 0 (infer). The following are cases in which the LFNST index value is inferred to be 0:
[0232] A. When a mode in which no transformation is applied (e.g., transform skip, BDPCM, lossless coding, etc.)
[0233] B. When the primary transformation is not DCT-2 (e.g., DST7 or DCT8), i.e., when the horizontal or vertical transformation is not DCT-2.
[0234] C. If the horizontal or vertical dimensions of the coding unit relative to the Luma block exceed the maximum size of the Luma conversion that can be converted, for example, if the maximum size of the Luma conversion that can be converted is 64, then LFNST cannot be applied if the size of the coding block relative to the Luma block is equivalent to 128 × 16.
[0235] In the case of a dual tree, it is determined whether the coding unit for the luma component and the coding unit for the chroma component will exceed the maximum luma conversion size. That is, it is checked whether the size of the luma block exceeds the maximum luma conversion size that can be converted, and for the chroma block, it is checked whether the length / width of the corresponding luma block for the color format and the size of the maximum luma conversion that can be converted exceed each other. For example, if the color format is 4:2:0, the length / width of the corresponding luma block will be twice that of the chroma block, and the size of the corresponding luma block conversion will be twice that of the chroma block. As another example, if the color format is 4:4:4, the length / width of the corresponding luma block and the size of the conversion will be the same as that of the corresponding chroma block.
[0236] A 64-length conversion or a 32-length conversion refers to a conversion applied to a horizontal or vertical dimension having lengths of 64 or 32, respectively, and "conversion size" refers to the respective lengths of 64 or 32.
[0237] If it is a single tree, check whether the size of the horizontally or vertically convertible luma block exceeds the maximum size of the luma conversion block that can be converted. If it does exceed this limit, LFNST index signaling may be omitted.
[0238] D. The LFNST index can only be sent if both the width and height of the coding unit are 4 or greater.
[0239] In the case of a dual tree, the LFNST index can only be signaled if both the horizontal and vertical dimensions for the relevant component (i.e., the luma or chroma component) are 4 or greater.
[0240] In the case of a single tree, the LFNST index can be signaled when both the horizontal and vertical lengths of the luma components are 4 or greater.
[0241] E. If the position of the last non-zero coefficient is not the DC position (top-left position of the block), in a dual-tree type luma block, if the position of the last non-zero coefficient is not the DC position, an LFNST index is sent. In a dual-tree type chroma block, if either the position of the last non-zero coefficient for Cb or the position of the last non-zero coefficient for Cr is not the DC position, the corresponding LNFST index is sent.
[0242] If it is a single-tree type, the LFNST index is sent if the position of the last non-zero coefficient of any of the Luma, Cb, or Cr components is not at the DC position.
[0243] Here, if the CBF (coded block flag) value, which indicates whether or not a conversion coefficient exists for a single conversion block, is 0, the position of the last non-zero coefficient for that conversion block is not checked in order to determine whether or not to perform LFNST index signaling. In other words, if the CBF value is 0, the conversion is not applied to that block, so when checking the conditions for LFNST index signaling, the position of the last non-zero coefficient does not need to be considered.
[0244] For example, 1) in a dual-tree type, if the component is luma and its CBF value is 0, the LFNST index is not signaled; 2) in a dual-tree type, if the component is chroma and its CBF value for Cb is 0 and its CBF value for Cr is 1, only the position of the last non-zero coefficient for Cr is checked and the corresponding LFNST index is sent; 3) in a single-tree type, only the position of the last non-zero coefficient is checked for components where the CBF value for luma, Cb, and Cr is 1.
[0245] If it is confirmed that a conversion coefficient exists at a location where an F.LFNST conversion coefficient cannot exist, LFNST index signaling can be omitted. For 4x4 and 8x8 conversion blocks, LFNST conversion coefficients exist at 8 locations from the DC position according to the conversion coefficient scan order in the VVC standard, and all remaining positions are filled with 0. For conversion blocks that are not 4x4 or 8x8, LFNST conversion coefficients exist at 16 locations from the DC position according to the conversion coefficient scan order in the VVC standard, and all remaining positions are filled with 0.
[0246] Therefore, if a non-zero conversion coefficient exists in the region that should be filled with zero values after residual coding, LFNST index signaling can be omitted.
[0247] On the other hand, ISP mode may apply only to luma blocks, or it may apply to both luma and chroma blocks. As mentioned above, when ISP prediction is applied, the coding unit is divided into two or four partition blocks for prediction, and the transformation is applied to each of these partition blocks. Therefore, when determining the conditions for signaling the LFNST index on a coding unit basis, the fact that LFNST can be applied to each of the partition blocks must be taken into consideration. Also, if ISP prediction mode is applied only to a specific component (e.g., luma blocks), the fact that the component is divided into partition blocks only must be taken into consideration when signaling the LFNST index. When ISP mode is applied, the possible LFNST index signaling methods can be summarized as follows.
[0248] 1. An LFNST index can be sent once per coding unit (CU), and in the case of a dual-tree, separate LFNST indexes can be signaled for the luma block and chroma block, respectively.
[0249] 2. If the LFNST index is not signaled, the LFNST index value is determined to the default value of 0 (infer). The following are cases in which the LFNST index value is inferred to 0:
[0250] A. When a mode in which no transformation is applied (e.g., transform skip, BDPCM, lossless coding, etc.)
[0251] B. If the horizontal or vertical dimensions of the coding unit relative to the Luma block exceed the maximum size of the Luma conversion that can be converted, for example, if the maximum size of the Luma conversion that can be converted is 64, then LFNST cannot be applied if the size of the coding block relative to the Luma block is the same as 128 × 16.
[0252] Alternatively, the decision to signal the LFNST index can be based on the size of the partition block instead of the coding unit. That is, if the horizontal or vertical length of the partition block relative to the given luma block exceeds the size of the maximum luma conversion that can be converted, the LFNST index signaling can be omitted, and the LFNST index value can be inferred as 0.
[0253] In the case of a dual tree, it is determined whether the coding unit or partition block for the luma component and the coding unit or partition block for the chroma component each exceed the maximum conversion block size. Specifically, the length and width of the coding unit or partition block for luma are compared with the maximum luma conversion size, and if either is greater than the maximum luma conversion size, LFNST is not applied. In the case of a coding unit or partition block for chroma, the width / length of the corresponding luma block for the color format is compared with the maximum possible luma conversion size. For example, if the color format is 4:2:0, the width / length of the corresponding luma block will be twice that of the chroma block, and the conversion size of the corresponding luma block will be twice that of the chroma block. As another example, if the color format is 4:4:4, the width / length and conversion size of the corresponding luma block will be the same as that of the corresponding chroma block.
[0254] In the case of a single tree, after checking whether the horizontal or vertical conversion of a luma block (coding unit or partition block) exceeds the maximum luma conversion block size that can be converted, LFNST index signaling may be omitted if it does.
[0255] C. If LFNST, which is included in the current VVC standard, is applied, the LFNST index can only be sent if both the width and height of the partition block are 4 or greater.
[0256] If we were to apply LFNST to 2×M(1×M) or M×2(M×1) blocks in addition to the LFNST currently included in the VVC standard, then the LFNST index could only be sent if the partition block size is greater than or equal to a 2×M(1×M) or M×2(M×1) block. Here, P×Q block being greater than or equal to an R×S block means that P≧R and Q≧S.
[0257] In summary, an LFNST index can only be sent if the partition block is greater than or equal to the minimum size for which LFNST is applicable. In the case of a dual tree, an LFNST index can only be signaled if the partition block for the luma or chroma component is greater than or equal to the minimum size for which LFNST is applicable. In the case of a single tree, an LFNST index can only be signaled if the partition block for the luma component is greater than or equal to the minimum size for which LFNST is applicable.
[0258] In this document, an M×N block being greater than or equal to a K×L block means that M is greater than or equal to K and N is greater than or equal to L. An M×N block being greater than a K×L block means that M is greater than or equal to K and N is greater than or equal to L, while M is greater than K or N is greater than L. An M×N block being less than or equal to a K×L block means that M is less than or equal to K and N is less than or equal to L, while M is less than or equal to K and N is less than or equal to L.
[0259] D. If the position of the last non-zero coefficient is not the DC position (top-left corner of the block), then in the case of a dual-tree type chroma block, an LFNST can be transmitted if the position of the last non-zero coefficient in at least one of the partition blocks is not the DC position. In the case of a dual-tree type chroma block, the LNFST index can be transmitted if the position of the last non-zero coefficient in all partition blocks for Cb (assuming there is one partition block if the ISP mode is not applied to the chroma component) and the position of the last non-zero coefficient in all partition blocks for Cr (assuming there is one partition block if the ISP mode is not applied to the chroma component) are not the DC position.
[0260] In the case of a single-tree type, if the position of the last non-zero coefficient in any one of the partition blocks for the Luma component, Cb component, and Cr component is not at the DC position, the corresponding LFNST index can be sent.
[0261] Here, if the CBF (coded block flag) value, which indicates whether or not a conversion coefficient exists for each partition block, is 0, the position of the last non-zero coefficient for that partition block is not checked in order to determine whether or not to perform LFNST index signaling. In other words, since no conversion is applied to that block when the CBF value is 0, the position of the last non-zero coefficient for that partition block is not considered when checking the conditions for LFNST index signaling.
[0262] For example, 1) in a dual-tree type with luma components, if the corresponding CBF value is 0 for each partition block, the corresponding partition block is excluded when deciding whether or not to perform LFNST index signaling; 2) in a dual-tree type with chroma components, if the CBF value for Cb is 0 and the CBF value for Cr is 1 for each partition block, only the position of the last non-zero coefficient for Cr is checked to decide whether or not to perform LFNST index signaling; 3) in a single-tree type, the position of the last non-zero coefficient is checked only for blocks where the CBF value is 1 for all partition blocks of luma, Cb, and Cr components to decide whether or not to perform LFNST index signaling.
[0263] In ISP mode, the video information may be configured so as not to check the position of the last non-zero coefficient, and embodiments relating to this are as follows:
[0264] i. In ISP mode, the check for the position of the last non-zero coefficient is omitted for both luma blocks and chroma blocks, and LFNST index signaling is allowed. That is, for all partition blocks, LFNST index signaling is allowed even if the position of the last non-zero coefficient is the DC position or the corresponding CBF value is 0.
[0265] ii. In ISP mode, the check regarding the position of the last non-zero coefficient is omitted only for luma blocks, while in chroma blocks, the check regarding the position of the last non-zero coefficient is performed using the method described above. For example, in the case of a dual-tree type and a luma block, LFNST index signaling is allowed without checking the position of the last non-zero coefficient, while in the case of a dual-tree type and a chroma block, the existence of a DC position for the position of the last non-zero coefficient is checked using the method described above to determine whether or not to signal the corresponding LFNST index.
[0266] iii. In the case of the ISP mode and the single-tree type, the method of item (i) or (ii) is applied. That is, when applying item (i) in the ISP mode and the single-tree type, the check regarding the position of the non-zero coefficient at the end is omitted for both the luma block and the chroma block, and the LFNST index signaling is allowed. Or, when applying item (ii), the check regarding the position of the non-zero coefficient at the end is omitted for the partition block for the luma component, and for the partition block for the chroma component (when ISP is not applied to the chroma component, it is regarded as having one partition block), the check regarding the position of the non-zero coefficient at the end is performed in the above-described manner to determine whether to perform the corresponding LFNST index signaling.
[0267] E. When it is confirmed that a conversion coefficient exists at a position where it should not exist at any position of one partition block among all the partition blocks, the LFNST index signaling can be omitted.
[0268] For example, in the case of 4×4 and 8×8 partition blocks, according to the conversion coefficient scan order in the VVC standard, LFNST conversion coefficients exist at eight positions starting from the DC position, and the remaining positions are all filled with 0. Also, when it is larger than or equal to 4×4 but not a 4×4 or 8×8 partition block, according to the conversion coefficient scan order in the VVC standard, LFNST conversion coefficients exist at 16 positions starting from the DC position, and the remaining positions are all filled with 0.
[0269] Therefore, after performing residual coding, if a non-zero conversion coefficient exists in the area where the 0 value should be filled, the LFNST index signaling can be omitted.
[0270] On the other hand, in ISP mode, the current VVC standard considers the length conditions independently for the horizontal and vertical directions, and applies DST-7 instead of DCT-2 without signaling to the MTS index. It is determined whether the horizontal or vertical length is equal to or greater than 4, and equal to or less than 16, and the primary conversion kernel is determined by the result of this determination. Therefore, in cases where LFNST can be applied while in ISP mode, the following conversion combinations are possible.
[0271] 1. When the LFNST index is 0 (including cases where the LFNST index is inferred to be 0), the determination conditions for the linear transformation when it is an ISP currently included in the VVC standard can be followed. That is, the satisfaction of the length conditions (equal to or greater than 4, and equal to or less than 16) can be checked independently for the horizontal and vertical directions. If the conditions are satisfied, DST-7 can be applied instead of DCT-2 for the linear transformation; otherwise, DCT-2 can be applied.
[0272] 2. When the LFNST index is greater than 0, the following two configurations are possible with a linear transformation:
[0273] A. DCT-2 can be applied to both horizontal and vertical directions.
[0274] B. The determination conditions for primary conversion when it is an ISP currently included in the VVC standard can be followed. That is, the satisfaction of the length conditions (equal to or greater than 4, and equal to or less than 16) can be checked independently for the horizontal and vertical directions. If the conditions are satisfied, DST-7 can be applied instead of DCT-2; otherwise, DCT-2 can be applied.
[0275] In ISP mode, the image information can be configured so that the LFNST index is sent per partition block rather than per coding unit. In such a case, the LFNST index signaling scheme described above can be used to determine whether or not to signal the LFNST index by assuming that there is only one partition block within the unit to which the LFNST index is sent.
[0276] On the other hand, below we will look at an example where LFNST is applied only to the Luma element in the case of a single tree.
[0277] The following table shows the encoding device syntax table related to the signaling of the LFNST index and MTS index for one example.
[0278] [Table 3]
[0279] The meanings of the main variables in the table above are as follows:
[0280] 1. cbWidth, cbHeight: The width and height of the current coding block.
[0281] 2. log2TbWidth, log2TbHeight: Base - 2 log values of the width and height of the current Transform Block. Zero-out is reflected, and non-zero coefficients can exist. This allows the block to be reduced to the upper left section.
[0282] 3. sps_lfnst_enabled_flag: This flag indicates whether LFNST is applicable (enable). A flag value of 0 indicates that LFNST is not applicable, and a flag value of 1 indicates that LFNST is applicable. It is defined in the Sequence Parameter Set (SPS).
[0283] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variable chType and the (x0,y0) position. chType has values of 0 and 1, where 0 indicates a luma element and 1 indicates a chroma element. The (x0,y0) position indicates the position on the picture, and the CuPredMode[chType][x0][y0] value can be MODE_INTRA (intra prediction) or MODE_INTER (inter prediction).
[0284] 5. IntraSubPartitionsSplitType: Indicates whether a certain ISP split was applied to the current encoding device. ISP_NO_SPLIT indicates that the encoding device was not split into partition blocks.
[0285] 6. intra_mip_flag[x0][y0]: The content for the (x0,y0) position is as described in item 4 above. intra_mip_flag is a flag that indicates whether or not the MIP (Matrix-based Intra Prediction) prediction mode has been applied. A flag value of 0 indicates that MIP cannot be applied, and a flag value of 1 indicates that MIP will be applied.
[0286] 7. cIdx: A value of 0 indicates luma, while values of 1 and 2 indicate the chroma elements Cb and Cr, respectively.
[0287] 8. Tree Type: Refers to single-tree, dual-tree, etc. (SINGLE_TREE: single-tree, DUAL_TREE_LUMA: dual-tree for luma components, DUAL_TREE_CHROMA: dual-tree for chroma components)
[0288] 9. lfnst_idx[x0][y0]: It is the LFNST index syntax element to be parsed. If it is not to be parsed, it can be inferred as a value of 0. That is, the default value is set to 0, indicating that LFNST is not applied.
[0289] The descriptions for the above-described syntax elements apply to the syntax elements shown in the following table.
[0290] In Table 3, when transform_skip_flag[x0][y0][0] == 0, it becomes one of the conditions for determining the signaling of the lfnst index based on the presence or absence of transform skip for luma components.
[0291] To this end, following an example, an encoding device syntax table is proposed as follows to remove the dependency between the transform skip flag for luma components and the LFNST index signaling for chroma components and to eliminate the worst-case delay related to LFNST application.
[0292]
Table 4
[0293] In the embodiments shown in Table 4, the signaling of the LFNST index for luma elements depends solely on the conversion skip flag for luma elements for all dual-tree and single-tree split modes. In dual-tree mode, the LFNST index for chroma elements is signaled based solely on the conversion skip flag for chroma elements. In single-tree split mode, LFNST is not applied to chroma elements to mitigate worst-case delays.
[0294] The variable LfnstTransformNotSkipFlag, shown in Table 4, is set by the transformation skip flag value for the current block's tree type and color component. The LFNST index is signaled only when its value is 1.
[0295] The variable LfnstTransformNotSkipFlag is set to 1 if the tree type is not a dual-tree chroma (tree Type != DUAL_TREE_CHROMA), that is, if the tree type is a single-tree or a dual-tree chroma, and the transformation skip flag value for the chroma element is 0. tree Type!=DUAL_TREE_CHROMA?transform_skip_flag[x0][y0][0]==0 :(transform_skip_flag[x0][y0][1]==0||transform_skip_flag[x0][y0][2]==0)), if the tree type is a dual-tree chroma, then when the transform skip flag value (transform_skip_flag[x0][y0][1]) is 0 for the chroma Cb element or 0 for the chroma Cr element, ( tree Type != DUAL_TREE_CHROMA? transform_skip_flag[x0][y0][0]==0: (transform_skip_flag[x0][y0][1]==0||transform_skip_flag[x0][y0][2]==0) ), set to 1.
[0296] In this document, the "x?y:z" operator indicates that if x is true, x is equal to y; otherwise, x is equal to z.
[0297] The following is the specification text for the conversion process, taking Table 4 into consideration.
[0298] [Table 5]
[0299] The variable LfnstZeroOutSigCoeffFlag in Table 3 is 0 if an effective coefficient exists at a position where zero-out occurs when LFNST is applied, and 1 otherwise. The variable LfnstZeroOutSigCoeffFlag is subsequently set by several conditions shown in Table 10.
[0300] The variable LfnstZeroOutSigCoeffFlag indicates whether or not a significant coefficient exists in the second section of the current block, excluding the top-left first section. This value is initially set to 1, and if a significant coefficient exists in the second section, its value is changed to 0. The LFNST index can only be parsed if the initially set value of the variable LfnstZeroOutSigCoeffFlag is maintained at 1. When determining and deriving whether or not the value of the variable LfnstZeroOutSigCoeffFlag is 1, the LFNST is applied to all luma or chroma elements of the current block, so the color index of the current block is not determined.
[0301] For example, the variable LfnstDcOnly in Table 3 becomes 1 if all the last effective coefficients in a transformation block with a corresponding CBF (Coded Block Flag, which is 1 if at least one effective coefficient exists in the block, or 0 if it is 1), and 0 otherwise. More specifically, in the case of a dual-tree luma, the position of the last effective coefficient is checked for one luma transformation block, and in the case of a dual-tree chroma, the position of the last effective coefficient is checked for all transformation blocks for Cb and Cr. In the case of a single tree, the position of the last effective coefficient can be checked for transformation blocks for luma, Cb, and Cr.
[0302] On the other hand, the syntax table for the encoding device that signals the LFNST index for other examples is as follows:
[0303] [Table 6]
[0304] In Table 6, the variable transform_skip_flag[x0][y0][cIdx] indicates whether or not transformation skipping is applied to the coding block for the element indicated by cIdx. cIdx has values of 0, 1, and 2, where 0 refers to a Luma element, and 1 and 2 refer to a Cb element and a Cr element, respectively. A transform_skip_flag[x0][y0][cIdx] value of 1 indicates that transformation skipping will be applied, and a value of 0 indicates that transformation skipping will not be applied.
[0305] In Table 6, the LfnstNotSkipFlag variable is set to 1 only if no conversion skip is applied to any element (all coding blocks) that make up the current encoding device; otherwise, it is set to 0. The LFNST index (lfnst_idx in Table 6) can only be signaled if LfnstNotSkipFlag is 1.
[0306] If the current encoding device is encoded in a single-tree structure (in Table 6, when the tree Type is SINGLE_TREE), all relevant elements consist of Y, Cb, and Cr. If the current encoding device is encoded in a separate tree structure for Luma (in Table 6, when the tree Type is DUAL_TREE_LUMA), all relevant elements consist of Y only. If the current encoding device is encoded in a separate tree structure for Chroma (in Table 6, when the tree Type is DUAL_TREE_CHROMA), all relevant elements consist of Cb and Cr.
[0307] In other words, if even one of the elements constituting the current encoding device is encoded by conversion skipping, the LFNST index will not be signaled, and the corresponding LFNST index value can be inferred to be 0. That is, LFNST will not be applied.
[0308] As shown in Table 6, in a structure that restricts LFNST from being applied if even one of the constituent elements (Y, Cb, Cr) is encoded with a conversion skip, as in the single tree, if several elements (Y, Cb, Cr) are encoded consecutively and during parsing, if an element is determined to be a conversion skip (for example, if Cb and Cr are determined to be conversion skips), the structure can be configured so that the corresponding conversion coefficients are not buffered for the element encoded with the conversion skip until the LFNST index value can be parsed.
[0309] For example, if it is determined that a certain element is encoded by a transformation skip, then it becomes certain that LFNST will not be applied, and inverse quantization, inverse transform, etc., can be performed immediately.
[0310] Instead of using Table 6, you can also describe it more concisely using a syntax table like Table 7.
[0311] [Table 7]
[0312] If, in the case of a single tree, the presence or absence of signaling in the LFNST index is determined by checking only whether or not there are conversion skips for the Luma element, then a syntax table for the encoding device can be constructed as follows.
[0313] [Table 8]
[0314] In Table 8, LfnstNotSkipFlag is determined the same way as in Table 6 or Table 7 if it is a separate tree (i.e., if it is a luma-separated tree, it is set to 1 if no conversion skip is applied to luma elements, and 0 otherwise; if it is a chroma-separated tree, it is set to 1 if no conversion skip is applied to any Cb and Cr elements, and 0 otherwise); and in the case of a single tree, it is set to 1 if no conversion skip is applied to luma elements, and 0 otherwise. Alternatively, a syntax table like Table 9 can be applied instead of Table 8.
[0315] [Table 9]
[0316] On the other hand, in the case of a single tree, the following is an example of how to determine the LFNST index signaling conditions when applying LFNST only to Luma elements.
[0317] Tables 3 and 4, and Tables 6, 7, 8, and 9, show that the LfnstDcOnly and LfnstZeroOutSigCoeffFlag variables are used as conditions for signaling the LFNST index. Basically, as shown in Table 6, the LfnstDcOnly and LfnstZeroOutSigCoeffFlag variables are all initialized to 1, and in the syntax table for residual coding as shown in Table 10, the values of the two corresponding variables are updated to 0. For reference, if an element is encoded with a transformation skip (to be Y, Cb, or Cr), another syntax table (transform_ts_coding) is called instead of the residual coding in Table 10, and the LfnstDcOnly and LfnstZeroOutSigCoeffFlag variables are not updated when LFNST index parsing is performed on the corresponding element.
[0318] [Table 10-1]
[0319] [Table 10-2]
[0320] In Table 10, lastSubBlock indicates the scan order position of the sub-block (Coefficient Group (CG)) where the last non-zero coefficient is located. 0 indicates a sub-block containing a DC element, while a value greater than 0 indicates a sub-block that does not contain a DC element.
[0321] `lastScanPos` indicates the position of the last effective coefficient in the scan order within a single subblock. If a subblock consists of 16 positions, the possible values are from 0 to 15.
[0322] LastSignificantCoeffX and LastSignificantCoeffY indicate the x and y coordinates in which the last significant coefficient is located within the transformation block. The x-coordinate starts at 0 and increases to the left and to the right, while the y-coordinate starts at 0 and increases from top to bottom. If the values of both variables are 0, it means that the last significant coefficient is located at DC.
[0323] For the examples in Table 4, Table 10 is basically applied to determine the LfnstDcOnly variable value and the LfnstZeroOutSigCoeffFlag variable value. When Table 10 is applied and encoded into a single tree, the residual coding presented in Table 10 is called for all elements. For example, if all Y, Cb, and Cr elements are not encoded into a transformation skip, residual coding is performed for each element.
[0324] Therefore, when Table 10 is applied and encoded into a single tree, if the last non-zero coefficient of any element is located anywhere other than the DC position (the top-left position of the corresponding transformation block), the LfnstDcOnly variable is updated to 0. If the last non-zero coefficient of any element is located in a section where no transformation coefficients can be located when LFNST is applied (i.e., in the current VVC standard, if it is a 4x4 or 8x8 transformation block, it is located in a section other than the 1st to 8th position by the forward transformation coefficient scan order, or if it is a transformation block to which LFNST can be applied, it is located in a section other than the top-left 4x4 section), the LfnstZeroOutSigCoeffFlag variable value is updated to 0. As shown in the encoding device syntax table in Table 4 and in Tables 6, 7, 8, and 9, the LFNST index is signaled only when the LfnstZeroOutSigCoeffFlag variable value is 1, and unless in ISP mode, the LFNST index is signaled only when the LfnstDcOnly variable value is 0.
[0325] However, in a single tree, if LFNST is applied only to chroma elements, it is possible to restrict the updating of the LfnstDcOnly and LfnstZeroOutSigCoeffFlag variables when performing residual coding on elements to which LFNST is not applied (chroma elements, Cb, or Cr). This is because the arrangement or distribution of transformation coefficients for elements to which LFNST is not applied may not logically determine whether or not the LFNST index is signaling, i.e., whether or not LFNST is applied.
[0326] Table 11 restricts the updating of the LfnstDcOnly and LfnstZeroOutSigCoeffFlag variables to only Luma elements in the case of a single tree. In the case of a non-single tree, the LfnstDcOnly and LfnstZeroOutSigCoeffFlag variables can be updated for all elements, as in Table 10.
[0327] [Table 11-1]
[0328] [Table 11-2]
[0329] In the residual coding presented in Table 11, the tree type has been added as a parameter compared to Table 10, so the syntax table for the conversion device is modified as shown in Table 12.
[0330] [Table 12]
[0331] Based on the contents of Table 4, some of the contents can be replaced with the examples in Tables 6 to 9, or the contents of Table 10 or Table 11 can be applied. Based on Tables 6 to 9 and Table 10 or Table 11, the following combinations are possible.
[0332] 1. Table 6 (or Table 7) + Table 10
[0333] 2. Table 6 (or Table 7) + Table 11
[0334] 3. Table 8 (or Table 9) + Table 10
[0335] 4. Table 8 (or Table 9) + Table 11
[0336] Below, we will examine how to apply a scaling list to chroma elements when LFNST is applied only to luma elements in a single tree.
[0337] Currently, VVC WD has a syntax element called `scaling_matrix_for_lfnst_disabled_flag`. If `scaling_matrix_for_lfnst_disabled_flag` is 1, the scaling list will not be applied when LFNST is applied; if it is 0, the scaling list can be applied when LFNST is applied.
[0338] Here, the scaling list is a matrix that specifies a specific weight value for each position of the transformation coefficient in the transformation block. By multiplying each transformation coefficient by the corresponding weight value, inverse quantization or quantization can be performed, and inverse quantization or quantization can be applied by subtracting according to the importance of the transformation coefficients.
[0339] In the case of a single tree, LFNST can only be applied to luma elements as shown in the example in Table 4. If the `scaling_matrix_for_lfnst_disabled_flag` value is 1 and LFNST is applied when encoded as a single tree, the scaling list is not applied to luma elements. In this case, the scaling list can be applied to chroma elements to which LFNST is not applied.
[0340] Table 13 shows an example of an inverse quantum and process (scaling process) that can implement the above-mentioned case.
[0341] [Table 13-1]
[0342] [Table 13-2]
[0343] In Table 13, `tree Type` indicates the tree type of the encoding device to which the currently processed transformation block belongs, with `SINGLE_TREE`, `DUAL_TREE_LUMA`, and `DUAL_TREE_CHROMA` representing the single tree, the separated tree for lumae, and the separated tree for chromae (dual-tree chroma), respectively.
[0344] In this embodiment, when it is a single tree, LFNST is applied only to luma elements. Therefore, when the scaling_matrix_for_lfnst_disabled_flag value is 1 and LFNST is applied, the scaling list is not applied to luma elements (corresponding to when the cIdx value is 0) (when the lfnst_idx[xTbY][yTbY] value is greater than 0).
[0345] On the other hand, for chroma elements (where the cIdx value is greater than 0), other conditions can be checked further (for example, by checking transform_skip_flag[xTbY][yTbY][cIdx]) to determine whether or not the scaling list is applied.
[0346] In the case of a separate tree, as with luma elements in a single tree, when the scaling_matrix_for_lfnst_disabled_flag value is 1 and LFNST is applied, the scaling list is not applied to luma and chroma elements (when the lfnst_idx[xTbY][yTbY] value is greater than 0).
[0347] Alternatively, in the case of a separate tree, other conditions can be checked (for example, by checking transform_skip_flag[xTbY][yTbY][cIdx]) to determine whether or not to apply the scaling list, similar to the case of chroma elements in a single tree.
[0348] Therefore, if the `scaling_matrix_for_lfnst_disabled_flag` value is 1, and LFNST can only be applied to luma elements in a single tree, then the scaling list will not be applied to luma elements, while it will be applied to chroma elements.
[0349] For example, a combination of Table 13 and the above-mentioned embodiment (a combination based on the contents of Table 4, with some contents replaced by the embodiments in Tables 6 to 9, or by applying the contents of Table 10 or Table 11) may be applied.
[0350] In this case, as per the specification text for "Transform process for scaled transform coefficients" in Table 5, LFNST can be configured to apply only to Luma elements in the case of a single tree.
[0351] The following drawings have been prepared to illustrate a specific example of this specification. The names of specific devices and signals / messages / fields shown in the drawings are presented as examples only, and the technical features of this specification are not limited to the specific names used in the following drawings.
[0352] Figure 12 is a flowchart illustrating the operation of a video decoding device according to one embodiment described in this document.
[0353] Each step disclosed in Figure 12 is based in part on the content described above in Figures 1 to 11. Therefore, in Figures 1 to 11, specific details that overlap with the content described above will be omitted or simplified.
[0354] In one embodiment, the decoding device 200 can receive flag information indicating the availability of the scaling list, the LFNST index for the current block, and residual information when LFNST is performed from the bitstream (S1210).
[0355] More specifically, the decoding device 200 can decode information about the quantized transformation coefficients for the current block from the bitstream, and can derive the quantized transformation coefficients for the target block based on the information about the quantized transformation coefficients for the current block. The information about the quantized transformation coefficients for the target block can be included in an SPS (Sequence Parameter Set) or a slice header, and may include at least one of the following: information on whether or not a simplified transformation (RST) is applied, information on the simplification factor, information on the minimum transformation size to apply the simplified transformation, information on the maximum transformation size to apply the simplified transformation, the inverse simplified transformation size, and information on a transformation index that points to any one of the transformation kernel matrices included in the transformation set.
[0356] Furthermore, the decoding device can receive information regarding the intra-predictive mode for the current block and information regarding whether or not ISP is applied to the current block. By receiving and parsing flag information indicating whether or not ISP encoding or ISP mode is applied, the decoding device can derive whether or not the current block is divided into a predetermined number of subpartitioned blocks. Here, the current block is the encoded block. The decoding device can also derive the size and number of subpartitioned blocks to be divided via flag information indicating the direction in which the current block is divided.
[0357] The LFNST index has values from 0 to 2 to specify the LFNST matrix when LFNST is applied to an inverse quadratic inseparable matrix. For example, an LFNST index value of 0 indicates that LFNST is not applied to the current block, an LFNST index value of 1 can indicate the first LFNST matrix, and an LFNST index value of 2 can indicate the second LFNST matrix.
[0358] ISP-related information and the LFNST index are received at the encoding device level.
[0359] When the decoding device receives and executes LFNST, flag information indicating the availability of the scaling list can be shown in `scaling_matrix_for_lfnst_disabled_flag` or `sps_scaling_matrix_for_lfnst_disabled_flag`, and is signaled in the sequence parameter set. If this flag value is 1, it indicates that the scaling list will not be applied when LFNST is applied, and if it is 0, it indicates that the scaling list can be applied when LFNST is applied. The scaling list is a matrix that specifies a specific weight value for each position of the transformation coefficient in the transformation block. By multiplying each transformation coefficient by the corresponding weight value, inverse quantization or quantization can be performed, and inverse quantization or quantization can be applied by subtracting according to the importance of the transformation coefficient.
[0360] The decoding device 200 can determine whether a scaling list is applied to the current block based on the flag information, the LFNST index, and the tree type of the current block in order to perform inverse quantization on the conversion coefficients for the current block (S1220).
[0361] If the current block's tree type is single-tree, the current block's color components can include a luma element, a first chroma element indicating chroma Cb, and a second chroma element indicating chroma Cr. If the current block's tree type is dual-tree luma, the current block can include a luma element. If the current block's tree type is dual-tree chroma, the current block's color components can include a first chroma element and a second chroma element.
[0362] Here, the current block is a transformation block, which is a transformation unit. If the tree type of the current block is a single tree, it can contain transformation blocks for luma elements, transformation blocks for the first chroma element, and transformation blocks for the second chroma element. If the tree type of the current block is a dual-tree luma, it can contain transformation blocks for luma elements. If the tree type of the current block is a dual-tree chroma, it can contain transformation blocks for the first and second chroma elements.
[0363] For example, if the current block is a single tree, LFNST can only be applied to luma elements. If it is encoded as a single tree and the scaling_matrix_for_lfnst_disabled_flag value is 1, and LFNST is applied, the scaling list will not be applied to luma elements. However, the scaling list can be applied to chroma elements to which LFNST is not applied.
[0364] In summary, if the flag information for the scaling list indicates that the scaling list is not available and the LFNST index is greater than 0 (i.e., LFNST is applicable), then if the current block's tree type is a single tree and it contains chroma elements, the scaling list will not be applied; however, if the current block's tree type is a single tree and it contains chroma elements, the scaling list can be applied.
[0365] For example, if the flag information indicates that the scaling list is not available and the LFNST index is greater than 0, then if the tree type of the current block is a dual-tree chroma, the LFNST can be applied to the current block and therefore the scaling list is not applied to the chroma elements.
[0366] For example, if the flag information indicates that the scaling list is not available and the LFNST index is greater than 0, then if the tree type of the current block is a dual-tree luma, then LFNST can be applied to the current block, and therefore the scaling list is not applied to the luma elements.
[0367] Subsequently, the decoding device derives a conversion coefficient for the current block from the residual information based on the judgment result (S1230).
[0368] The derived transformation coefficients can be arranged in 4x4 block units according to the reverse diagonal scan order, and the transformation coefficients within each 4x4 block are also arranged according to the reverse diagonal scan order. In other words, the transformation coefficients after inverse quantization are arranged according to the reverse scan order applied in video codecs such as VVC and HEVC.
[0369] The decoding device can derive modified transformation coefficients from transformation coefficients based on the LFNST index and the LFNST matrix for LFNST, that is, by applying LFNST (S1240).
[0370] Unlike linear transformations, which separate the coefficients to be transformed vertically or horizontally, LFNST is a non-separated transformation that applies the transformation without separating the coefficients in a specific direction. Such a non-separated transformation is a low-frequency non-separated transformation that applies the forward transformation only to low-frequency sections, not to entire block sections.
[0371] The decoding device can derive various variables to apply LFNST and can determine whether or not LFNST should be applied based on the current block's tree type and size, among other things.
[0372] The decoding device can derive a first variable (variable LfnstDcOnly) indicating whether or not an effective coefficient exists at a position other than the DC element of the current block, and a second variable (variable LfnstZeroOutSigCoeffFlag) indicating whether or not the conversion coefficient exists in the second section excluding the first section in the upper left of the current block.
[0373] These first and second variables are initially set to 1. If an effective coefficient exists in a position other than the DC element of the current block, the first variable is updated to 0, and if a conversion coefficient exists in the second section, the second variable is updated to 0.
[0374] If the first variable is updated to 0 and the second variable remains at 1, then LFNST is applied to the current block.
[0375] On the other hand, for Luma blocks to which intra-subpartition (ISP) mode can be applied, the LFNST index can be parsed without deriving the variable LfnstDcOnly.
[0376] Specifically, when ISP mode is applied and the transform_skip_flag[x0][y0][0] value for Luma elements is 0, the LFNST index is signaled regardless of the LfnstDcOnly variable value when the tree type of the current block is a single tree or a dual tree for Luma.
[0377] On the other hand, for chroma elements to which ISP mode is not applied, the value of the variable LfnstDcOnly can be set to 0 by the transform_skip_flag[x0][y0][1], which is a transformation skip flag for chroma Cb elements, and the transform_skip_flag[x0][y0][2], which is a transformation skip flag for chroma Cr elements. In other words, in transform_skip_flag[x0][y0][cIdx], when the cIdx value is 1, the variable LfnstDcOnly can be set to 0 only when the transform_skip_flag[x0][y0][1] value is 0, and when the cIdx value is 2, the variable LfnstDcOnly can be set to 0 only when the transform_skip_flag[x0][y0][2] value is 0. If the value of the variable LfnstDcOnly is 0, the decoding device can parse the LFNST index, and other LFNST indices are not signaled and can be inferred to have a value of 0.
[0378] The second variable is LfnstZeroOutSigCoeffFlag, which indicates whether zero-out was performed when LFNST was applied. The second variable is initially set to 1 and changes to 0 if effective coefficients exist in the second section.
[0379] The variable LfnstZeroOutSigCoeffFlag is derived to 0 if the index of the subblock containing the last non-zero coefficient is greater than 0, the width and height of the transform block are all equal to or greater than 4, the last position of the non-zero coefficient within the subblock containing the last non-zero coefficient is greater than 7, and the size of the transform block is 4x4 or 8x8. A subblock refers to a 4x4 block used as the coding unit in residual coding, and can also be called a CG (Coefficient Group). An index of 0 in a subblock refers to the top-left 4x4 subblock.
[0380] In other words, in a transformation block, if a non-zero coefficient is derived in any section other than the top-left section where an LFNST transformation coefficient can exist, or if a non-zero coefficient exists in a 4x4 block or 8x8 block that is located away from the 8th position in the scan order, the variable LfnstZeroOutSigCoeffFlag is set to 0.
[0381] The decoding device determines an LFNST set, including an LFNST matrix, based on the intra-prediction mode derived from the intra-prediction mode information, and can select one of several LFNST matrices based on the LFNST set and the LFNST index.
[0382] In this case, the same LFNST set and the same LFNST index are applied to the subpartitioned transform blocks that were divided in the current block. That is, since the same intra-prediction mode is applied to the subpartitioned transform blocks, the LFNST set determined based on the intra-prediction mode is also applied in the same way to all subpartitioned transform blocks. Furthermore, since the LFNST index is signaled at the encoding device level, the same LFNST matrix is applied to the subpartitioned transform blocks that were divided in the current block.
[0383] On the other hand, as mentioned above, the transformation set is determined by the intra-prediction mode of the transformation block to be transformed, and the inverse LFNST is performed based on one of the transformation kernel matrices, i.e., LFNST matrices, included in the transformation set indicated by the LFNST index. The matrix applied to the inverse LFNST is called the inverse LFNST matrix or LFNST matrix, and the name does not matter as long as such a matrix is in a transform relationship with the matrix used for the forward LFNST.
[0384] In one example, the inverse LFNST matrix is a non-square matrix in which the number of columns is less than the number of rows.
[0385] The decoding device can derive the residual sample for the current block based on a linear inverse transform of the corrected transformation coefficients (S1250).
[0386] In this case, the inverse linear transformation is performed using the standard separable transform, and the MTS described above may also be used.
[0387] Subsequently, the decoding device 200 can generate a reconstructed sample based on the residual sample for the current block and the predicted sample for the current block.
[0388] The following drawings have been prepared to illustrate a specific example of this specification. The names of specific devices and signals / messages / fields shown in the drawings are presented as examples only, and the technical features of this specification are not limited to the specific names used in the following drawings.
[0389] Figure 13 is a flowchart illustrating the operation of a video encoding device according to one embodiment described in this document.
[0390] Each step disclosed in Figure 13 is based in part on the content described above in Figures 4 to 11. Therefore, in Figures 2 and 4 to 11, specific details that overlap with the content described above will be omitted or simplified.
[0391] An encoding device 100 according to one embodiment can derive predicted samples for the current block based on an intra-prediction mode applied to the current block.
[0392] The encoding device can perform predictions on a per-subpartition conversion block basis if an ISP is applied to the current block.
[0393] The encoding device can determine whether or not to apply ISP encoding or ISP mode to the current block, i.e., the encoding block. Based on this determination, it can determine in which direction the current block will be divided and derive the size and number of subblocks to be divided.
[0394] The same intra-prediction mode is applied to the subpartitioned transform blocks divided within the current block, and the encoding device can derive prediction samples for each subpartitioned transform block. That is, the encoding device performs intra-prediction sequentially depending on the division pattern of the subpartitioned transform block, for example, horizontally or vertically, from left to right, or from top to bottom. For the leftmost or topmost subblock, it refers to the reconstructed pixels of the already encoded block, as in the normal intra-prediction method. Furthermore, for each edge of a subsequent internal subpartitioned transform block that is not adjacent to a previously encoded subpartitioned transform block, it refers to the reconstructed pixels of the adjacent encoded block, as in the normal intra-prediction method, in order to derive the reference pixels adjacent to that edge.
[0395] The encoding device 100 can derive residual samples for the current block based on the predicted samples (S1310).
[0396] The encoding device 100 can apply at least one of LFNST or MTS to the residual sample to derive conversion coefficients for the current block, and arrange the conversion coefficients in a predetermined scan order.
[0397] The encoding device can derive conversion coefficients for the current block based on a conversion process such as a linear and / or quadratic conversion of the residual sample (S1320).
[0398] The first-order transformation can be performed via multiple transformation kernels, as in MTS, in which case the transformation kernel is selected based on the intra-predictive mode.
[0399] The encoding device 100 can determine whether to perform a quadratic transformation or a non-separable transformation, specifically LFNST, on the transformation coefficients for the current block, and can derive modified transformation coefficients by applying LFNST to the transformation coefficients.
[0400] Unlike linear transformations, which separate the coefficients to be transformed vertically or horizontally, LFNST is a non-separated transformation that applies the transformation without separating the coefficients in a specific direction. Such a non-separated transformation is a low-frequency non-separated transformation that applies the transformation only to the low-frequency section, not to the entire target block.
[0401] The encoding device can derive various variables to apply LFNST and can determine whether or not LFNST should be applied based on the current block's tree type and size, among other things.
[0402] The encoding device can derive a first variable (variable LfnstDcOnly) indicating whether or not an effective coefficient exists at a position other than the DC element of the current block, and a second variable (variable LfnstZeroOutSigCoeffFlag) indicating whether or not the conversion coefficient exists in the second section excluding the first section in the upper left of the current block.
[0403] These first and second variables are initially set to 1. If an effective coefficient exists in a position other than the DC element of the current block, the first variable is updated to 0, and if a conversion coefficient exists in the second section, the second variable is updated to 0.
[0404] LFNST is applied to the current block if the first variable is updated to 0 and the second variable remains at 1.
[0405] On the other hand, in the case of a Luma block to which intra-subpartition (ISP) mode can be applied, LFNST can be applied without deriving the variable LfnstDcOnly.
[0406] Specifically, when ISP mode is applied and the transform_skip_flag[x0][y0][0] value for the Luma element is 0, LFNST is applied regardless of the LfnstDcOnly variable value when the tree type of the current block is a single tree or a dual tree for Luma.
[0407] On the other hand, for chroma elements to which ISP mode is not applied, the value of the variable LfnstDcOnly can be set to 0 by the transform_skip_flag[x0][y0][1], which is a transformation skip flag for chroma Cb elements, and the transform_skip_flag[x0][y0][2], which is a transformation skip flag for chroma Cr elements. In other words, in transform_skip_flag[x0][y0][cIdx], when the cIdx value is 1, the variable LfnstDcOnly can be set to 0 only when the transform_skip_flag[x0][y0][1] value is 0, and when the cIdx value is 2, the variable LfnstDcOnly can be set to 0 only when the transform_skip_flag[x0][y0][2] value is 0. If the variable LfnstDcOnly is 0, the encoding device can apply LFNST, but may not apply other LFNSTs.
[0408] The second variable is LfnstZeroOutSigCoeffFlag, which indicates whether zero-out was performed when LFNST was applied. The second variable is initially set to 1 and changes to 0 if effective coefficients exist in the second section.
[0409] The variable LfnstZeroOutSigCoeffFlag is derived to 0 if the index of the subblock containing the last non-zero coefficient is greater than 0 and the width and height of the transform block are all equal to or greater than 4, or if the last position of the non-zero coefficient within the subblock containing the last non-zero coefficient is greater than 7 and the position of the transform block is 4x4 or 8x8. A subblock refers to a 4x4 block used as the coding unit in residual coding, and can also be called a CG (Coefficient Group). An index of 0 in a subblock refers to the top-left 4x4 subblock.
[0410] In other words, in a transformation block, if a non-zero coefficient is derived in any section other than the top-left section where an LFNST transformation coefficient can exist, or if a non-zero coefficient exists in a 4x4 block or 8x8 block that is located away from the 8th position in the scan order, the variable LfnstZeroOutSigCoeffFlag is set to 0.
[0411] The encoding device determines an LFNST set, including an LFNST matrix, based on the intra-prediction mode derived from the intra-prediction mode information, and can select one of several LFNST matrices.
[0412] In this case, the same LFNST set and the same LFNST index are applied to the subpartitioned transform blocks that were divided in the current block. That is, since the same intra-prediction mode is applied to the subpartitioned transform blocks, the LFNST set determined based on the intra-prediction mode is also applied in the same way to all subpartitioned transform blocks. Furthermore, since the LFNST index is signaled at the encoding device level, the same LFNST matrix is applied to the subpartitioned transform blocks that were divided in the current block.
[0413] On the other hand, as mentioned above, the transformation set is determined by the intra-prediction mode of the transformation block to be transformed, and LFNST is performed based on one of the transformation kernel matrices included in the LFNST transformation set, i.e., the LFNST matrices. The matrix applied to LFNST is called the LFNST matrix, and the name of such a matrix does not matter as long as it is in a transform relationship with the matrix used for inverse LFNST.
[0414] In one example, the LFNST matrix is a non-square matrix in which the number of rows is less than the number of columns.
[0415] The encoding device can determine whether LFNST is performed in the transformation process to quantize the transformation coefficients based on the scaling list, and whether the scaling list is applied to the current block based on the tree type of the current block (S1330).
[0416] The scaling list is a matrix that specifies a specific weight value for each position of the transformation coefficient in the transformation block. By multiplying each transformation coefficient by the corresponding weight value, inverse quantization or quantization can be performed, and inverse quantization or quantization can be applied by subtracting based on the importance of the transformation coefficients.
[0417] For example, the encoding device can choose not to apply the scaling list if the current block's tree type is a single tree and contains chroma elements, but can apply the scaling list if the current block's tree type is a single tree and contains chroma elements.
[0418] If the current block's tree type is single-tree, the current block's color components can include a luma element, a first chroma element indicating chroma Cb, and a second chroma element indicating chroma Cr. If the current block's tree type is dual-tree luma, the current block can include a luma element. If the current block's tree type is dual-tree chroma, the current block's color components can include a first chroma element and a second chroma element.
[0419] Here, the current block is a transformation block, which is a transformation unit. If the tree type of the current block is a single tree, it can contain transformation blocks for luma elements, transformation blocks for the first chroma element, and transformation blocks for the second chroma element. If the tree type of the current block is a dual-tree luma, it can contain transformation blocks for luma elements. If the tree type of the current block is a dual-tree chroma, it can contain transformation blocks for the first and second chroma elements.
[0420] For example, if the current block is a single tree, the encoding device can apply LFNST only to luma elements, and if LFNST is applied, it does not apply the scaling list to luma elements. However, it can apply the scaling list to chroma elements to which LFNST is not applied.
[0421] In summary, if the encoding device has an LFNST index greater than 0 (i.e., LFNST is applied), it will not apply the scaling list if the current block's tree type is a single tree and it contains chroma elements, but it can apply the scaling list if the current block's tree type is a single tree and it contains chroma elements.
[0422] For example, if the encoding device has an LFNST index greater than 0, and the current block's tree type is a dual-tree chroma, then LFNST can be applied to the current block and the scaling list will not be applied to the chroma elements.
[0423] For example, if the encoding device has an LFNST index greater than 0, and the tree type of the current block is a dual-tree luma, then LFNST can be applied to the current block and the scaling list will not be applied to the luma elements.
[0424] The encoding device can quantize the conversion coefficients based on the aforementioned decision, that is, whether or not the scaling list is applied to the current block (S1340).
[0425] In other words, the encoding device can quantize the conversion coefficients using a scaling list for conversion blocks to which LFNST is not applied, and quantize the conversion coefficients without using a scaling list for conversion blocks to which LFNST is applied.
[0426] The encoding device can encode and output residual information and flag information indicating the availability of the scaling list when LFNST is performed (S1350).
[0427] When LFNST is executed, the flag information indicating the availability of the scaling list can be indicated as `scaling_matrix_for_lfnst_disabled_flag` or `sps_scaling_matrix_for_lfnst_disabled_flag`, and is signaled in the sequence parameter set. A flag value of 1 indicates that the scaling list will not be applied when LFNST is applied, and a value of 0 indicates that the scaling list will be applied when LFNST is applied.
[0428] If the encoding device has an LFNST index greater than 0 and the current block is a single tree, LFNST is applied to the Luma element, allowing the flag value to be encoded to 1.
[0429] However, when the encoding device has an LFNST index greater than 0 and the current block is a single tree, LFNST is not applied to the chroma elements, so the image information can be configured so that a scaling list is applied.
[0430] If the encoding device has an LFNST index greater than 0 and the current block's tree type is a dual-tree chroma, then LFNST can be applied to the current block, and the flag value can be encoded to 1 so that the scaling list is not applied to the chroma elements.
[0431] For example, if the encoding device has an LFNST index greater than 0 and the current block's tree type is a dual-tree luma, then LFNST can be applied to the current block, and the flag value can be encoded to 1 so that the luma elements are not subjected to a scaling list.
[0432] The encoding device can perform quantization based on the modified transformation coefficients for the current block, derive the quantized transformation coefficients, and encode the LFNST index.
[0433] The encoding device can generate residual information containing information about the quantized conversion coefficients. This residual information may include the aforementioned conversion-related information / syntax elements. The encoding device can encode the image / video information containing the residual information and output it in the form of a bitstream.
[0434] More specifically, the encoding device 100 can generate information about quantized conversion coefficients and encode the generated information about quantized conversion coefficients.
[0435] The syntax element of the LFNST index according to this embodiment may indicate whether (reverse) LFNST is applied and which of the LFNST matrices included in the LFNST set is used. If the LFNST set includes two transformation kernel matrices, there may be three possible values for the syntax element of the LFNST index.
[0436] For example, if the current partitioning tree structure for a block is of the dual-tree type, then an LFNST index can be encoded for both the luma block and the chroma block.
[0437] In one embodiment, the syntax element value for the transformation index can be derived as 0, indicating that (reverse) LFNST is not applied to the current block; 1, indicating the first LFNST matrix among the LFNST matrices; and 2, indicating the second LFNST matrix among the LFNST matrices.
[0438] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficient may be called a transformation coefficient. If transformation / inverse transformation is omitted, the transformation coefficient may also be called a coefficient or residual coefficient, or may still be called a transformation coefficient for consistency of expression.
[0439] Furthermore, in this document, quantized transformation coefficients and transformation coefficients may be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients, and such information about the transformation coefficients may be signaled via residual coding syntax. Based on the residual information (or information about the transformation coefficients), transformation coefficients can be derived, and scaled transformation coefficients can be derived via inverse transformation (scaling) of the transformation coefficients. Based on inverse transformation (transformation) of the scaled transformation coefficients, residual samples can be derived. This can be similarly applied / expressed in other parts of this document.
[0440] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks; however, this document is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of this document.
[0441] The method described in this document above can be implemented in software form, and the encoding and / or decoding device described in this document may be included in, for example, image processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.
[0442] In this document, when embodiments are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0443] Furthermore, decoding and encoding devices to which this document applies may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction equipment, real-time communication equipment such as video communications, mobile streaming equipment, storage media, camcorders, customized video (VoD) service providers, OTT video (Over the Top Video) equipment, internet streaming service providers, 3D video equipment, image telephone video equipment, and medical video equipment, and may be used to process video signals or data signals. For example, OTT video (Over the Top Video) equipment may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.
[0444] Furthermore, the processing methods to which this document applies can be produced in the form of programs executed by a computer and stored on a computer-readable recording medium. Multimedia data having the data structure relating to this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), general-purpose serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored on computer-readable recording media or transmitted over a wireless network. Furthermore, embodiments of this document can be embodied in computer program products in the form of program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.
[0445] Figure 14 schematically shows an example of a video / image coding system to which this document can be applied.
[0446] Referring to Figure 14, a video / image coding system may include a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device in the form of a file or streaming via a digital storage medium or network.
[0447] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may consist of a separate device or external component.
[0448] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.
[0449] An encoding device can encode input video / images. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0450] The transmitting unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0451] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.
[0452] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0453] Figure 15 illustrates the structure of a content streaming system to which this document applies.
[0454] Furthermore, the content streaming system to which this document applies may broadly include encoding servers, streaming servers, web servers, media storage, user devices, and multimedia input devices.
[0455] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting it to the streaming server. As an alternative example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted. The bitstream can be generated by the encoding method or bitstream generation method to which this document applies, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving it.
[0456] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, which then transmits multimedia data to the user. The content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.
[0457] The streaming server can receive content from media storage and / or encoding servers. For example, if it receives content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0458] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage. Each server in the content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.
[0459] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined to embody an apparatus, and the technical features of the apparatus claims herein can be combined to embody a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody a method.
Claims
1. A step of determining whether a scaling list is applied to the current block based on the tree type of the current block, Based on the above decision, the steps are to derive a conversion coefficient for the current block from the residual information, The step includes deriving a residual sample based on the inverse transformation of the transformation coefficients, Based on the fact that the tree type of the current block is a single tree, the scaling list does not apply to the luma component of the current block. Based on the fact that the tree type of the current block is the single tree, the scaling list is applied to the chroma component of the current block. Based on the fact that the tree type of the current block is a dual-tree luma and the LFNST (low frequency non-separable transform) index of the current block is greater than 0, the scaling list is not applied to the luma component of the current block. A method in which, based on the fact that the tree type of the current block is a dual-tree chroma and the LFNST index is greater than 0, the scaling list is not applied to the chroma component of the current block.
2. The method according to claim 1, wherein LFNST is not applied to the chromatic component of the current block, based on the fact that the tree type of the current block is the single tree.
3. The present block includes a conversion block, according to claim 1.
4. The steps include: deriving a conversion coefficient for the current block based on the conversion process for the residual sample of the current block; The steps include: quantizing the transformation coefficients based on whether the scaling list applies to the current block; Whether the scaling list applies to the current block is determined based on the tree type of the current block. Based on the fact that the tree type of the current block is a single tree, the scaling list does not apply to the luma component of the current block. Based on the fact that the tree type of the current block is the single tree, the scaling list is applied to the chroma component of the current block. Based on the fact that the tree type of the current block is a dual-tree luma and that an LFNST (low frequency non-separable transform) is performed on the current block, the scaling list is not applied to the luma component of the current block. A method in which, based on the fact that the tree type of the current block is a dual-tree chroma and that the LFNST is performed on the current block, the scaling list is not applied to the chroma component of the current block.
5. The method of claim 4, wherein, based on the fact that the tree type of the current block is the single tree, the LFNST is not performed on the chroma component of the current block.
6. The method according to claim 4, wherein the current block includes a conversion block.
7. A step of generating a bitstream relating to image information, wherein the bitstream is The steps include: deriving a conversion coefficient for the current block based on the conversion process for the residual sample of the current block; A step of quantizing the transformation coefficients based on whether the scaling list is applied to the current block, A step of encoding residual information related to the conversion coefficient, and a step of generating based on, The step of transmitting data including the bitstream, Whether the scaling list applies to the current block is determined based on the tree type of the current block. Based on the fact that the tree type of the current block is a single tree, the scaling list does not apply to the luma component of the current block. Based on the fact that the tree type of the current block is the single tree, the scaling list is applied to the chroma component of the current block. Based on the fact that the tree type of the current block is a dual-tree luma and that an LFNST (low frequency non-separable transform) is performed on the current block, the scaling list is not applied to the luma component of the current block. A method in which, based on the fact that the tree type of the current block is a dual-tree chroma and that the LFNST is performed on the current block, the scaling list is not applied to the chroma component of the current block.