Video coding method and apparatus based on conversion
The video coding method and apparatus address the challenge of high-resolution video compression by using LFNST and MTS to enhance coding efficiency, particularly for LFNST indices, resulting in improved compression and storage efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-08-21
- Publication Date
- 2026-07-29
AI Technical Summary
The increasing demand for high-resolution and high-quality video content, including immersive media, has led to higher transmission and storage costs due to the increased amount of information, necessitating highly efficient video compression technologies.
A video coding method and apparatus that utilizes LFNST and MTS to improve the efficiency of video coding, particularly through the coding of LFNST indices, by deriving context information based on tree types and applying truncated unary codes to syntax elements.
Enhances overall video compression efficiency, improves LFNST index coding, and increases the efficiency of secondary conversion processes.
Smart Images

Figure 0007897403000032 
Figure 0007897403000033 
Figure 0007897403000034
Abstract
Description
[Technical Field]
[0001] This document relates to video coding technology, and more specifically, to a video coding method and apparatus based on transformation in a video coding system. [Background technology]
[0002] Recently, demand for high-resolution, high-quality video / video, such as 4K or 8K or higher UHD (Ultra High Definition) video / video, has been increasing in various fields. As video / video data becomes higher resolution and higher quality, the amount of information or bits transmitted relative to existing video / video data increases. Therefore, when transmitting video / video data using existing wired or wireless broadband lines, or storing video / video data using existing storage media, transmission and storage costs increase.
[0003] Furthermore, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality), and holograms have recently increased, and there has been a rise in broadcasting of video content with different visual characteristics from real-world footage, such as game footage.
[0004] Therefore, highly efficient video compression technology is required to effectively compress, transmit, store, and play back high-resolution, high-quality video information with the diverse characteristics described above. [Overview of the project] [Problems that the invention aims to solve]
[0005] The technical objective of this document is to provide a method and apparatus for improving the efficiency of video coding.
[0006] Another technical objective of this paper is to provide a method and apparatus for improving the efficiency of LFNST index coding.
[0007] Another technical objective of this paper is to provide a method and apparatus for improving the efficiency of quadratic conversion through the coding of LFNST indices. [Means for solving the problem]
[0008] According to one embodiment of this document, a video decoding method is provided that is performed by a decoding device. The method includes the steps of: applying at least one of LFNST or MTS to a conversion coefficient to derive a residual sample; and generating a restored picture based on the residual sample, wherein the LFNST is performed based on an LFNST conversion set, an LFNST kernel included in the LFNST conversion set, and an LFNST index pointing to the LFNST kernel, the first bin of the syntax element bin string for the LFNST index is derived based on context information which differs depending on the tree type of the current block, and the second bin of the syntax element bin string is derived based on previously set context information.
[0009] The process further includes the steps of: deriving context information for a syntax element for the LFNST index; decoding the bin of the syntax element bin string for the LFNST index based on the context information; and deriving the value of the syntax element for the LFNST index.
[0010] If the tree type of the current block is a single tree, the first bin is derived as first context information; if the tree type of the current block is not a single tree, the first bin is derived as second context information.
[0011] The context information for the second bin is derived as a third context information that is different from the first and second context information.
[0012] The LFNST conversion set includes two LFNST kernels, and the value for the syntax element includes one of the following: 0, indicating that the LFNST is not applied to the current block; 1, indicating the first LFNST kernel; or 2, indicating the second LFNST kernel.
[0013] The values for the syntax elements are binary-encoded using truncated unalicode, where a value of 0 for the syntax element becomes "0", a value of 1 for the syntax element becomes "10", and a value of 2 for the syntax element becomes "11".
[0014] According to one embodiment of this document, a video encoding method is provided that is performed by an encoding device. The method includes the steps of: applying at least one of LFNST or MTS to a residual sample to derive a conversion coefficient for the current block; encoding an LFNST index pointing to an LFNST kernel and quantized residual information, wherein the LFNST is performed based on an LFNST conversion set and the LFNST kernel contained in the LFNST conversion set, the first bin of the syntax element bin string for the LFNST index is derived based on context information which differs depending on the tree type of the current block, and the second bin of the syntax element bin string is derived based on previously set context information.
[0015] According to another embodiment of this document, a digital storage medium is provided which stores video data containing encoded video information and a bitstream generated by a video encoding method performed by an encoding device.
[0016] According to another embodiment of this document, there is provided a digital storage medium storing video data including encoded video information and a bitstream for causing a decoding device to execute the video decoding method.
Advantages of the Invention
[0017] According to this document, the overall video / video compression efficiency can be increased.
[0018] According to this document, the efficiency of LFNST index coding can be increased.
[0019] According to this document, the efficiency of the secondary conversion can be increased through the coding of the LFNST index.
[0020] <00 [Figure 3] This diagram schematically illustrates the configuration of a video / image decoding device to which this document can be applied.
[0024] [Figure 4] An illustrative diagram of a content streaming system structure to which this document applies is shown below.
[0025] [Figure 5] A schematic representation of a multiplexing technique according to one embodiment of this document is provided below.
[0026] [Figure 6] Sixty-five intra-directional modes for prediction directions are shown as examples.
[0027] [Figure 7] This is a diagram illustrating the RST according to one embodiment of this document.
[0028] [Figure 8] This diagram shows the order in which the output data of a forward linear transformation is arranged into a one-dimensional vector, as an example.
[0029] [Figure 9] This diagram shows the order in which the output data of a forward quadratic transformation is arranged in a two-dimensional block, as an example.
[0030] [Figure 10] This figure shows a wide-angle intra-prediction mode according to one embodiment of this document.
[0031] [Figure 11] This diagram shows the block pattern to which LFNST is applied.
[0032] [Figure 12] This figure shows the array of output data for a forward LFNST as an example.
[0033] [Figure 13]This diagram shows a zero-out in a block to which a 4x4 LFNST is applied, as an example.
[0034] [Figure 14] This diagram shows a zero-out in a block to which an 8x8 LFNST is applied, as an example.
[0035] [Figure 15] This figure shows a block diagram of a CABAC encoding system according to one embodiment.
[0036] [Figure 16] This is a flowchart illustrating the operation of a video decoding device according to one embodiment of this document.
[0037] [Figure 17] This is a flowchart illustrating the operation of a video encoding device according to one embodiment described herein. [Modes for carrying out the invention]
[0038] This document is subject to various modifications and can have many different embodiments, and specific embodiments are illustrated in the drawings for detailed explanation. However, this document is not intended to limit itself to any particular embodiment. The terms used herein are used solely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this document, terms such as "includes" or "has" indicate the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should not be understood as excluding the existence or possibility of adding one or more other features, figures, steps, actions, components, parts, or combinations thereof.
[0039] On the other hand, each configuration shown in the diagrams described in this document is illustrated independently for the convenience of explaining its distinct characteristic functions, and does not mean that each configuration is embodied in separate hardware or separate software. For example, two or more of the configurations may be combined to form a single configuration, and a single configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included within the scope of the rights of this document, as long as they do not deviate from the essence of this document.
[0040] The preferred embodiments of this document will be described in more detail below with reference to the attached drawings. The same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components will be omitted.
[0041] This document relates to video / image coding. For example, the methods / examples disclosed in this document are related to the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), next-generation video / image coding standards after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), EVC (essential video coding) standard, AVS2 standard, etc.).
[0042] This document presents various implementations of video / image coding, and unless otherwise noted, these implementations can be combined and implemented in conjunction with each other.
[0043] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a single image representing a specific time period, while "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more CTUs (coding tree units). A single picture can consist of one or more slices or tiles. A single picture can consist of one or more tile groups. A tile group can contain one or more tiles.
[0044] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" can be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, or it can represent only the luma component pixel / pixel value, or only the chroma component pixel / pixel value. Alternatively, a sample can refer to a pixel value in the spatial domain, and if such a pixel value is converted to the frequency domain, it can also refer to the conversion coefficient in the frequency domain.
[0045] A unit can represent a basic unit of image processing. A unit can contain at least one of the following: a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block can contain a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0046] In this document, the terms " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Similarly, "A, B, C" also means "at least one of A, B, and / or C."
[0047] In addition, in this document, "or" is interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) both "A and B." Alternatively, "or" in this document can mean "additionally or alternatively."
[0048] In this specification, "at least one of A and B" can mean "just A," "just B," or "both A and B." Furthermore, in this specification, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B."
[0049] Furthermore, in this specification, "at least one of A, B and C" may mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."
[0050] Furthermore, parentheses used in this specification can mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra prediction," but rather "intra prediction" is proposed as an example of "prediction." Similarly, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" is proposed as an example of "prediction."
[0051] In this specification, technical features described individually within a single drawing may be represented individually or simultaneously.
[0052] Figure 1 schematically shows an example of a video / image coding system to which this document can be applied.
[0053] Referring to Figure 1, a video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.
[0054] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may consist of a separate device or external component.
[0055] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or video / image archives containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.
[0056] An encoding device can encode input video / image data. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.
[0057] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0058] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.
[0059] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0060] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the term "video encoding device" may include image encoding devices.
[0061] Referring to Figure 2, the encoding device 200 may be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a reconstructor or a reconstructed block generator. The aforementioned video splitting unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0062] The video splitting unit 210 can split the input video (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure and / or the ternary structure. Alternatively, the binary tree structure may be applied first. The coding procedure according to this document can be executed based on the final coding unit that cannot be further split. In this case, based on coding efficiency due to video characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit is a unit of sample prediction, and the conversion unit is a unit that derives a conversion coefficient and / or a unit that derives a residual signal from the conversion coefficient.
[0063] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.
[0064] The subtraction unit 231 can generate a residual signal (residual block, residual sample, or residual sample array) by subtracting the predicted signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input video signal (original block, original sample, or original sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform a prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes a predicted sample for the current block. The prediction unit 220 can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. As will be described later in the explanation of each prediction mode, the prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240. The information related to prediction can be encoded by the entropy encoding unit 240 and output in bitstream form.
[0065] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. Depending on the prediction mode, the referenced sample may be located adjacent to or far from the current block. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of prediction direction. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 may also use the prediction modes applied to adjacent blocks to determine the prediction mode to be applied to the current block.
[0066] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks that exist in the current picture and temporally adjacent blocks that exist in the reference picture. The reference picture containing the reference block and the reference picture containing the temporally adjacent block may be the same or different. The temporally adjacent block may be called a collocated reference block or colCU, and the reference picture containing the temporally adjacent block may be called a collocated picture (colPic). For example, the interpretation unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, no residual signal is transmitted.In motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0067] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra-prediction (CIIP). The prediction unit can also perform intra-block copy (IBC) for predictions on blocks. The intra-block copy can be used for content video / moving video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in a manner similar to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document.
[0068] The predicted signals generated via the interpretation unit 221 and / or intrapretation unit 222 can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can apply transformation techniques to the residual signal to generate transformation coefficients. For example, transformation techniques may include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph, assuming that the relationship information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a predicted signal using all previously reconstructed pixels and based on that. The transformation process can also be applied to pixel blocks of the same size that are square, or to non-square blocks of variable size.
[0069] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). The entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. Signaling / transmitted information and / or syntax elements described later in this document can be encoded via the encoding procedure described above and included in the bitstream.The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.
[0070] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the prediction unit 220. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, and can also be used for inter-prediction of the next picture after filtering, as described later.
[0071] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.
[0072] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream format.
[0073] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve encoding efficiency.
[0074] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.
[0075] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which this document can be applied.
[0076] Referring to Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-predictor 332 and an intra-predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The aforementioned entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering device 350 can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware component may further include memory 360 as an internal / external component.
[0077] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Thus, the decoding processing unit is, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit by a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed video signal decoded and output via the decoding device 300 can then be reproduced via a playback device.
[0078] The decoding device 300 can receive the signal output from the encoding device shown in Figure 2 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image restoration and the quantized values of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of adjacent and decoded blocks or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element.In this case, the CABAC entropy decoding method can update the context model after determining the context model by utilizing the information of the decoded symbol / bin for the context model of the next symbol / bin. Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit 330, and information for the residual on which entropy decoding was performed by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the inverse quantization unit 321. In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device described in this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, prediction unit 330, addition unit 340, filtering unit 350, and memory 360.
[0079] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.
[0080] The inverse conversion unit 322 performs an inverse conversion on the conversion coefficients to obtain a residual signal (residual block, residual sample array).
[0081] The prediction unit can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode.
[0082] The prediction unit can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra-prediction (CIIP). The prediction unit can also perform intra-block copying (IBC) for predictions on blocks. This intra-block copying can be used for content video / moving video coding, such as in SCC (screen content coding), for example, in games. IBC basically performs prediction within the current picture, but can be performed in a manner similar to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document.
[0083] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. Depending on the prediction mode, the referenced sample may be located adjacent to the current block or at a distance. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also use the prediction mode applied to the adjacent block to determine the prediction mode to be applied to the current block.
[0084] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.
[0085] The summing unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit 330. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.
[0086] The addition unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.
[0087] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0088] The filtering unit 350 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 350 can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0089] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 332 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.
[0090] In this specification, the embodiments described for the prediction unit 330, inverse quantization unit 321, inverse conversion unit 322, and filtering unit 350 of the decoding device 300 can be applied identically or correspondingly to the prediction unit 220, inverse quantization unit 234, inverse conversion unit 235, and filtering unit 260 of the encoding device 200, respectively.
[0091] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived in both the encoding and decoding devices, and the encoding device can improve video coding efficiency by signaling the decoding device information about the residual between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can combine the residual block and the predicted block to generate a restored block containing restored samples, and can generate a restored picture containing the restored block.
[0092] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, perform a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, perform a quantization procedure on the transformation coefficients to derive quantized transformation coefficients, and signal the associated residual information (via a bitstream) to a decoding device. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can perform an inverse quantization / inverse transformation procedure based on the residual information to derive a residual sample (or residual block). The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.
[0093] Figure 4 illustrates an example of a content streaming system structure to which this document applies.
[0094] Furthermore, the content streaming systems to which this document applies may include a wide range of components, such as encoding servers, streaming servers, web servers, media storage facilities, user equipment, and multimedia input devices.
[0095] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted. The bitstream can be generated by the encoding method or bitstream generation method to which this document applies, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving it.
[0096] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.
[0097] The streaming server can receive content from a media storage and / or encoding server. For example, if it starts receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream over a certain period of time.
[0098] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, and digital signage. Each server within the content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.
[0099] Figure 5 schematically illustrates the multiple transformation technique described in this document.
[0100] Referring to Figure 5, the conversion unit can correspond to the conversion unit in the encoding device shown in Figure 2, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device shown in Figure 2 or the inverse conversion unit in the decoding device shown in Figure 3.
[0101] The transformation unit can perform a linear transformation based on the residual samples (residual sample array) in the residual block to derive (linear) transformation coefficients (S510). Such a linear transformation (primary transform) can be called a core transform. Here, the linear transformation can be based on Multiple Transform Selection (MTS), and when a multiple transform is applied as the linear transformation, it can be called a multiple core transform.
[0102] Multiple core transforms can represent a transformation method that additionally uses a Discrete Cosine Transform (DCT) type 2 and a Discrete Sine Transform (DST) type 7, a Discrete Sine Transform (DST) type 8, and / or a Discrete Sine Transform (DST) type 1. That is, the multiple core transforms can represent a transformation method that transforms a spatial domain residual signal (or residual block) into frequency domain transformation coefficients (or first-order transformation coefficients) based on a plurality of transformation kernels selected from the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the first-order transformation coefficients can be called temporary transformation coefficients from the perspective of the transformer.
[0103] That is, when an existing conversion method is applied, a spatial-domain to frequency-domain conversion of a residual signal (or residual block) can be applied based on DCT type 2 to generate conversion coefficients. In contrast, when the multiple core conversion is applied, a spatial-domain to frequency-domain conversion of a residual signal (or residual block) can be applied based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate conversion coefficients (or first-order conversion coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc., can be called conversion types, conversion kernels, or conversion cores. Such DCT / DST conversion types can be defined based on basis functions.
[0104] When the multi-core transformation is performed, a vertical transformation kernel and a horizontal transformation kernel can be selected from the transformation kernels for the target block, and a vertical transformation can be performed on the target block based on the vertical transformation kernel, and a horizontal transformation can be performed on the target block based on the horizontal transformation kernel. Here, the horizontal transformation may represent a transformation to the horizontal component of the target block, and the vertical transformation may represent a transformation to the vertical component of the target block. The vertical transformation kernel / horizontal transformation kernel can be adaptively determined based on the prediction mode and / or transformation index of the target block (CU or subblock) including the residual block.
[0105] Furthermore, for example, when applying MTS to perform a linear transformation, a specific basis function can be set to a predetermined value, and the mapping relationship to the transformation kernel can be set by combining which basis function is applied when it is a vertical or horizontal transformation. For example, if the horizontal transformation kernel is represented by trTypeHor and the vertical transformation kernel is represented by trTypeVer, then a trTypeHor or trTypeVer value of 0 can be set to DCT2, a trTypeHor or trTypeVer value of 1 can be set to DST7, and a trTypeHor or trTypeVer value of 2 can be set to DCT8.
[0106] In this case, the MTS index information can be encoded and signaled to the decoding device to indicate one of several sets of conversion kernels. For example, if the MTS index is 0, it can indicate that both trTypeHor and trTypeVer values are 0; if the MTS index is 1, it can indicate that both trTypeHor and trTypeVer values are 1; if the MTS index is 2, it can indicate that the trTypeHor value is 2 and the trTypeVer value is 1; if the MTS index is 3, it can indicate that the trTypeHor value is 1 and the trTypeVer value is 2; and if the MTS index is 4, it can indicate that both trTypeHor and trTypeVer values are 2.
[0107] As an example, the conversion kernel set based on MTS index information is shown in the table below.
[0108] [Table 1]
[0109] The transformation unit can perform a quadratic transformation based on the (primary) transformation coefficients to derive modified (secondary) transformation coefficients (S520). The primary transformation is a transformation from the spatial domain to the frequency domain, and the secondary transformation means a transformation in a more compressed representation by utilizing the correlations that exist between the (primary) transformation coefficients. The secondary transformation may include a non-separable transform. In this case, the secondary transformation may be called a non-separable secondary transform (NSST) or MDNSST (mode-dependent non-separable secondary transform). The non-separable secondary transform may represent a transformation that generates modified transformation coefficients (or quadratic transformation coefficients) for a residual signal by quadratic transforming the (primary) transformation coefficients derived via the primary transformation based on a non-separable transform matrix. Here, based on the non-separable transformation matrix, the transformation can be applied at once to the (primary) transformation coefficients without separately applying vertical and horizontal transformations (or horizontal and vertical transformations independently). That is, the non-separable quadratic transformation can be applied to the (primary) transformation coefficients not separately in the vertical and horizontal directions, but rather, for example, a transformation method can be shown in which a two-dimensional signal (transformation coefficients) is realigned to a one-dimensional signal via a specific direction (e.g., row-first direction or column-first direction), and then a modified transformation coefficient (or quadratic transformation coefficient) is generated based on the non-separable transformation matrix. For example, row-first ordering means arranging the M×N blocks in a column in the order of the 1st row, 2nd row, ..., Nth row, and column-first ordering means arranging the M×N blocks in a column in the order of the 1st column, 2nd column, ..., Mth column. The aforementioned non-separable quadratic transformation can be applied to the top-left region of a block composed of (linear) transformation coefficients (hereinafter referred to as a transformation coefficient block).For example, if both the width (W) and height (H) of the conversion coefficient block are 8 or greater, an 8x8 inseparable quadratic transformation can be applied to the 8x8 region at the upper left end of the conversion coefficient block. Also, if both the width (W) and height (H) of the conversion coefficient block are 4 or greater, and either the width (W) or height (H) of the conversion coefficient block is less than 8, a 4x4 inseparable quadratic transformation can be applied to the min(8,W) x min(8,H) region at the upper left end of the conversion coefficient block. However, the embodiments are not limited to these, and for example, even if only the condition that both the width (W) or height (H) of the conversion coefficient block are 4 or greater is met, a 4x4 inseparable quadratic transformation can also be applied to the min(8,W) x min(8,H) region at the upper left end of the conversion coefficient block.
[0110] Specifically, for example, when a 4x4 input block is used, the unseparated quadratic transform can be performed as follows:
[0111] The aforementioned 4x4 input block X is shown as follows:
[0112]
number
[0113] When X is expressed in vector form, JPEG0007897403000003.jpg13104 is shown as follows:
[0114]
number
[0115] As shown in equation 2, vector JPEG0007897403000005.jpg12101 rearranges the 2D block of X in equation 1 into a 1D vector using row-first order.
[0116] In this case, the quadratic inseparable transform can be calculated as follows.
[0117]
number
[0118] Here, JPEG0007897403000007.jpg1197 shows the transformation coefficient vector, and T shows the 16×16 (non-separable) transformation matrix.
[0119] 16 × 1 transformation coefficient vector via the above formula 3 JPEG0007897403000008.jpg1093 can be derived, and the above JPEG0007897403000009.jpg12107 can be reorganized into 4x4 blocks via a scan order (horizontal, vertical, diagonal, etc.). However, the calculation described above is merely illustrative, and to reduce the computational complexity of the non-separable quadratic transform, methods such as HyGT (Hypercube-Givens Transform) can also be used for calculating the non-separable quadratic transform.
[0120] On the other hand, the non-separated quadratic transform can be a mode-dependent transform kernel (or transform core, transform type), where the mode can include intra-predictive mode and / or inter-predictive mode.
[0121] As described above, the non-separable quadratic transformation can be performed based on an 8x8 transformation or a 4x4 transformation determined based on the width (W) and height (H) of the transformation coefficient block. An 8x8 transformation refers to a transformation that can be applied to an 8x8 region contained within the transformation coefficient block when both W and H are greater than or equal to 8, and the 8x8 region is the upper left corner 8x8 region within the transformation coefficient block. Similarly, a 4x4 transformation refers to a transformation that can be applied to a 4x4 region contained within the transformation coefficient block when both W and H are greater than or equal to 4, and the 4x4 region is the upper left corner 4x4 region within the transformation coefficient block. For example, an 8x8 transformation kernel matrix can be a 64x64 / 16x64 matrix, and a 4x4 transformation kernel matrix can be a 16x16 / 8x16 matrix.
[0122] In this case, for mode-based conversion kernel selection, two non-separable quadratic conversion kernels can be configured for each conversion set for non-separable quadratic conversions for both 8x8 and 4x4 conversions, resulting in four conversion sets. That is, four conversion sets can be configured for 8x8 conversions, and four conversion sets can be configured for 4x4 conversions. In this case, each of the four conversion sets for 8x8 conversions can contain two 8x8 conversion kernels, and each of the four conversion sets for 4x4 conversions can contain two 4x4 conversion kernels.
[0123] However, the size of the transformation, i.e., the size of the region to which the transformation is applied, is merely illustrative, and sizes other than 8x8 or 4x4 can be used, and the number of sets is n, with k transformation kernels in each set.
[0124] The aforementioned set of transformations may be called an NSST set or an LFNST set. The selection of a particular set from among the aforementioned set of transformations can be performed, for example, based on the intra-prediction mode of the current block (CU or sub-block). LFNST (Low-Frequency Non-Separable Transform) is an example of a reduced non-separable transform described later, and represents a non-separable transform for low-frequency components.
[0125] For reference, for example, an intraprediction mode may include two non-directinoal (or non-angular) intraprediction modes and 65 directional (or angular) intraprediction modes. The non-directinoal intraprediction mode may include a planar intraprediction mode (number 0) and a DC intraprediction mode (number 1), and the directional intraprediction mode may include 65 intraprediction modes (numbers 2 through 66). However, this is merely an example, and this document may also apply to cases with a different number of intraprediction modes. On the other hand, a 67th intraprediction mode may be used in some cases, and this 67th intraprediction mode may represent a linear model (LM) mode.
[0126] Figure 6 illustrates 65 intradirectional modes for predicting directions.
[0127] Referring to Figure 6, we can distinguish between intra-prediction modes with horizontal directionality and intra-prediction modes with vertical directionality, centered around intra-prediction mode 34, which has a right-downward diagonal prediction direction. In Figure 6, H and V represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate a displacement of 1 / 32 units on the sample grid position. This can be used to indicate an offset relative to the mode index value. Intra-prediction modes 2 through 33 have horizontal directionality, while intra-prediction modes 34 through 66 have vertical directionality. On the other hand, intra-prediction mode 34 can be considered neither strictly horizontal nor vertical, but it can be classified as belonging to the horizontal directionality in terms of determining the transformation set of the quadratic transformation. This is because the input data is transposed for the vertical directionality modes which are symmetrical with respect to intra-prediction mode 34, while the input data alignment method for the horizontal directionality modes is used for intra-prediction mode 34. Transposing input data means that for 2D block data M×N, rows become columns and columns become rows to form N×M data. Intra-prediction modes 18 and 50 represent the horizontal intra-prediction mode and the vertical intra-prediction mode, respectively. Intra-prediction mode 2 can be called the diagonal intra-prediction mode because it predicts in the diagonal direction to the right with a left reference pixel. In the same context, intra-prediction mode 34 can be called the diagonal intra-prediction mode to the right and down, and intra-prediction mode 66 can be called the diagonal intra-prediction mode to the left and down.
[0128] For example, the mapping of four transformation sets by the intra-predictive mode is shown in the table below.
[0129] [Table 2]
[0130] As shown in Table 2, the intra prediction mode allows any one of the four transformation sets, i.e., lfnstTrSetIdx to be mapped to any one of the four values from 0 to 3.
[0131] On the other hand, if it is determined that a specific set is to be used for an inseparable transformation, one of the k transformation kernels in the specific set can be selected via an inseparable quadratic transformation index. The encoding device can derive an inseparable quadratic transformation index that points to a specific transformation kernel based on an RD (rate-distortion) check and signal the decoding device to the inseparable quadratic transformation index. The decoding device can select one of the k transformation kernels in the specific set based on the inseparable quadratic transformation index. For example, lfnst index value 0 can point to the first inseparable quadratic transformation kernel, lfnst index value 1 can point to the second inseparable quadratic transformation kernel, and lfnst index value 2 can point to the third inseparable quadratic transformation kernel. Alternatively, lfnst index value 0 can indicate that the first inseparable quadratic transformation is not applied to the target block, and lfnst index values 1 to 3 can point to the three transformation kernels.
[0132] The transformation unit can perform the non-separable quadratic transformation based on the selected transformation kernel to obtain modified (quadratic) transformation coefficients. The modified transformation coefficients can be derived as quantized transformation coefficients via the quantization unit, as described above, and can be encoded and transmitted to the decoding unit for signaling and the inverse quantization / inverse transformation unit within the encoding unit.
[0133] On the other hand, as mentioned above, if the quadratic transformation is omitted, the (primary) transformation coefficients, which are the output of the primary (separated) transformation, can be derived as quantized transformation coefficients via the quantization unit, as described above, and can be encoded and transmitted to the decoding device for signaling and to the inverse quantization / inverse transformation unit within the encoding device.
[0134] The inverse transform unit can perform a series of steps in the reverse order of the steps performed by the transform unit described above. The inverse transform unit can receive the (inversely quantized) transform coefficients, perform a quadratic (inverse) transform to derive (primary) transform coefficients (S550), and perform a primary (inverse) transform on the (primary) transform coefficients to obtain a residual block (residual samples) (S560). Here, the primary transform coefficients can be called modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding and decoding devices can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.
[0135] On the other hand, the decoding device may further include a quadratic inverse transform applicability determination unit (or an element that determines whether a quadratic inverse transform is applicable) and a quadratic inverse transform determination unit (or an element that determines the quadratic inverse transform). The quadratic inverse transform applicability determination unit can determine whether a quadratic inverse transform is applicable. For example, the quadratic inverse transform may be NSST, RST, or LFNST, and the quadratic inverse transform applicability determination unit can determine whether a quadratic inverse transform is applicable based on a quadratic transform flag parsed from the bitstream. As another example, the quadratic inverse transform applicability determination unit may also determine whether a quadratic inverse transform is applicable based on the transformation coefficients of the residual block.
[0136] The quadratic inverse transform determination unit can determine the quadratic inverse transform. In this case, the quadratic inverse transform determination unit can determine the quadratic inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified by the intra-prediction mode. In one embodiment, the quadratic transform determination method can be determined depending on the linear transform determination method. Various combinations of linear and quadratic transforms can be determined by the intra-prediction mode. Also, as an example, the quadratic inverse transform determination unit can determine the region to which the quadratic inverse transform is applied based on the size of the current block.
[0137] On the other hand, as mentioned above, if the quadratic (inverse) transformation is omitted, a residual block (residual sample) can be obtained by receiving the (inversely quantized) transformation coefficients and performing the primary (separated) inverse transformation. As mentioned above, the encoding and decoding devices can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.
[0138] On the other hand, in this document, in order to reduce the computational complexity and memory requirements of unseparable quadratic transforms, we can apply RST (reduced secondary transform), which is a reduced version of the transformation matrix (kernel) based on the NSST concept.
[0139] On the other hand, the conversion kernel, conversion matrix, and coefficients constituting the conversion kernel matrix described in this document—that is, kernel coefficients or matrix coefficients—can be represented using 8 bits. This is one condition for implementation in decoding and encoding devices, and it reduces the memory requirements for storing the conversion kernel, along with a reasonably acceptable performance degradation compared to existing 9-bit or 10-bit representations. Furthermore, representing the kernel matrix with 8 bits allows for the use of smaller multipliers, making it more compatible with SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.
[0140] In this specification, RST can mean a transformation performed on a residual sample for a target block based on a transform matrix whose size has been reduced by a simplification factor. When a simplification transformation is performed, the amount of computation required during the transformation can be reduced by reducing the size of the transform matrix. That is, RST can be used to resolve the complexity problem that arises when transforming large blocks or during non-separable transformations.
[0141] RST can be referred to by a variety of terms, including reduced transform, reduced secondary transform, reduction transform, simplified transform, and simple transform, and the name RST is not limited to the examples listed. Alternatively, since RST primarily occurs in the low-frequency region with non-zero coefficients in the transform block, it can also be called LFNST (Low-Frequency Non-Separable Transform). The aforementioned transform index can be named the LFNST index.
[0142] On the other hand, when the quadratic inverse transformation is performed based on RST, the inverse transformation unit 235 of the encoding device 200 and the inverse transformation unit 322 of the decoding device 300 may include an inverse RST unit that derives corrected transformation coefficients based on the inverse RST of the transformation coefficients, and an inverse linear transformation unit that derives a residual sample for the target block based on the inverse linear transformation of the corrected transformation coefficients. The inverse linear transformation means the inverse transformation of the linear transformation applied to the residual. In this document, deriving transformation coefficients based on a transformation may mean deriving transformation coefficients by applying the transformation in question.
[0143] Figure 7 is a diagram illustrating an RST according to one embodiment described in this document.
[0144] In this specification, “target block” may mean the current block, residual block, or transformed block on which coding is performed.
[0145] In one embodiment of the RST, an N-dimensional vector can be mapped to an R-dimensional vector located in a different space, and a reduced transformation matrix can be determined, where R is less than N. N can represent the square of the side length of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplified factor can represent the R / N value. The simplified factor can be called by various terms such as reduced factor, reduction factor, simplified factor, or simple factor. On the other hand, R can be called the reduced coefficient, but in some cases the simplified factor can also represent R. Also, in some cases the simplified factor can represent the N / R value.
[0146] In one embodiment, the simplification factor or simplification coefficient may be signaled via a bitstream, but the embodiment is not limited thereto. For example, predefined values for the simplification factor or simplification coefficient may be stored in each encoding device 200 and decoding device 300, in which case the simplification factor or simplification coefficient is not signaled separately.
[0147] The size of the simplified transformation matrix according to one embodiment is R×N, which is smaller than the size N×N of a normal transformation matrix, and can be defined as shown in the following equation 4.
[0148]
number
[0149] The matrix T in the Reduced Transform block shown in Figure 7(a) can be interpreted as the matrix TR × N in Equation 4. As shown in Figure 7(a), when the simplified transformation matrix TR × N is multiplied by the residual sample for the target block, the transformation coefficients for the target block can be derived.
[0150] In one embodiment, if the size of the block to which the transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), then the RST shown in Figure 7(a) can be expressed by a matrix operation as shown in Equation 5 below. In this case, the memory and multiplication operations can be reduced to approximately 1 / 4 by the simplification factor.
[0151] In this document, matrix operations can be understood as operations in which a matrix is placed to the left of a column vector and multiplied by the column vector to obtain a new column vector.
[0152]
number
[0153] In equation 5, r1 to r 64 This can represent the residual sample for the target block, and more specifically, it is the transformation coefficient generated by applying a linear transformation. The result of the calculation in Equation 5 is the transformation coefficient c for the target block. i This can be derived, and c i The derivation process is as shown in equation 6.
[0154]
number
[0155] The result of the calculation in formula 6, and the conversion coefficients c1 to c for the target block. R The following can be derived: That is, when R=16, the conversion coefficients c1 to c for the target block are 16 The following can be derived. If a transformation matrix of size 64×64 (N×N) is applied instead of RST and multiplied by a residual sample of size 64×1 (N×1), 64 (N) transformation coefficients will be derived for the target block. However, because RST was applied, only 16 (R) transformation coefficients will be derived for the target block. The total number of transformation coefficients for the target block decreases from N to R, and the amount of data that the encoding device 200 sends to the decoding device 300 decreases, thus increasing the transmission efficiency between the encoding device 200 and the decoding device 300.
[0156] From the perspective of the size of the transformation matrix, the size of a normal transformation matrix is 64 × 64 (N × N), while the size of the simplified transformation matrix is reduced to 16 × 64 (R × N). Therefore, compared to performing a normal transformation, the memory usage when executing RST can be reduced to an R / N ratio. Also, compared to the number of multiplication operations N × N when using a normal transformation matrix, the number of multiplication operations when using a simplified transformation matrix can be reduced to an R / N ratio (R × N).
[0157] In one embodiment, the conversion unit 232 of the encoding device 200 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the residual samples for the target block. Such conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 300. The inverse conversion unit 322 of the decoding device 300 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) for the conversion coefficients, and can derive residual samples for the target block based on an inverse primary conversion for the modified conversion coefficients.
[0158] Inverse RST matrix T according to one embodiment N×R has a size of N×R, which is smaller than the size N×N of the normal inverse conversion matrix, and is in a transpose relationship with the simplified conversion matrix T shown in Equation 4 R×N .
[0159] The matrix T in the Reduced Inv.Transform block shown in (b) of FIG. 7 t can mean the inverse RST matrix T R×N T (the superscript T means transpose). As shown in (b) of FIG. 7, when the inverse RST matrix T R×N T is multiplied by the conversion coefficients for the target block, modified conversion coefficients for the target block or residual samples for the target block can be derived. The inverse RST matrix T R×N T can also be expressed as (T R×N ) T N×R .
[0160] More specifically, when inverse RST is applied as the secondary inverse conversion, the inverse RST matrix T R×N TWhen multiplied by this, the modified transformation coefficients for the target block can be derived. On the other hand, the inverse RST can be applied as an inverse linear transformation, in which case the inverse RST matrix T can be applied to the transformation coefficients for the target block. R×N T When multiplied, the residual sample for the target block can be derived.
[0161] In one embodiment, when the size of the block to which the inverse transform is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the RST shown in Figure 7(b) can be expressed by a matrix operation as shown in the following equation 7.
[0162]
number
[0163] In formula 7, c1 to c 16 This can show the conversion coefficient for the target block. The result of the calculation in Equation 7, the modified conversion coefficient for the target block, or r showing the residual sample for the target block. i This can be derived, and r i The derivation process is as shown in equation 8.
[0164]
number
[0165] The result of the calculation in Equation 8, the modified conversion coefficient for the target block, or r1 to r representing the residual sample for the target block. NThis can be derived. Considering the size of the inverse transform matrix, the size of a normal inverse transform matrix is 64 × 64 (N × N), while the size of the simplified inverse transform matrix is reduced to 64 × 16 (N × R). Therefore, compared to performing a normal inverse transform, memory usage when performing the inverse RST can be reduced to an R / N ratio. Also, compared to the number of multiplication operations N × N when using a normal inverse transform matrix, the number of multiplication operations can be reduced to an R / N ratio (N × R) when using a simplified inverse transform matrix.
[0166] On the other hand, the transformation set configuration shown in Table 2 can also be applied to 8x8 RST. That is, the corresponding 8x8 RST can be applied using the transformation set in Table 2. Since one transformation set consists of two or three transformations (kernels) depending on the prediction mode on the screen, it can be configured to select one from a maximum of four transformations, including not applying a quadratic transformation. When a quadratic transformation is not applied, it can be considered as if the identity matrix is applied. If we assign indices 0, 1, 2, and 3 to the four transformations respectively (for example, index 0 can be assigned to the identity matrix, i.e., when a quadratic transformation is not applied), the transformation to be applied can be specified by signaling a transformation index or lfnst index, which is a syntax element, for each transformation coefficient block. That is, via the transformation index, 8x8 RST can be specified for the 8x8 top-left block in the RST configuration, or 8x8 lfnst can be specified if LFNST is applied. 8×8 lfnst and 8×8 RST refer to transformations that can be applied to an 8×8 region contained within a transformation coefficient block when both W and H of the target block to be transformed are greater than or equal to 8, and the relevant 8×8 region is the upper left corner 8×8 region within the transformation coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to transformations that can be applied to a 4×4 region contained within a transformation coefficient block when both W and H of the target block are greater than or equal to 4, and the relevant 4×4 region is the upper left corner 4×4 region within the transformation coefficient block.
[0167] On the other hand, according to one embodiment of this document, in the encoding process, it is possible to select only 48 data points from 64 data points constituting an 8x8 region, rather than using a 16x64 transformation kernel matrix, and apply a maximum of a 16x48 transformation kernel matrix to them. Here, "maximum" means that the maximum value of m is 16 for an mx48 transformation kernel matrix that can generate m coefficients. That is, when RST is executed by applying an mx48 transformation kernel matrix (m ≤ 16) to an 8x8 region, it is possible to generate m coefficients from 48 data inputs. If m is 16, it generates 16 coefficients from 48 data inputs. That is, if the 48 data points form a 48x1 vector, a 16x1 vector can be generated by sequentially multiplying the 16x48 matrix and the 48x1 vector. At this time, the 48 data points constituting an 8x8 region can be appropriately arranged to construct a 48x1 vector. For example, a 48x1 vector can be constructed based on 48 data points that constitute the region excluding the bottom right 4x4 region of the 8x8 region. At this time, when a matrix operation is performed by applying a transformation kernel matrix of up to 16x48, 16 modified transformation coefficients are generated, and these 16 modified transformation coefficients can be placed in the upper left 4x4 region according to the scanning order, while the upper right 4x4 region and the lower left 4x4 region can be filled with 0.
[0168] For the inverse transformation of the decoding process, the transposed matrix of the transformation kernel matrix described above can be used. That is, when inverse RST or LFNST is performed as the inverse transformation process executed by the decoding device, the input coefficient data to which inverse RST is applied is composed of one-dimensional vectors in a predetermined order, and the modified coefficient vectors obtained by multiplying the one-dimensional vectors by the corresponding inverse RST matrix on the left side can be arranged in two-dimensional blocks in a predetermined order.
[0169] To summarize, when RST or LFNST is applied to an 8x8 region during the transformation process, a matrix operation is performed between the 48 transformation coefficients from the top-left, top-right, and bottom-left regions of the 8x8 region (excluding the bottom-right region) and a 16x48 transformation kernel matrix. For this matrix operation, the 48 transformation coefficients are input as a one-dimensional array. After this matrix operation is performed, 16 modified transformation coefficients are derived, which can then be arranged in the top-left region of the 8x8 region.
[0170] Conversely, when inverse RST or LFNST is applied to an 8x8 region during the inverse transformation process, the 16 transformation coefficients corresponding to the upper left corner of the 8x8 region can be input in a one-dimensional array form according to the scanning order and performed matrix operations with a 48x16 transformation kernel matrix. That is, the matrix operation in such a case can be expressed as (48x16 matrix) * (16x1 transformation coefficient vector) = (48x1 modified transformation coefficient vector). Here, the nx1 vector can be interpreted as an nx1 matrix, and can therefore also be expressed as an nx1 column vector. Also, * signifies matrix multiplication. When such a matrix operation is performed, 48 modified transformation coefficients can be derived, and these 48 modified transformation coefficients can be arranged in the upper left, upper right, and lower left regions of the 8x8 region, excluding the lower right region.
[0171] On the other hand, when the quadratic inverse transformation is performed based on RST, the inverse transformation unit 235 of the encoding device 200 and the inverse transformation unit 322 of the decoding device 300 may include an inverse RST unit that derives corrected transformation coefficients based on the inverse RST of the transformation coefficients, and an inverse linear transformation unit that derives a residual sample for the target block based on the inverse linear transformation of the corrected transformation coefficients. The inverse linear transformation means the inverse transformation of the linear transformation applied to the residual. In this document, deriving transformation coefficients based on a transformation may mean deriving transformation coefficients by applying the transformation in question.
[0172] Specifically, regarding the detailed non-separated transform, LFNST, the following applies: LFNST can include a forward transform by the encoding device and an inverse transform by the decoding device.
[0173] The encoding device takes the result (or part of the result) derived after applying a primary (core) transform as input and applies a secondary transform.
[0174]
number
[0175] In the above equation 9, x and y are the input and output of the quadratic transformation, respectively, G is the matrix representing the quadratic transformation, and the transform basis vector consists of column vectors. In the case of inverse LFNST, the dimension of the transformation matrix G is expressed as [number of rows × number of columns], and in the case of forward LFNST, the transposed matrix G is G T It becomes that dimension.
[0176] In the case of the reverse LFNST, the dimensions of matrix G are [48×16], [48×8], [16×16], and [16×8], and the [48×8] matrix and the [16×8] matrix are submatrices obtained by sampling eight transformation basis vectors from the left side of the [48×16] matrix and the [16×16] matrix, respectively.
[0177] In contrast, in the case of forward LFNST, matrix G T The dimensions are [16×48], [8×48], [16×16], and [8×16], and the [8×48] matrix and the [8×16] matrix are submatrices obtained by sampling 8 transformation basis vectors from the upper side of the [16×48] matrix and the [16×16] matrix, respectively.
[0178] Therefore, in the case of forward LFNST, the input x can be a [48×1] vector or a [16×1] vector, and the output y can be a [16×1] vector or an [8×1] vector. Since the output of the forward linear transformation in video coding and decoding is two-dimensional (2D) data, in order to construct a [48×1] vector or a [16×1] vector as input x, the 2D data that is the output of the forward transformation must be appropriately arranged to construct a one-dimensional vector.
[0179] Figure 8 shows the sequence for arranging the output data of a forward linear transformation into a one-dimensional vector, as an example. The left-hand diagrams of Figure 8(a) and (b) show the sequence for creating a [48×1] vector, and the right-hand diagrams of Figure 8(a) and (b) show the sequence for creating a [16×1] vector. In the case of LFNST, a one-dimensional vector x can be obtained by sequentially arranging the 2D data in the order shown in Figure 8(a) and (b).
[0180] The orientation of the output data of such a forward linear transformation can be determined by the intra-prediction mode of the current block. For example, if the intra-prediction mode of the current block is horizontal with respect to the diagonal direction, the output data of the forward linear transformation can be arranged in the order shown in Figure 8(a), and if the intra-prediction mode of the current block is vertical with respect to the diagonal direction, the output data of the forward linear transformation can be arranged in the order shown in Figure 8(b).
[0181] For example, a different ordering can be applied to the arrangement in Figures 8(a) and (b). To derive the same result (y vector) as when applying the ordering in Figures 8(a) and (b), one simply rearranges the column vectors of matrix G to match the desired ordering. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0182] Since the output y derived via equation 9 is a one-dimensional vector, if a configuration that processes the result of a forward quadratic transformation as input, such as a configuration that performs quantization or residual coding, requires two-dimensional data as input data, then the output y vector from equation 9 must again be appropriately arranged in 2D data.
[0183] Figure 9 shows the order in which the output data of a forward quadratic transformation is arranged in a two-dimensional block, as an example.
[0184] In the case of LFNST, the output can be placed in a 2D block according to a predetermined scan order. Figure 9(a) shows that when the output y is a [16×1] vector, the output values are placed in 16 positions in the 2D block according to the diagonal scan order. Figure 9(b) shows that when the output y is an [8×1] vector, the output values are placed in 8 positions in the 2D block according to the diagonal scan order, and the remaining 8 positions are filled with 0. In Figure 9(b), X is shown to be filled with 0.
[0185] As in other examples, the order in which the output vector y is processed by a configuration that performs quantization or residual coding can be a pre-defined order, so the output vector y may not be placed in a 2D block as shown in Figure 9. However, in the case of residual coding, data coding can be performed in units of 2D blocks (e.g., 4x4) such as a Coefficient Group (CG), in which case the data can be arranged in a specific order, as in the diagonal scan order of Figure 9.
[0186] On the other hand, the decoding device can construct a one-dimensional input vector y by arranging the two-dimensional data output via an inverse quantization process or the like in a pre-set scan order for the reverse transformation. The input vector y can be output as the input vector x by the following formula.
[0187]
number
[0188] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. In the case of inverse LFNST, the output vector x is a [48×1] vector or a [16×1] vector.
[0189] The output vector x is arranged in a 2D block in the order shown in Figure 8, and this 2D data becomes the input data (or part of the input data) for the inverse linear transformation.
[0190] Therefore, the inverse quadratic transformation is generally the opposite of the forward quadratic transformation process. In the case of the inverse transformation, unlike the forward transformation, the inverse quadratic transformation is applied first, followed by the inverse linear transformation.
[0191] In reverse LFNST, one transformation matrix G can be selected from eight [48×16] matrices and eight [16×16] matrices. Which of the [48×16] and [16×16] matrices to apply is determined by the size and pattern of the block.
[0192] Furthermore, the eight matrices can be derived from four transformation sets, as shown in Table 2 above, and each transformation set can consist of two matrices. Which of the four transformation sets to use is determined by the intra-prediction mode, and more specifically, the transformation set is determined based on the extended intra-prediction mode value, which also takes into account Wide Angle Intra Prediction (WAIP). Which of the two matrices that make up the selected transformation set is selected is derived via index signaling. More specifically, the transmitted index value can be 0, 1, or 2, where 0 indicates that LFNST should not be applied, and 1 and 2 can indicate either of the two transformation matrices that make up the transformation set selected based on the intra-prediction mode value.
[0193] Figure 10 shows a wide-angle intra-prediction mode according to one embodiment of this document.
[0194] Typical intra-predictive mode values can range from 0 to 66 and from 81 to 83, while, as illustrated, intra-predictive mode values extended by WAIP can range from -14 to 83. Values from 81 to 83 refer to the CCLM (Cross Component Linear Model) mode, while values from -14 to -1 and from 67 to 80 refer to intra-predictive mode values extended by WAIP application.
[0195] When the width of a predicted block is greater than its height, the upper reference pixel is generally closer to the predicted position within the block. Therefore, predicting in the bottom-left direction is more accurate than predicting in the top-right direction. Conversely, when the height of a block is greater than its width, the left reference pixel is generally closer to the predicted position within the block. Therefore, predicting in the top-right direction is more accurate than predicting in the bottom-left direction. Thus, it is advantageous to remap the index in the wide-angle intra-prediction mode, i.e., apply a mode index transformation.
[0196] When wide-angle intra-prediction is applied, information can be signaled to existing intra-predictions, and after the information has been parsed, it can be remapped with the index of the wide-angle intra-prediction mode. Therefore, the total number of intra-prediction modes for a particular block (e.g., a non-square block of a particular size) does not change; that is, the total number of intra-prediction modes is 67, and the intra-prediction mode coding for the particular block does not change.
[0197] Table 3 below shows the process of remapping the intra prediction mode to the wide-angle intra prediction mode to derive the modified intra mode.
[0198] [Table 3]
[0199] In Table 3, the extended intra-prediction mode value is ultimately stored in the predModeIntra variable, ISP_NO_SPLIT indicates that the CU block is not split into subpartitions by the Intra Sub Partitions (ISP) technology now adopted in the VVC standard, and the cIdx variable values of 0, 1, and 2 indicate that the component is Luma, Cb, or Cr, respectively. The Log2 function shown in Table 3 returns a log value with a base of 2, and the Abs function returns the absolute value.
[0200] The input values for the wide-angle intra prediction mode mapping process include the variable `predModeIntra`, which indicates the intra prediction mode, and the height and width of the transformation block. The output value is the modified intra prediction mode (`predModeIntra`). The height and width of the transformation block or coding block can become the height and width of the current block for the intra prediction mode remapping. In this case, the variable `whRatio`, which reflects the ratio of width to height, can be set to `Abs(Log2(nW / nH))`.
[0201] For non-square blocks, the intra-prediction mode can be modified in two separate cases.
[0202] First, if all of the following conditions are met: (1) the current block width is greater than the height, (2) the pre-correction intra prediction mode is greater than or equal to 2, and (3) the intra prediction mode is (8 + 2 * whRatio) if the variable whRatio is greater than 1, and 8 if the variable whRatio is less than or equal to 1, and is less than the derived value [predModeIntra is less than (whRatio>1)?(8 + 2 * whRatio): 8], then the intra prediction mode is set to a value 65 greater than the current intra prediction mode [predModeIntra is set equal to (predModeIntra + 65)].
[0203] If the above conditions are not met, then the intra-prediction mode is set to a value 67 less than the current intra-prediction mode [predModeIntra is set equal to (predModeIntra-67)] if all of the following conditions are met: (1) the current block height is greater than the width, (2) the pre-correction intra-prediction mode is less than or equal to 66, and (3) the intra-prediction mode is (60-2*whRatio) if the variable whRatio is greater than 1, and 60 if the variable whRatio is less than or equal to 1.
[0204] Table 2, mentioned above, shows how the transformation set is selected in LFNST based on the intra-prediction mode values extended by WAIP. As shown in Figure 10, modes 14-33 and modes 35-80 are symmetrical with respect to the prediction direction, with respect to mode 34. For example, modes 14 and 54 are symmetrical with respect to the direction corresponding to mode 34. Therefore, modes located in directions that are symmetrical with respect to each other will be subject to the same transformation set, and this symmetry is reflected in Table 2.
[0205] However, we assume that the forward LFNST input data for mode 54 is symmetrical to the forward LFNST input data for mode 14. For example, for modes 14 and 54, the 2D data is rearranged into 1D data according to the arrangement order shown in Figure 8(a) and Figure 8(b), respectively, and we can see that the arrangement patterns shown in Figure 8(a) and Figure 8(b) are symmetrical with respect to the direction pointed to by mode 34 (diagonal term).
[0206] On the other hand, as mentioned above, the choice of which transformation matrix—the [48×16] matrix or the [16×16] matrix—to apply to LFNST is determined by the size of the block to be transformed.
[0207] Figure 11 shows block patterns to which LFNST is applied. Figure 11(a) shows a 4x4 block, (b) shows 4x8 and 8x4 blocks, (c) shows a 4xN or Nx4 block where N is 16 or greater, (d) shows an 8x8 block, and (e) shows an MxN block where M≧8, N≧8, and N>8 or M>8.
[0208] In Figure 11, the blocks with thick borders indicate the regions to which LFNST is applied. For blocks (a) and (b) in Figure 11, LFNST is applied to the top-left 4x4 region, and for block (c) in Figure 11, LFNST is applied to each of the two consecutively arranged top-left 4x4 regions. In Figures (a), (b), and (c), LFNST is applied in units of 4x4 regions, so such LFNST will be named "4x4LFNST" below, and the corresponding transformation matrix can be a base [16x16] or [16x8] matrix for G in equations 9 and 10.
[0209] More specifically, a [16x8] matrix is applied to the 4x4 block (4x4TU or 4x4CU) in Figure 11(a), and a [16x16] matrix is applied to the blocks in Figure 11(b) and (c). This is to match the computational complexity for the worst case to 8 multiplications per sample.
[0210] For Figures 11(d) and (e), LFNST is applied to the upper left 8x8 region, and such LFNST will be referred to as "8x8LFNST" below. The corresponding transformation matrix can be either a [48x16] or [48x8] matrix. In the case of forward LFNST, a [48x1] vector (the x vector in equation 9) is input as input data, so not all sample values in the upper left 8x8 region are used as input values for forward LFNST. That is, as can be seen in the left-hand order of Figure 8(a) or Figure 8(b), the bottom-right 4x4 block is left as is, and a [48x1] vector can be constructed based on the samples belonging to the remaining three 4x4 blocks.
[0211] A [48x8] matrix can be applied to the 8x8 block (8x8TU or 8x8CU) in Figure 11(d), and a [48x16] matrix can be applied to the 8x8 block in Figure 11(e). This is also to match the computational complexity for the worst case to 8 multiplications per sample.
[0212] Depending on the block pattern, applying the corresponding forward LFNST (4x4 LFNST or 8x8 LFNST) will generate 8 or 16 output data (y vectors in Equation 9, [8x1] or [16x1] vectors), and due to the properties of the matrix GT in forward LFNST, the number of output data will be the same as or less than the number of input data.
[0213] Figure 12 shows an example of the output data array for a forward LFNST, illustrating the blocks in which the output data for the forward LFNST is arranged according to the block pattern.
[0214] In Figure 12, the shaded area at the upper left corner of the block corresponds to the area where the output data of the forward LFNST is located. Positions marked with 0 indicate samples filled with the value 0, and the remaining area indicates the area that is not modified by the forward LFNST. In the area not modified by the LFNST, the output data of the forward linear transformation remains unchanged.
[0215] As mentioned above, the number of output data points also changes because the dimension of the transformation matrix applied changes depending on the block pattern. As shown in Figure 12, there are cases where the output data of the forward LFNST cannot completely fill the upper left 4x4 block. In the cases of Figure 12(a) and (d), the blocks or parts of the area inside the blocks shown by the thick lines are to which the [16x8] matrix and the [48x8] matrix are applied, respectively, to generate an [8x1] vector as the output of the forward LFNST. That is, according to the scan order shown in Figure 9(b), only 8 output data points are filled as in Figure 12(a) and (d), and the remaining 8 positions can be filled with 0. In the case of the LFNST applied block in Figure 11(d), the two 4x4 blocks at the upper right and lower left corners adjacent to the upper left 4x4 block are also filled with 0 values, as shown in Figure 12(d).
[0216] As described above, the LFNST index is basically signaled to specify whether LFNST can be applied and which transformation matrix to apply. As shown in Figure 12, when LFNST is applied, the number of output data for forward LFNST may be the same as or less than the number of input data, resulting in a region filled with zero values as shown below.
[0217] 1) As shown in Figure 12(a), within the 4x4 block at the top left, the positions from the 8th in the scan order, i.e., samples from the 9th to the 16th.
[0218] 2) As shown in Figures 12(d) and (e), a [16×48] matrix or an [8×48] matrix is applied to two 4×4 blocks adjacent to the top-left 4×4 block, or the second and third 4×4 blocks in the scan order.
[0219] Therefore, if checking the areas in 1) and 2) above reveals the presence of non-zero data, it is certain that LFNST will not be applied, and thus the signaling of the corresponding LFNST index can be omitted.
[0220] For example, in the case of LFNST adopted in the VVC standard, signaling of the LFNST index is performed after residual coding, so the encoding device can know whether or not there is non-zero data (effective factor) for all positions within the TU or CU block via residual coding. Therefore, the encoding device can decide whether or not to perform signaling for the LFNST index based on the presence or absence of non-zero data, and the decoding device can decide whether or not to parse the LFNST index. If there is no non-zero data in the area specified in 1) and 2) above, then signaling of the LFNST index will be performed.
[0221] On the other hand, the following simplification method can be applied to the adopted LFNST.
[0222] (i) For example, the number of output data for a forward LFNST can be limited to a maximum of 16.
[0223] In the case of Figure 11(c), a 4x4 LFNST can be applied to each of the two 4x4 regions adjacent to the upper left corner, in which case a maximum of 32 LFNST output data can be generated. If the number of output data for forward LFNST is limited to a maximum of 16, then for 4xN / Nx4 (N≧16) blocks (TU or CU), a 4x4 LFNST can be applied only to one 4x4 region located at the upper left corner, and LFNST can be applied only once to all blocks in Figure 11. This simplifies the implementation for video coding.
[0224] (ii) For example, zero-out can be additionally applied to regions to which LFNST is not applied. In this document, zero-out can mean filling the values of all positions belonging to a particular region with the value 0. That is, zero-out can also be applied to regions that maintain the result of the forward linear transformation without being modified by LFNST. As mentioned above, since LFNST is divided into 4×4LFNST and 8×8LFNST, zero-out can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.
[0225] (ii)-(A)When 4×4LFNST is applied, areas to which 4×4LFNST is not applied can be zeroed out. Figure 13 shows an example of zeroing out in a block to which 4×4LFNST is applied.
[0226] As shown in Figure 13, for blocks to which 4x4 LFNST is applied, that is, for blocks (a), (b), and (c) in Figure 12, the regions to which LFNST is not applied can all be filled with 0.
[0227] On the other hand, Figure 13(d) shows that when the maximum number of output data for the forward LFNST is limited to 16 as an example, zero-out is performed on the remaining blocks to which the 4x4 LFNST is not applied.
[0228] (ii)-(B) When 8×8LFNST is applied, areas to which 8×8LFNST is not applied can be zeroed out. Figure 14 shows an example of zeroing out in a block to which 8×8LFNST is applied.
[0229] As shown in Figure 14, for blocks to which 8x8 LFNST is applied, that is, for blocks (d) and (e) in Figure 12, the regions to which LFNST is not applied can all be filled with 0.
[0230] (iii) When LFNST is applied by the zero-out method presented in (ii) above, the area filled with zeros can change. Therefore, the zero-out method proposed in (ii) above allows checking for the presence of non-zero data over a wider area than in the case of LFNST in Figure 12.
[0231] For example, when applying (ii)-(B), after checking whether non-zero data exists in the areas filled with zero values in Figure 12 (d) and (e), and further in the areas additionally filled with zeros in Figure 14, signaling to the LFNST index can be performed only if no non-zero data exists.
[0232] Of course, even when applying the zero-out proposed in (ii) above, it is still possible to check whether non-zero data exists, just as with existing LFNST index signaling. That is, LFNST index signaling can be applied by checking whether non-zero data exists for blocks filled with zeros in Figure 12. In such a case, zero-out can be performed only on the encoding device, and LFNST index parsing can be performed on the decoding device without assuming the zero-out, that is, by checking whether non-zero data exists only for areas explicitly indicated as zeros in Figure 12.
[0233] A variety of embodiments can be derived by applying combinations of the simplification methods ((i), (ii)-(A), (ii)-(B), (iii)) to the aforementioned LFNST. Of course, the combinations for the simplification methods are not limited to the embodiments described below, and any combination can be applied to the LFNST.
[0234] Examples
[0235] - Limit the number of output data for forward LFNST to a maximum of 16 → (i)
[0236] When -4×4LFNST is applied, all areas to which 4×4LFNST is not applied are zeroed out → (ii)-(A)
[0237] When -8×8LFNST is applied, all areas to which 8×8LFNST is not applied are zeroed out → (ii)-(B)
[0238] - Check whether non-zero data exists in the regions filled with existing zero values and the regions filled with zeros by additional zero-outs ((ii)-(A), (ii)-(B)), and only if no non-zero data exists, perform LFNST indexing signaling → (iii)
[0239] In the above embodiment, when LFNST is applied, the area in which non-zero output data can exist is limited to the upper left 4x4 area. More specifically, in the cases of Figure 13(a) and Figure 14(a), the 8th position in the scan sequence is the last position in which non-zero data can exist, and in the cases of Figure 13(b) and (d) and Figure 14(b), the 16th position in the scan sequence (i.e., the outermost position at the lower right of the upper left 4x4 block) is the last position in which non-zero data can exist.
[0240] Therefore, when LFNST is applied, it is possible to determine whether LFNST index signaling is permitted after checking whether non-zero data exists at locations where the residual coding process is not permitted (locations beyond the last location).
[0241] In the case of the zero-out method proposed in (ii), the amount of data that ultimately occurs when both the linear transformation and LFNST are applied is reduced, thereby reducing the computational load required when executing the entire transformation process. That is, when LFNST is applied, zero-out is also applied to the forward linear transformation output data in the region where LFNST is not applied, so there is no need to generate data for the region that will be zero-out from the time the forward linear transformation is executed. Therefore, the computational load required to generate the relevant data can be saved. The additional effects of the zero-out method proposed in (ii) can be summarized as follows.
[0242] Firstly, as mentioned above, the amount of computation required to execute the entire transformation process is reduced.
[0243] In particular, when (ii)-(B) is applied, the computational complexity in the worst case is reduced, and the conversion process can be made lighter. To elaborate, generally, a large amount of computation is required when performing a large-size linear conversion, and when (ii)-(B) is applied, the number of data items derived as a result of the forward LFNST execution can be reduced to 16 or less, and the effect of reducing the computational complexity of the conversion increases even more as the overall block (TU or CU) size increases.
[0244] Secondly, the amount of computation required for the entire conversion process is reduced, thereby lowering the power consumption required to perform the conversion.
[0245] Third, reduce the latency associated with the conversion process.
[0246] Quadratic transformations like LFNST add computational complexity to existing linear transformations, thus increasing the overall delay time associated with the transformation execution. In particular, in the case of intra-prediction, since reconstruction data from adjacent blocks is used in the prediction process, the increased delay time due to the quadratic transformation during encoding leads to an increased delay time until reconstruction, which can result in an overall increase in the delay time of intra-prediction encoding.
[0247] However, by applying the zero-out method described in (ii), the delay time for the primary conversion execution can be significantly reduced when LFNST is applied. As a result, the overall delay time for the conversion execution remains the same or decreases, making it easier to implement the encoding device.
[0248] On the other hand, conventional intra-prediction treated the block currently to be encoded as a single encoding unit and performed encoding without division. However, ISP (Intra Sub-Paritions) coding means dividing the block currently to be encoded horizontally or vertically and performing intra-predictive coding. In this case, encoding / decoding is performed on each divided block to generate a reconstructed block, and the reconstructed block can be used as a reference block for the next divided block. For example, during ISP coding, one coding block can be divided into two or four sub-blocks and coded, and in ISP, one sub-block performs intra-prediction by referencing the reconstructed pixel value of the adjacent sub-block located to its left or adjacent upper side. Hereafter, the term "coding" can be used as a concept that includes both coding performed by the encoding device and decoding performed by the decoding device.
[0249] ISP (Individual Partitioning) is the process of dividing a block, as predicted by the Lambda Intra, into two or four subpartitions vertically or horizontally, depending on the block size. For example, the minimum block size to which ISP can be applied is 4x8 or 8x4. If the block size is larger than 4x8 or 8x4, the block will be divided into four subpartitions.
[0250] When ISP is applied, subblocks are coded sequentially depending on the division configuration, for example, horizontally or vertically, from left to right or top to bottom. After the inverse transformation and intra-prediction process for one subblock are performed, and the restoration process is carried out, coding for the next subblock can proceed. For the leftmost or topmost subblock, the system references the restored pixels of the already coded block, as in a typical intra-prediction method. Furthermore, for each edge of a subsequent internal subblock, if it is not adjacent to a previous subblock, the system references the restored pixels of the adjacent coding block that has already been coded, as in a typical intra-prediction method, to derive the reference pixels adjacent to that edge.
[0251] In ISP coding mode, all subblocks can be coded in the same intra-predictive mode, and flags indicating whether to use ISP coding and in which direction (horizontal or vertical) to split can be signaled. In this case, the number of subblocks can be adjusted to 2 or 4 depending on the block pattern, and if the size (width x height) of a single subblock is less than 16, splitting to that subblock can be prohibited, or ISP coding itself can not be applied.
[0252] On the other hand, in ISP prediction mode, one coding unit is divided into two or four partition blocks, i.e., subblocks, for prediction, and the same prediction mode within the same screen is applied to each of the two or four partition blocks.
[0253] As mentioned above, both horizontal and vertical division directions are possible (when an M×N coding unit with horizontal and vertical lengths M and N respectively is divided horizontally, it is divided into M×(N / 2) blocks if it is divided into 2, and into M×(N / 4) blocks if it is divided into 4) and vertical division directions (when an M×N coding unit is divided vertically, it is divided into (M / 2)×N blocks if it is divided into 2, and into (M / 4)×N blocks if it is divided into 4). When divided horizontally, the partition blocks are coded in the order from top to bottom, and when divided vertically, the partition blocks are coded in the order from left to right. The partition block currently being coded can be predicted by referring to the restored pixel values of the upper (left) partition block when the division is horizontal (vertical).
[0254] A transformation can be applied to the residual signal generated by the ISP prediction method on a partition block basis. In addition to the existing DCT-2, a DST-7 / DCT-8 combination-based MTS (Multiple Transform Selection) technique can be applied to the forward-referenced core transform (or primary transform). The transformation coefficients generated by the linear transformation can then be subjected to a forward LFNST (Low Frequency Non-Separable Transform) to produce the final modified transformation coefficients.
[0255] That is, LFNST can also be applied to the partition blocks that are divided when the ISP prediction mode is applied. As described above, the same intra prediction mode is applied to the divided partition blocks. Therefore, when selecting the LFNST set derived based on the intra prediction mode, the LFNST sets derived for all partition blocks can be applied. That is, since the same intra prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks accordingly.
[0256] On the other hand, by way of an example, LFNST can only be applied to transform blocks where both the horizontal and vertical lengths are 4 or more. Therefore, when the horizontal or vertical length of the partition block divided by the ISP prediction method is less than 4, LFNST is not applied and the LFNST index is not signaled either. Also, when applying LFNST to each partition block, the corresponding partition block can be regarded as one transform block. Of course, when the ISP prediction method is not applied, LFNST can be applied to the coding block.
[0257] Looking specifically at applying LFNST to each partition block, it is as follows.
[0258] By way of an example, after applying forward LFNST to an individual partition block, only up to 16 (8 or 16) coefficients are left in the left-upper 4×4 area according to the transform coefficient scanning order, and then zero-out where the remaining positions and areas are all filled with 0 values can be applied.
[0259] Or, by way of an example, when the length of one side of the partition block is 4, LFNST is only applied to the left-upper 4×4 area, and when the lengths of all sides of the partition block, that is, the width and height, are 8 or more, LFNST can be applied to the remaining 48 coefficients excluding the right-lower 4×4 area inside the left-upper 8×8 area.
[0260] Alternatively, as an example, to match the worst-case computational complexity with 8 multiplications per sample, if each partition block is 4x4 or 8x8, only 8 transformation coefficients can be output after applying forward LFNST. That is, if the partition block is 4x4, an 8x16 matrix can be applied to the transformation matrix, and if the partition block is 8x8, an 8x48 matrix can be applied to the transformation matrix.
[0261] On the other hand, in the current VVC standard, LFNST index signaling is performed on a coding unit basis. Therefore, in ISP prediction mode, if LFNST is applied to all partition blocks, the same LFNST index value can be applied to those partition blocks. That is, once an LFNST index value is sent at the coding unit level, that LFNST index can be applied to all partition blocks within the coding unit. As mentioned above, LFNST index values can have values of 0, 1, or 2, where 0 indicates that LFNST is not applied, and 1 and 2 refer to the two transformation matrices that exist within a single LFNST set when LFNST is applied.
[0262] As described above, the LFNST set is determined by the intra-prediction mode, and in the case of ISP prediction mode, all partition blocks within the coding unit are predicted in the same intra-prediction mode, so the partition blocks can refer to the same LFNST set.
[0263] As another example, while LFNST index signaling is still performed on a coding unit basis, in ISP prediction mode, instead of uniformly deciding whether to apply LFNST to all partition blocks, it is possible to decide whether to apply the LFNST index value signaled at the coding unit level to each partition block or not, based on separate conditions. Here, the separate conditions can be signaled in the form of a flag for each partition block via a bitstream, where a flag value of 1 applies the LFNST index value signaled at the coding unit level, and a flag value of 0 does not apply LFNST.
[0264] The following describes how to maintain worst-case computational complexity when applying LFNST to ISP mode.
[0265] In ISP mode, when LFNST is applied, the application of LFNST can be limited to maintain the number of multiplications per sample (or per coefficient, per position) below a certain value. Depending on the size of the partition block, LFNST can be applied as follows to maintain the number of multiplications per sample (or per coefficient, per position) at 8 or less.
[0266] 1. If the width and height of the partition block are both 4 or greater, the same method as the worst-case calculation complexity adjustment method for LFNST in the current VVC standard can be applied.
[0267] In other words, if the partition block is a 4x4 block, instead of a 16x16 matrix, an 8x16 matrix obtained by sampling the top 8 rows from a 16x16 matrix can be applied in the forward direction, and a 16x8 matrix obtained by sampling the left 8 columns from a 16x16 matrix can be applied in the reverse direction. Also, if the partition block is an 8x8 block, instead of a 16x48 matrix in the forward direction, an 8x48 matrix obtained by sampling the top 8 rows from a 16x48 matrix can be applied, and in the reverse direction, instead of a 48x16 matrix, a 48x8 matrix obtained by sampling the left 8 columns from a 48x16 matrix can be applied.
[0268] For 4×N or N×4 (N>4) blocks, when performing a forward transformation, the 16 coefficients generated after applying a 16×16 matrix only to the top-left 4×4 block are placed in the top-left 4×4 region, and the remaining regions can be filled with zeros. Similarly, when performing a reverse transformation, the 16 coefficients located in the top-left 4×4 block can be arranged in scanning order to construct an input vector, which is then multiplied by a 16×16 matrix to generate 16 output data. The generated output data are placed in the top-left 4×4 region, and the remaining regions excluding the top-left 4×4 region can be filled with zeros.
[0269] For 8×N or N×8 (N>8) blocks, when performing a forward transformation, the 16 coefficients generated after applying a 16×48 matrix only to the ROI region within the upper-leftmost 8×8 block (the remaining region after excluding the lower-rightmost 4×4 block from the upper-leftmost 8×8 block) can be placed in the upper-leftmost 4×4 region, and all other regions can be filled with zero values. Similarly, when performing a reverse transformation, the 16 coefficients located in the upper-leftmost 4×4 block can be arranged in scanning order to construct an input vector, which can then be multiplied by a 48×16 matrix to generate 48 output data. The generated output data can fill the ROI region, and all remaining regions can be filled with zero values.
[0270] As another example, to maintain the number of multiplications per sample (or per coefficient, per position) below a certain value, the number of multiplications per sample (or per coefficient, per position) can be kept below 8 based on the ISP coding unit size, which is not the size of the ISP partition block. If there is only one ISP partition block that satisfies the conditions for LFNST to be applied, the complexity calculation for the worst-case LFNST can be applied based on the size of that coding unit, which is not the size of the partition block. For example, if a Luma coding block for a coding unit is divided into four 4x4 partition blocks and coded by ISP, and there are no non-zero conversion coefficients for two of these partition blocks, the other two partition blocks can be set to generate 16 conversion coefficients each (not 8, based on the encoder).
[0271] Below, we will discuss how to signal the LFNST index when using ISP mode.
[0272] As mentioned above, the LFNST index can have values of 0, 1, or 2. 0 indicates that LFNST should not be applied, while 1 and 2 indicate one of the two LFNST kernel matrices included in the selected LFNST set. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The current method by which the LFNST index is transmitted in the VVC standard is as follows:
[0273] 1. An LFNST index can be sent once for each coding unit (CU), and in the case of a dual-tree, separate LFNST indexes can be signaled for the luma block and chroma block, respectively.
[0274] 2. When the LFNST index is not signaled, the LFNST index value is determined (inferred) to be the default value of 0. When the LFNST index value is inferred to be 0, it is as follows.
[0275] A. When it is a mode to which no conversion is applied (for example, transform skip, BDPCM, lossless coding, etc.)
[0276] B. When the primary conversion is not DCT-2 (DST7 or DCT8), that is, when the horizontal conversion or the vertical conversion is not DCT-2
[0277] C. When the horizontal or vertical length of the luma block of the coding unit exceeds the size of the maximum luma conversion that can be performed. For example, when the size of the maximum luma conversion that can be performed is 64, if the size of the luma block of the coding block is 128×16, etc., LFNST cannot be applied.
[0278] In the case of dual tree, it is determined whether each of the coding unit for the luma component and the coding unit for the chroma component exceeds the size of the maximum luma conversion. That is, it is checked whether the horizontal / vertical length of the luma block exceeds the size of the maximum luma conversion that can be performed, and it is checked whether the horizontal / vertical length of the corresponding luma block for the color format and the size of the maximum conversion that can be performed exceed the size of the maximum luma conversion that can be performed for the chroma block. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luma block is twice that of the corresponding chroma block, respectively, and the conversion size of the corresponding luma block is twice that of the corresponding chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical length and the conversion size of the corresponding luma block are the same as those of the corresponding chroma block.
[0279] 64-length conversion or 32-length conversion refers to a conversion applied to a horizontal or vertical length of 64 or 32, respectively, and "conversion size" can mean the corresponding length of 64 or 32.
[0280] If it is a single tree, after checking whether the horizontal or vertical length of the Luma block exceeds the maximum Luma conversion block size that can be converted, LFNST index signaling can be omitted if it does.
[0281] D. The LFNST index can only be sent if the width and height of the coding unit are both 4 or greater.
[0282] In the case of a dual tree, the LFNST index can only be signaled if both the horizontal and vertical lengths for the relevant component (i.e., the luma or chroma component) are 4 or greater.
[0283] In the case of a single tree, the LFNST index can be signaled if the horizontal and vertical lengths of the luma components are all 4 or greater.
[0284] E. If the position of the last non-zero coefficient is not the DC position (top-left corner of the block), in the case of a dual-tree type luma block, the LFNST index is sent when the position of the last non-zero coefficient is not the DC position. In the case of a dual-tree type chroma block, the corresponding LNFST index is sent when either the position of the last non-zero coefficient for Cb or the position of the last non-zero coefficient for Cr is not the DC position.
[0285] In the case of a single-tree type, the LFNST index is sent if the position of the last non-zero coefficient among the Luma component, Cb component, and Cr component is not at the DC position.
[0286] Here, if the CBF (coded block flag) value, which indicates whether a conversion coefficient exists for a particular conversion block, is 0, the position of the last non-zero coefficient for that conversion block is not checked in order to determine whether LFNST index signaling is possible. In other words, if the CBF value is 0, the conversion is not applied to that block, so the position of the last non-zero coefficient is not considered when checking the conditions for LFNST index signaling.
[0287] For example, 1) if it is a dual-tree type and is a luma component, if the corresponding CBF value is 0, the LFNST index will not be signaled; 2) if it is a dual-tree type and is a chroma component, if the CBF value for Cb is 0 and the CBF value for Cr is 1, only the position of the last non-zero coefficient for Cr will be checked and the corresponding LFNST index will be sent; 3) if it is a single-tree type, the position of the last non-zero coefficient will be checked only for components where the CBF value for luma, Cb, and Cr is 1.
[0288] If it is confirmed that a conversion coefficient exists at a position where an F.LFNST conversion coefficient cannot exist, LFNST index signaling can be omitted. In the case of 4x4 and 8x8 conversion blocks, the conversion coefficient scanning order in the VVC standard allows LFNST conversion coefficients to exist at 8 positions from the DC position, and the remaining positions are all filled with 0. In the case of non-4x4 and non-8x8 conversion blocks, the conversion coefficient scanning order in the VVC standard allows LFNST conversion coefficients to exist at 16 positions from the DC position, and the remaining positions are all filled with 0.
[0289] Therefore, if, after performing residual coding, a non-zero conversion coefficient exists in the region where the zero value should be satisfied, LFNST index signaling can be omitted.
[0290] On the other hand, ISP mode can be applied only to luma blocks, or it can be applied to both luma and chroma blocks. As mentioned above, when ISP prediction is applied, the coding unit in question is divided into two or four partition blocks for prediction, and the transformation can be applied to each of these partition blocks. Therefore, when determining the conditions for signaling the LFNST index on a coding unit basis, the fact that LFNST can be applied to each of the partition blocks in question must be taken into consideration. Also, when ISP prediction mode is applied only to a specific component (e.g., luma blocks), the fact that the component is divided into partition blocks only must be taken into consideration when signaling the LFNST index. The possible LFNST index signaling methods when in ISP mode are summarized below.
[0291] 1. An LFNST index can be sent once for each coding unit (CU), and in the case of a dual-tree, separate LFNST indexes can be signaled for the luma block and chroma block, respectively.
[0292] 2. If the LFNST index is not signaled, the LFNST index value is set to the default value of 0 (infer). The cases in which the LFNST index value is inferred to 0 are as follows:
[0293] A. When the mode does not apply a transformation (e.g., transform skip, BDPCM, lossless coding, etc.)
[0294] B. LFNST cannot be applied if the horizontal or vertical length of the coding unit relative to the Luma block exceeds the size of the maximum Luma conversion that can be converted. For example, if the size of the coding block relative to the Luma block is 128 × 16 when the size of the maximum Luma conversion that can be converted is 64.
[0295] Alternatively, the signaling of the LFNST index can be determined based on the size of the partition block instead of the coding unit. That is, if the horizontal or vertical length of the partition block for the relevant luma block exceeds the size of the maximum luma conversion that can be performed, the LFNST index signaling can be omitted and the LFNST index value can be inferred to be 0.
[0296] In the case of a dual tree, it is determined whether the maximum conversion block size is exceeded for each coding unit or partition block for the luma component and for each coding unit or partition block for the chroma component. Specifically, the width and height of the coding unit or partition block for luma are compared to the maximum luma conversion size, and if either is greater than the maximum luma conversion size, LFNST is not applied. In the case of a coding unit or partition block for chroma, the width / height of the corresponding luma block for the color format is compared to the maximum possible luma conversion size. For example, if the color format is 4:2:0, the width / height of the corresponding luma block will be twice that of the corresponding chroma block, and the conversion size of the corresponding luma block will be twice that of the corresponding chroma block. As another example, if the color format is 4:4:4, the width / height and conversion size of the corresponding luma block will be the same as that of the corresponding chroma block.
[0297] In the case of a single tree, after checking whether the horizontal or vertical length of a Luma block (coding unit or partition block) exceeds the maximum Luma conversion block size that can be converted, LFNST index signaling can be omitted if it does.
[0298] C. If LFNST, which is included in the current VVC standard, is applied, an LFNST index can only be sent if the width and height of the partition block are both 4 or greater.
[0299] If we were to apply LFNST to 2×M (1×M) or M×2 (M×1) blocks in addition to the LFNST currently included in the VVC standard, we could only send LFNST indices for partition blocks that are larger than or equal to 2×M (1×M) or M×2 (M×1) blocks. Here, P×Q blocks being larger than or equal to R×S blocks means that P≧R and Q≧S.
[0300] In summary, an LFNST index can only be sent for partition blocks that are larger than or equal to the minimum size for which LFNST is applicable. In the case of a dual tree, an LFNST index can only be signaled for partition blocks for luma or chroma components if they are larger than or equal to the minimum size for which LFNST is applicable. In the case of a single tree, an LFNST index can only be signaled for partition blocks for luma components if they are larger than or equal to the minimum size for which LFNST is applicable.
[0301] In this document, when an M×N block is greater than or equal to a K×L block, it means that M is greater than or equal to K and N is greater than or equal to L. When an M×N block is greater than a K×L block, it means that M is greater than or equal to K, N is greater than or equal to L, and M is greater than K or N is greater than L. When an M×N block is less than or equal to a K×L block, it means that M is less than or equal to K and N is less than or equal to L. When an M×N block is smaller than a K×L block, it means that M is less than or equal to K, N is less than or equal to L, and M is less than K or N is less than L.
[0302] D. If the position of the last non-zero coefficient is not the DC position (top-left corner of the block), in the case of a dual-tree type chroma block, the LFNST index can be sent if the position of the last non-zero coefficient in any one of the partition blocks is not the DC position. In the case of a dual-tree type chroma block, the LNFST index can be sent if the position of the last non-zero coefficient in any of the partition blocks for Cb (assuming there is one partition block if the ISP mode is not applied to the chroma component) and the position of the last non-zero coefficient in any of the partition blocks for Cr (assuming there is one partition block if the ISP mode is not applied to the chroma component) are not the DC position.
[0303] In the case of a single-tree type, the corresponding LFNST index can be transmitted if the position of the last non-zero coefficient in any of the partition blocks for the Luma, Cb, and Cr components is not at the DC position.
[0304] Here, if the CBF (coded block flag) value, which indicates whether a conversion coefficient exists for each partition block, is 0, the position of the last non-zero coefficient for that partition block is not checked in order to determine whether LFNST index signaling is possible. In other words, if the CBF value is 0, the conversion is not applied to that block, so when checking the conditions for LFNST index signaling, the position of the last non-zero coefficient for that partition block is not considered.
[0305] For example, 1) if it is a dual-tree type and contains luma components, when the corresponding CBF value is 0 for each partition block, the corresponding partition block is excluded when determining whether LFNST index signaling is possible. 2) if it is a dual-tree type and contains chroma components, when the CBF value for Cb is 0 and the CBF value for Cr is 1 for each partition block, only the position of the last non-zero coefficient for Cr is checked to determine whether LFNST index signaling is possible. 3) if it is a single-tree type, the position of the last non-zero coefficient is checked only for blocks where the CBF value is 1 for all partition blocks of luma, Cb, and Cr components to determine whether LFNST index signaling is possible.
[0306] In ISP mode, the video information can be configured so that the position of the last non-zero coefficient is not checked, and an example of this is shown below.
[0307] i. In ISP mode, LFNST index signaling can be allowed by omitting the check for the position of the last non-zero coefficient for both luma blocks and chroma blocks. That is, LFNST index signaling can be allowed even if the position of the last non-zero coefficient is the DC position for all partition blocks, or if the corresponding CBF value is 0.
[0308] ii. In ISP mode, the check for the position of the last non-zero coefficient is omitted only for luma blocks, while in chroma blocks, the check for the position of the last non-zero coefficient using the method described above can be performed. For example, in the case of a dual-tree type and a luma block, LFNST index signaling can be allowed without checking the position of the last non-zero coefficient, while in the case of a dual-tree type and a chroma block, the presence or absence of a DC position at the position of the last non-zero coefficient can be checked using the method described above to determine whether or not signaling for the corresponding LFNST index is permitted.
[0309] iii. In the case of ISP mode and single-tree type, method i or ii above can be applied. That is, in the case of ISP mode and single-tree type, when method i is applied, LFNST index signaling can be permitted by omitting the check for the position of the last non-zero coefficient for both luma blocks and chroma blocks. Alternatively, by applying method ii, the check for the position of the last non-zero coefficient can be omitted for partition blocks of luma components, and for partition blocks of chroma components (the number of partition blocks can be considered as 1 if ISP is not applied to chroma components), the check for the position of the last non-zero coefficient can be performed in the aforementioned method to determine whether LFNST index signaling is permitted.
[0310] E. If it is confirmed that a conversion coefficient exists in a location that is not a possible location for any of the partition blocks, then LFNST index signaling can be omitted.
[0311] For example, in the case of a 4x4 partition block and an 8x8 partition block, the conversion coefficient scanning order in the VVC standard allows LFNST conversion coefficients to exist at 8 positions from the DC position, and the remaining positions are all filled with 0. Also, if the partition block is larger than or equal to 4x4, and is not a 4x4 or 8x8 partition block, the conversion coefficient scanning order in the VVC standard allows LFNST conversion coefficients to exist at 16 positions from the DC position, and the remaining positions are all filled with 0.
[0312] Therefore, if a non-zero conversion coefficient exists in the region where the value should be zero after residual coding has been performed, LFNST index signaling can be omitted.
[0313] On the other hand, in ISP mode, the current VVC standard checks the length conditions independently for the horizontal and vertical directions and applies DST-7 instead of DCT-2 without signaling to the MTS index. It is determined whether the horizontal or vertical length is greater than or equal to 4, and less than or equal to 16, and the primary conversion kernel is determined by the result of this determination. Therefore, for ISP mode and when LFNST can be applied, the following conversion combination configurations are possible.
[0314] 1. When the LFNST index is 0 (including cases where the LFNST index is inferred to be 0), the linear transformation determination conditions for ISPs currently included in the VVC standard can be followed. That is, the length conditions (greater than or equal to 4, and less than or equal to 16) are checked independently for the horizontal and vertical directions. If the conditions are met, DST-7 can be applied instead of DCT-2 for the linear transformation; otherwise, DCT-2 can be applied.
[0315] 2. For cases where the LFNST index is greater than 0, the following two configurations are possible with a linear transformation.
[0316] A. DCT-2 can be applied to both horizontal and vertical directions.
[0317] B. The primary conversion determination conditions can be followed when the ISP is currently included in the VVC standard. That is, the length conditions (greater than or equal to 4, and less than or equal to 16) are checked independently for the horizontal and vertical directions. If the conditions are met, DST-7 can be applied instead of DCT-2; otherwise, DCT-2 can be applied.
[0318] In ISP mode, the video information can be configured so that the LFNST index is transmitted per partition block, rather than per coding unit. In such cases, the LFNST index signaling can be enabled or disabled by assuming that there is only one partition block within the unit in which the LFNST index is transmitted using the aforementioned LFNST index signaling method.
[0319] On the other hand, the following section examines the signaling order of the LFNST index and the MTS index.
[0320] For example, the LFNST index, signaled by residual coding, can be coded after the coding position for the last non-zero coefficient position, and the MTS index can be coded immediately after the LFNST index. In such a configuration, the LFNST index can be signaled for each conversion unit. Alternatively, even without signaling by residual coding, the LFNST index can be coded after the coding for the last effective coefficient position, and the MTS index can be coded after the LFNST index.
[0321] The syntax for a typical residual coding example is as follows:
[0322] [Table 4-1]
[0323] [Table 4-2]
[0324] The meanings of the major variables shown in Table 4 are as follows:
[0325] 1. cbWidth, cbHeight: The current width and height of the coding block.
[0326] 2. log2TbWidth, log2TbHeight: Base - 2 log values of the current width and height of the Transform Block. Zero-out is reflected, and non-zero coefficients can exist in the upper left region.
[0327] 3. sps_lfnst_enabled_flag: This flag indicates whether LFNST is applicable (enable). A flag value of 0 indicates that LFNST is not applicable, and a flag value of 1 indicates that LFNST is applicable. It is defined in the Sequence Parameter Set (SPS).
[0328] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponds to the variable chType and the (x0, y0) position. chType can have values of 0 or 1, where 0 indicates the luma component and 1 indicates the chroma component. The (x0, y0) position indicates the position on the picture, and the CuPredMode[chType][x0][y0] value allows for MODE_INTRA (intra prediction) and MODE_INTER (inter prediction).
[0329] 5. IntraSubPartitionsSplit[x0][y0]: The content for the (x0, y0) position is the same as in 4 above. It indicates what kind of ISP splitting was applied at the (x0, y0) position, and ISP_NO_SPLIT indicates that the coding unit corresponding to the (x0, y0) position is not split into partition blocks.
[0330] 6. intra_mip_flag[x0][y0]: The content for the (x0, y0) position is the same as in 4 above. intra_mip_flag is a flag that indicates whether the MIP (Matrix-based Intra Prediction) prediction mode is applied. A flag value of 0 indicates that MIP is not applicable, and a flag value of 1 indicates that MIP is applied.
[0331] 7. cIdx: A value of 0 indicates luma, while values of 1 and 2 indicate the chromatic components Cb and Cr, respectively.
[0332] 8. treeType: Refers to single-tree and dual-tree types (SINGLE_TREE: single-tree, DUAL_TREE_LUMA: dual-tree for luma component, DUAL_TREE_CHROMA: dual-tree for chroma component)
[0333] 9. tu_cbf_cb[x0][y0]: The content for the position (x0, y0) is the same as in 4 above. It indicates the CBF (Coded Block Flag) for the Cb component. If its value is 0, it means that there are no non-zero coefficients in the corresponding transformation unit for the Cb component, and if it is 1, it means that there are non-zero coefficients in the corresponding transformation unit for the Cb component.
[0334] 10. lastSubBlock: Indicates the scan order position of the sub-block (Coefficient Group (CG)) where the last effective coefficient (lastnon-zero coefficient) is located. 0 indicates a sub-block containing the DC component, while a value greater than 0 indicates a sub-block that does not contain the DC component.
[0335] 11. lastScanPos: Indicates the position of the last effective coefficient in the scan order within a subblock. If a subblock consists of 16 positions, the possible values are from 0 to 15.
[0336] 12.lfnst_idx[x0][y0]: This is the LFNST index syntax element to be parsed. If it is not parsed, it is inferred to a value of 0. That is, the default value is set to 0, indicating that LFNST will not be applied.
[0337] 13. LastSignificantCoeffX, LastSignificantCoeffY: Indicates the x and y coordinates in which the last significant coefficient is located within the transformation block. The x coordinate increases from left to right, starting from 0, and the y coordinate increases from top to bottom, starting from 0. If the values of both variables are all 0, it means that the last significant coefficient is located at DC.
[0338] 14. cu_sbt_flag: This flag indicates whether the SubBlock Transform (SBT) currently included in the VVC standard is applicable. A flag value of 0 indicates that SBT is not applicable, and a flag value of 1 indicates that SBT is applicable.
[0339] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: These flags indicate whether explicit MTS has been applied to the interCU and intraCU, respectively. A flag value of 0 indicates that MTS cannot be applied to the interCU or intraCU, while a value of 1 indicates that it can be applied.
[0340] 16. tu_mts_idx[x0][y0]: This is the MTS index syntax element to be parsed. If parsed, it is inferred to a value of 0. That is, the default value is set to 0, indicating that DCT-2 is applied to all horizontal and vertical directions.
[0341] As shown in Table 4, in the case of a single tree, the signaling of the LFNST index can be determined by only the condition of the position of the last effective coefficient for Luma. That is, if the position of the last effective coefficient is not DC, and the last effective coefficient is located inside the upper left subblock (CG), for example, a 4x4 block, the LFNST index is signaled. In the case of the 4x4 and 8x8 transformation blocks, the LFNST index is signaled only if the last effective coefficient is located within the upper left subblock at a position between 0 and 7.
[0342] In the case of a dual tree, the lumana and chromana are each independently signaled with LFNST indices. In the case of chromana, the LFNST index can be signaled by applying the last effective coefficient position condition only to the Cb component. The condition is not checked for the Cr component, and if the CBF value for Cb is 0, the LFNST index can be signaled by applying the last effective coefficient position condition to the Cr component.
[0343] In Table 4, "Min(log2TbWidth, log2TbHeight)>=2" can be expressed as "Min(tbWidth, tbHeight)>=4", and "Min(log2TbWidth, log2TbHeight)>=4" can be expressed as "Min(tbWidth, tbHeight)>=16".
[0344] In Table 4, log2ZoTbWidth and log2ZoTbHeight represent the base-2 (base-2) logarithmic values of the width and height of the upper-left region where the last effective coefficient can exist due to zeroing out.
[0345] As shown in Table 4, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places: firstly, before the MTS index or LFNST index values are parsed, and secondly, after the MTS index has been parsed.
[0346] The first update occurs before the MTS index (tu_mts_idx[x0][y0]) value is parsed, so log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.
[0347] After the MTS index is parsed, log2ZoTbWidth and log2ZoTbHeight are set if the MTS index value is greater than 0 (i.e., a DST-7 / DCT-8 combination). When applying DST-7 / DCT-8 independently to the horizontal and vertical directions in a linear transformation, there can be up to 16 effective coefficients per row or column for each direction. That is, after applying DST-7 / DCT-8 for a length of 32 or more, up to 16 transformation coefficients can be derived per row or column from the left or top. Therefore, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions of a 2D block, effective coefficients can only exist up to a maximum of 16x16 area at the top left.
[0348] Furthermore, when DCT-2 is applied independently to the horizontal and vertical directions in a linear transformation, there can be up to 32 effective coefficients per row or column in each direction. That is, when applying DCT-2 to a length of 64 or more, up to 32 transformation coefficients can be derived from the left or top side for each row or column. Therefore, when DCT-2 is applied to both the horizontal and vertical directions of a 2D block, effective coefficients can only exist up to a maximum of 32x32 area at the top left.
[0349] Furthermore, when DST-7 / DCT-8 is applied to one direction and DCT-2 is applied to the other direction, 16 effective coefficients can exist in the former direction and 32 effective coefficients can exist in the latter direction. For example, in a 64x8 conversion block, where DCT-2 is applied to the horizontal direction and DST-7 is applied to the vertical direction (which can occur in situations where implicit MTS is applied), effective coefficients can exist in a maximum 32x8 area at the upper left corner.
[0350] If log2ZoTbWidth and log2ZoTbHeight are updated in two places, as shown in Table 4, that is, before MTS index parsing, then the ranges of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight, as shown in the table below.
[0351] [Table 5]
[0352] In such cases, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values in the binary evolution process for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.
[0353] [Table 6]
[0354] On the other hand, for example, if ISP mode is enabled and LFNST is applied, the spec text can be constructed as shown in Table 7 when the signaling in Table 4 is applied. Compared with Table 4, the condition for signaling the LFNST index only when ISP mode is not enabled (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT in Table 4) has been removed.
[0355] In the case of a single tree, if the LFNST index sent when it is a luma (when cIdx=0) is to be reused when it is a chroma, the LFNST index sent for the first ISP partition block where the effective coefficient exists can be applied to the chroma conversion block. Alternatively, even in the case of a single tree, the LFNST index can be signaled separately for the chroma component from the luma component. Explanations for the variables listed in Table 7 are as shown in Table 4.
[0356] [Table 7]
[0357] On the other hand, as an example, the LFNST index and / or the MTS index can be signaled at the coding unit level. As mentioned above, the LFNST index can have three values: 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate the first and second candidates, respectively, of the two LFNST kernel candidates included in the selected LFNST set. The LFNST index is coded via truncated unary binarization, and the values 0, 1, and 2 can be coded in bin strings 0, 10, and 11, respectively.
[0358] For example, LFNST can only be applied when DCT-2 is applied to both the horizontal and vertical directions in a linear transformation. Therefore, if the MTS index is signaled after LFNST index signaling, the MTS index can only be signaled if the LFNST index value is 0. If the LFNST index is not 0, the linear transformation can be performed by applying DCT-2 to both the horizontal and vertical directions without signaling the MTS index.
[0359] The MTS index value can have values of 0, 1, 2, 3, and 4, where 0, 1, 2, 3, and 4 indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, and DCT-8 / DCT-8 are applied to the horizontal and vertical directions, respectively. The MTS index can also be coded via truncated unali dimorphism, where the values 0, 1, 2, 3, and 4 can be coded in bin strings 0, 10, 110, 1110, and 1111, respectively.
[0360] LFNST and MTS indexes can be signaled at the coding unit level, and the MTS index can be coded after the LFNST index at the coding unit level. The coding unit syntax table for this is as follows:
[0361] [Table 8]
[0362] The variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag in Table 8 can be set as shown in Table 9 below.
[0363] The variable LfnstDcOnly is set to 1 if the last effective coefficients of a transformation block with a corresponding CBF (Coded Block Flag, 1 if at least one effective coefficient exists in the block, 0 otherwise) value are all located at the DC position (top left corner), and 0 otherwise. More specifically, in the case of a dual-tree luma, the position of the last effective coefficient is checked for one luma transformation block, and in the case of a dual-tree chroma, the position of the last effective coefficient is checked for both the transformation block for Cb and the transformation block for Cr. In the case of a single tree, the position of the last effective coefficient can be checked for the transformation blocks for luma, Cb, and Cr.
[0364] The variable LfnstZeroOutSigCoeffFlag is 0 if an effective coefficient exists at a position where LFNST would result in a zero-out when applied, and 1 otherwise.
[0365] In Table 8 and the following tables, lfnst_idx[x0][y0] indicates the LFNST index for the corresponding coding unit, and tu_mts_idx[x0][y0] indicates the MTS index for the corresponding coding unit.
[0366] As shown in Table 8, the conditions for signaling lfnst_idx[x0][y0] can include a condition (!transform_skip_flag[x0][y0]) that checks whether the transform_skip_flag[x0][y0] value is 0, in which case the condition for checking whether the existing tu_mts_idx[x0][y0] value is 0 (i.e., checking whether both the horizontal and vertical directions are DCT-2) can be omitted.
[0367] The transform_skip_flag[x0][y0] indicates whether the coding unit is coded into a transformation skip mode where transformations are omitted, and this flag is signaled before the MTS index and the LFNST index. That is, because lfnst_idx[x0][y0] is signaled before the tu_mtx_idx[x0][y0] value, only the condition on the transform_skip_flag[x0][y0] value can be checked.
[0368] As shown in Table 8, when coding tu_mts_idx[x0][y0], various conditions are checked, and as mentioned above, tu_mts_idx[x0][y0] is signaled only when the lfnst_idx[x0][y0] value is 0.
[0369] Furthermore, tu_cbf_luma[x0][y0] is a flag indicating whether an effective coefficient exists for the luma component, and cbWidth and cbHeight represent the width and height of the coding unit for the luma component, respectively.
[0370] According to Table 8, when both the width and height of the coding unit for the luma component are 32 or less, tu_mts_idx[x0][y0] is signaled, meaning that the applicability of MTS is determined by the width and height of the coding unit for the luma component.
[0371] In other cases where TU tiling occurs (for example, if the maximum transformation size is set to 32, a 64x64 coding unit is divided into four 32x32 transformation blocks for coding), the MTS index can be signaled based on the size of each transformation block. For example, when both the width and height of a transformation block are 32 or less, the same MTS index value can be applied to all transformation blocks within a coding unit, and the same primary transformation can be applied. Also, when TU tiling occurs, the tu_cbf_luma[x0][y0] value in Table 8 is the CBF value for the top-left transformation block, or it can be set to 1 if at least one transformation block has a corresponding CBF value of 1.
[0372] Furthermore, as shown in Table 8, even in ISP mode, it is possible to configure the system to signal (IntraSubPartitionsSplitType!=ISP_NO_SPLIT)lfnst_idx[x0][y0], and the same LFNST index value can be applied to all ISP partition blocks.
[0373] On the other hand, tu_mts_idx[x0][y0] can only be signaled if it is not in ISP mode (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT).
[0374] As shown in Table 8, when the MTS index is signaled immediately after the LFNST index, information for the linear transformation cannot be known when performing residual coding. That is, the MTS index is signaled after the residual coding. Therefore, the part of the residual coding section that performs zero-out, leaving only 16 coefficients for a 32-length DST-7 or DCT-8, can be modified as shown in Table 9 below.
[0375] [Table 9-1]
[0376] [Table 9-2]
[0377] As shown in Table 9, the part of the process of determining log2ZoTbWidth and log2ZoTbHeight (where log2ZoTbWidth and log2ZoTbHeight represent the base-2 log values of the width and height of the upper-left region remaining after zeroing out, respectively) in which the tu_mts_idx[x0][y0] value is checked can be omitted.
[0378] The binary representations for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 9 can be determined based on log2ZoTbWidth and log2ZoTbHeight, as shown in Table 6.
[0379] Furthermore, as shown in Table 9, when determining log2ZoTbWidth and log2ZoTbHeight using residual coding, a condition can be added to check sps_mts_enable_flag.
[0380] For example, if information regarding the last effective coefficient position for a Luma transform block is recorded during the residual coding process, the MTS index can also be signaled as shown in Table 10.
[0381] [Table 10]
[0382] In Table 10, LumaLastSignificantCoeffX and LumaLastSignificantCoeffY represent the X and Y coordinates of the last effective coefficient position for the Luma transformation block, respectively. A condition has been added to Table 10 that LumaLastSignificantCoeffX and LumaLastSignificantCoeffY must all be less than 16. If even one of them is 16 or greater, DCT-2 is applied to both the horizontal and vertical directions. Therefore, it can be inferred that signaling for tu_mts_idx[x0][y0] is omitted, and DCT-2 is applied to both the horizontal and vertical directions.
[0383] The fact that LumaLastSignificantCoeffX and LumaLastSignificantCoeffY are both less than 16 means that the last significant coefficient lies within the upper left 16x16 region, indicating that, if the 32-length DST-7 or DCT-8 is applied in the current VVC standard, there is a possibility that zero-out was applied, leaving only the leftmost or topmost 16 transformation coefficients. Therefore, tu_mts_idx[x0][y0] can be signaled to indicate the transformation kernel used for the linear transformation.
[0384] On the other hand, as in other examples, the coding unit syntax table, translation unit syntax table, and residual coding syntax table are as shown in the table below. According to Table 11, the MTS index moves to the coding unit level syntax at the translation unit level and is signaled after the LFNST index signaling. Also, the restriction that does not allow LFNST when ISP is applied to the coding unit is removed. Since the restriction that does not allow LFNST when ISP is applied to the coding unit is removed, LFNST can be applied to all intra-prediction blocks. In addition, both the MTS index and the LFNST index are conditionally signaled to the last part at the coding unit level.
[0385] [Table 11]
[0386] [Table 12]
[0387] [Table 13]
[0388] In Table 11, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed by the residual coding in Table 13. The variable MtsZeroOutSigCoeffFlag is changed from 1 to 0 if there is an effective coefficient in the region that should be filled with 0 by zero-out (LastSignificantCoeffX>15||LastSignificantCoeffY>15), in which case the MTS index is not signaled, as in Table 11.
[0389] On the other hand, as shown in Table 11, if tu_cbf_luma[x0][y0] is 0, the mts_idx[x0][y0] coding can be omitted. That is, if the CBF value of the luma component is 0, the transformation is not applied, so there is no need to signal the MTS index, and the MTS index coding can be omitted.
[0390] For example, the aforementioned technical features can be embodied in other conditional constructs. For instance, after MTS is executed, a variable can be derived indicating whether an effective coefficient exists in the region excluding the DC region of the current block, and if the variable indicates that an effective coefficient exists in the region excluding the DC region, the MTS index can be signaled. That is, the existence of an effective coefficient in the region excluding the DC region of the current block indicates that the tu_cbf_luma[x0][y0] value is 1, in which case the MTS index can be signaled.
[0391] The aforementioned variable can be represented as MtsDcOnly, which is initially set to 1 at the coding unit level, and then changed to 0 at the residual coding level if it indicates that an effective coefficient exists in the area excluding the DC area of the current block. When the variable MtsDcOnly is 0, the video information can be configured so that the MTS index is signaled.
[0392] If tu_cbf_luma[x0][y0] is 0, the variable MtsDcOnly will retain its initial value of 1 because no call to the residual coding syntax is made at the conversion unit level in Table 12. In such a case, since the variable MtsDcOnly was not changed to 0, the video information can be configured so that the MTS index is not signaled. That is, the MTS index is not parsed and signaled.
[0393] On the other hand, the decoding device can determine the color index (cIdx) of the conversion coefficients in order to derive the variable MtsZeroOutSigCoeffFlag in Table 13. A color index (cIdx) of 0 means that the luma component is present.
[0394] For example, since MTS can only be applied to the luma component of the current block, the decoding device can determine whether the color index is luma when deriving the variable MtsZeroOutSigCoeffFlag, which determines whether the MTS index can be parsed (ifcIdx==0, MtsZeroOutSigCoeffFlag=0).
[0395] The variable MtsZeroOutSigCoeffFlag indicates whether zero-out was performed when MTS was applied. It shows whether a conversion coefficient exists in an area other than the upper-left region where the last effective coefficient can exist after MTS execution due to zero-out, i.e., the upper-left 16x16 region. The variable MtsZeroOutSigCoeffFlag is initially set to 1 at the coding unit level (MtsZeroOutSigCoeffFlag=1) as shown in Table 11, and if a conversion coefficient exists in an area other than the 16x16 region, its value can be changed from 1 to 0 at the residual coding level (MtsZeroOutSigCoeffFlag=0) as shown in Table 13. If the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.
[0396] As shown in Table 13, at the residual coding level, a non-zero-out region can be set where non-zero conversion coefficients can exist depending on whether zero-out associated with MTS was performed. In this case as well, if the color index (cIdx) is 0, the non-zero-out region can be set to the top-left 16x16 region of the current block.
[0397] Thus, when deriving the variable that determines whether the MTS index can be parsed, we determine whether the color component is luma or chroma. However, since LFNST can be applied to both the luma and chroma components of the current block, we do not determine the color component when deriving the variable that determines whether the LFNST index can be parsed.
[0398] For example, Table 11 shows the variable LfnstZeroOutSigCoeffFlag, which can indicate whether zero-out was performed when LFNST was applied. The variable LfnstZeroOutSigCoeffFlag indicates whether an effective coefficient exists in the second region, excluding the first region at the top left of the current block. This value is initially set to 1, and if an effective coefficient exists in the second region, its value can be changed to 0. The LFNST index can only be parsed if the initially set value of the variable LfnstZeroOutSigCoeffFlag is maintained at 1. When determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, the color index of the current block is not determined because LFNST can be applied to both the luma and chroma components of the current block.
[0399] Figure 15 shows a block diagram of a CABAC encoding system according to one embodiment, and shows a block diagram of CABAC (context-adaptive biary arithmetic coding) for encoding a single syntactic element.
[0400] First, the CABAC encoding process converts the input signal to a binary value via binary code if the input signal is a syntactic element that is not a binary value. If the input signal is already a binary value, it is bypassed without going through binary code, i.e., it is input to the encoding engine. Here, each binary digit 0 or 1 that makes up the binary value is called a bin. For example, if the binary string after binary code is 110, then 1, 1, and 0 are each called one bin. The bins for a syntactic element can represent the value of that syntactic element.
[0401] The binary-coded bin is then fed into either a regular coding engine or a bypass coding engine.
[0402] The canonical coding engine assigns a contextual model that reflects probability values to the given bin, and then encodes the bin based on the assigned contextual model. The canonical coding engine can update the probability model for each bin after encoding it. A bin encoded in this way is called a context-coded bin.
[0403] The bypass coding engine omits the steps of estimating probabilities for the input bin and updating the probability model applied to the bin after coding. It improves coding speed by coding the input bin using a uniform probability distribution instead of assigning context. A bin coded in this way is called a bypass bin.
[0404] Entropy coding allows switching between coding paths, determining whether to perform coding via a normal coding engine or a bypass coding engine. Entropy decoding performs the same process as coding, but in reverse order.
[0405] When a truncated unary code is applied to an LFNST index using a binary method, the LFNST index consists of a maximum of two bins, and the binary codes assigned to the possible LFNST index values 0, 1, and 2 are 0, 10, and 11, respectively.
[0406] For example, context-based CABAC coding can be applied to the first bin of the LFNST index (regular coding), while bypass coding can be applied to the second bin.
[0407] Furthermore, other examples show that context-based CABAC coding can be applied to both the first and second bins of an LFNST index. The assignment of ctxInc to an LFNST index using such context-coded bins can be represented in the following table.
[0408] [Table 14]
[0409] As shown in Table 14, for the first bin (binIdx=0), context 0 can be applied in the case of a single tree, and context 1 can be applied if it is not a single tree. Also, as shown in Table 14, context 2 can be applied to the second bin (binIdx=1). In other words, two contexts can be assigned to the first bin, and one context can be assigned to the second bin, and each context can be distinguished by the ctxInc value (0, 1, 2).
[0410] Here, a single tree means that the luma and chroma components are coded using the same coding structure. After a coding unit is divided using the same coding structure, if the size of the coding unit falls below a certain threshold and the luma and chroma components are coded using separate tree structures, the coding unit can be considered a dual tree, and the context of the first bin can be determined accordingly. That is, the first context can be assigned as shown in Table 14.
[0411] Alternatively, you can code using context 0 if the value of the variable treeType is assigned to the first bin as SINGLE_TREE, and context 1 otherwise.
[0412] For example, when the encoding process attempts only the first (or second) of two LFNST kernel candidates, the LFNST index will have two fixed bin values for that candidate (if only the first candidate is attempted, it can be coded with 10; if only the second candidate is attempted, it can be coded with 11). In this case, if the second bin is coded by bypass, even though the second bin has a fixed value (0 or 1), a fixed amount of bits will be generated for the second bin, which can significantly increase the coding cost for the second bin. If the second bin is coded in a contextual manner without bypass, when only one fixed LFNST kernel candidate is attempted, the probability of the corresponding fixed value (0 or 1) occurring for the second bin is updated to 100%, which can significantly reduce the coding cost. In summary, by coding both bins for LFNST index coding using a context-based approach, even if only one LFNST kernel candidate is fixed and applied during the encoding process, the loss of coding efficiency compared to applying both candidates is minimized while also reducing encoding complexity. This provides substantial flexibility in finding the performance-complexity trade-off during the encoding process.
[0413] The following drawings have been prepared to illustrate a specific example of this specification. The names of specific devices and signals / messages / fields shown in the drawings are presented illustratively, and the technical features of this specification are not limited to the specific names used in the following drawings.
[0414] Figure 16 is a flowchart illustrating the operation of a video decoding device according to one embodiment described in this document.
[0415] Each step disclosed in Figure 16 is based in part on the details described in Figures 5 through 15. Therefore, specific details that overlap with those described in Figures 3 and 5 through 15 are omitted or simplified in their explanation.
[0416] The decoding device 300 according to one embodiment can receive information for the intra prediction mode, residual information, and an LFNST index from the bitstream (S1610).
[0417] This type of information is received as syntax information, which is received as a binary-coded bin string containing 0s and 1s.
[0418] The decoding device can derive binary information for the syntax elements of the LFNST index. This involves generating a candidate set of binary values that the syntax elements of the received transformed index may have, and in this embodiment, the syntax elements of the LFNST index can be binary-coded into a truncated unalicode scheme.
[0419] The syntax element of the conversion index in this embodiment can indicate whether LFNST is applied and one of the LFNST kernels included in the LFNST conversion set. If the LFNST conversion set includes two LFNST kernels, the syntax element of the LFNST index has three values.
[0420] In other words, according to one embodiment, the syntax element value for an LFNST index may include 0, indicating that LFNST is not applied to the current block; 1, indicating the first LFNST kernel among the LFNST kernels; and 2, indicating the second LFNST kernel among the LFNST kernels.
[0421] In this case, the syntax element values for the three LFNST indices can be coded as 0, 10, and 11 using the truncated unalicode method. That is, the value 0 for the syntax element can be binary coded as "0", the value 1 for the syntax element as "10", and the value 2 for the syntax element as "11".
[0422] Furthermore, the decoding device can perform residual coding based on the received residual information to derive conversion coefficients (S1620).
[0423] The decoding device 300 can decode information about the quantized transformation coefficients for the current block from the bitstream and derive the quantized transformation coefficients for the target block based on the information about the quantized transformation coefficients for the current block. The information about the quantized transformation coefficients for the target block can be contained in an SPS (Sequence Parameter Set) or a slice header and may include at least one of the following: information on whether a simplified transformation (RST) is applied, information on the simplification factor, information on the minimum transformation size to apply the simplified transformation, information on the maximum transformation size to apply the simplified transformation, the inverse simplified transformation size, and information on a transformation index that points to one of the transformation kernel matrices contained in the transformation set.
[0424] The decoding device 300 can perform inverse quantization on the current block's residual information, i.e., the quantized conversion coefficients, to derive the conversion coefficients, and can arrange the derived conversion coefficients in a predetermined scanning order.
[0425] Specifically, the derived transformation coefficients can be arranged in a 4x4 block unit according to the reverse diagonal scan order, and the transformation coefficients within a 4x4 block can also be arranged according to the reverse diagonal scan order. In other words, the transformation coefficients after inverse quantization can be arranged according to the reverse scan order applied in VVC and HEVC video codecs.
[0426] The transformation coefficients derived based on such residual information may be inversely quantized transformation coefficients as described above, or they may be quantized transformation coefficients. In other words, the transformation coefficients should be data that can be checked to see if they are non-zero data in the current block, regardless of whether they can be quantized or not.
[0427] The decoding device can derive residual samples by applying an inverse transform to the quantized transformation coefficients.
[0428] As mentioned above, the decoding device can derive residual samples by applying either the non-separating transform LFNST or the separating transform MTS, respectively, which can be performed based on an LFNST kernel, i.e., an LFNST index that points to the LFNST matrix and an MTS index that points to the MTS kernel.
[0429] On the other hand, the decoding device can derive context information for the first bin of the LFNST index and pre-configured context information for the second bin based on the tree type of the current block (S1630).
[0430] The context information for the first bin can be derived as a different value depending on the tree type of the current block. For example, if the tree type of the current block is a single tree, the first bin can be derived as the first context information, and if the tree type of the current block is not a single tree, the first bin can be derived as the second context information.
[0431] On the other hand, context information for the second bin can be derived as a single pre-configured value, regardless of the current block's tree type. That is, for example, the bin string of an LFNST index can be decoded based on a context model where both bins are not bypass-based.
[0432] The decoding device decodes the bin of the syntax element bin string based on context information, and by context-based decoding, it can derive the value of the syntax element for the LFNST index that is applied to the current block from among the binary values that the syntax element of the LFNST index can have (S1640).
[0433] To summarize, the decoding device receives a bin string that has been binary-encoded into a truncated unalicode scheme and decodes the syntax elements of the LFNST index based on contextual information.
[0434] In other words, it is possible to derive whether one of the LFNST indices 0, 1, or 2 is currently applied to the target block.
[0435] The decoding device 300 determines the LFNST conversion set based on the mapping relationship by the intra-prediction mode applied to the currently applied block, and can derive the corrected conversion coefficients by applying LFNST based on the LFNST conversion set, the LFNST kernels included in the LFNST conversion set, and the LFNST index (S1650).
[0436] Subsequently, the decoding device can derive a residual sample for the current block based on the inverse linear transformation of the corrected transformation coefficients (S1660), and generate a restored picture based on the residual sample (S1670).
[0437] The following drawings have been prepared to illustrate a specific example of this specification. The names of specific devices and signals / messages / fields shown in the drawings are presented illustratively, and the technical features of this specification are not limited to the specific names used in the following drawings.
[0438] Figure 17 is a flowchart illustrating the operation of a video encoding device according to one embodiment described in this document.
[0439] Each step disclosed in Figure 17 is based in part on the details described in Figures 5 through 15. Therefore, specific details that overlap with those described in Figures 2 and 5 through 15 are omitted or simplified in their explanation.
[0440] An encoding device 200 according to one embodiment can derive predicted samples for the current block based on an intra-prediction mode applied to the current block (S1710).
[0441] The encoding device 200 can derive the residual sample for the current block based on the predicted sample (S1720).
[0442] The encoding device 200 can apply at least one of LFNST or MTS to the residual sample to derive a conversion coefficient for the current block, and arrange the conversion coefficients in a predetermined scanning order.
[0443] The encoding device can derive conversion coefficients for the current block based on a linear conversion for residual samples (S1730).
[0444] The first-order transformation can be performed via multiple transformation kernels, similar to MTS, in which case the transformation kernel can be selected based on the intra-prediction mode.
[0445] After applying MTS to derive the conversion coefficients, the encoding device can zero out the remaining area of the current block, excluding a specific upper-left region of the current block, for example, a 16x16 region.
[0446] Furthermore, the encoding device 200 can determine whether to perform a quadratic transformation or a non-separable transformation, specifically LFNST, on the transformation coefficients for the current block, and can derive modified transformation coefficients by applying LFNST to the transformation coefficients. Specifically, the encoding device can derive modified transformation coefficients by applying LFNST based on an LFNST transformation set and an LFNST kernel included in the LFNST transformation set (S1740).
[0447] Unlike linear transformations, which separate the coefficients to be transformed vertically or horizontally, LFNST is a non-separated transformation that applies the transformation without separating the coefficients in a specific direction. Such a non-separated transformation is a low-frequency non-separated transformation that applies the transformation only to the low-frequency region, not to the entire target block being transformed.
[0448] The encoding device can encode at least one of the following: an LFNST index pointing to an LFNST kernel or an MTS index pointing to an MTS kernel.
[0449] First, the encoding device derives the syntax element value for the LFNST index and can binary the syntax element value for the LFNST index (S1750).
[0450] The syntax element of the LFNST index in this embodiment can indicate whether LFNST is applied and one of the LFNST kernels included in the LFNST conversion set. If the LFNST conversion set includes two LFNST kernels, the syntax element of the LFNST index has three values.
[0451] In one embodiment, the syntax element value for an LFNST index can be derived as follows: 0, indicating that LFNST is not applied to the current block; 1, indicating the first LFNST kernel among the LFNST kernels; and 2, indicating the second LFNST kernel among the LFNST kernels.
[0452] The encoding device can binary the syntax element values for the three conversion indices using the truncated unalicode method with 0, 10, and 11. That is, the value 0 for the syntax element can be binaryized as "0", the value 1 for the syntax element can be binaryized as "10", and the value 2 for the syntax element can be binaryized as "11", and the encoding device can binary the syntax element for the derived conversion indices with any one of "0", "10", and "11".
[0453] The encoding device can derive context information for the first bin of the LFNST index based on the current block's tree type, derive pre-configured context information for the second bin, and encode the bin of the syntax element bin string based on the context information (S1760).
[0454] The context information for the first bin can be derived as a different value depending on the tree type of the current block. For example, if the tree type of the current block is a single tree, the first bin can be derived as the first context information, and if the tree type of the current block is not a single tree, the first bin can be derived as the second context information.
[0455] On the other hand, context information for the second bin can be derived as a single value regardless of the current block's tree type. That is, for example, the bin string of an LFNST index can be encoded based on a context model where both bins are not bypass-based.
[0456] To summarize, the encoding device binary-codes the bin string of the LFNST index into a truncated unalicode scheme and encodes the syntax elements of the LFNST index for the corresponding binary-coded value based on different contextual information.
[0457] The encoding device can configure and output video information such that at least one of the LFNST index and the MTS index is signaled at the coding unit level, and the MTS index is signaled immediately after the LFNST index is signaled, and can output after quantized residual information encoding (S1770).
[0458] Furthermore, the encoding device can perform quantization based on the conversion coefficients for the current block or modified conversion coefficients to derive quantized conversion coefficients, and encode and output video information including information about the quantized conversion coefficients.
[0459] The encoding device can generate residual information containing information about quantized conversion coefficients. This residual information may include the detailed conversion-related information / syntax elements. The encoding device can encode the video information containing this residual information and output it in bitstream format.
[0460] More specifically, the encoding device can generate information about the quantized transformation coefficients and encode the generated information about the quantized transformation coefficients.
[0461] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If transformation / inverse transformation is omitted, the transformation coefficients may also be called coefficients or residual coefficients, or for consistency of expression, they may still be called transformation coefficients.
[0462] Furthermore, in this document, quantized transformation coefficients and transformation coefficients can be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients, and such information about the transformation coefficients can be signaled via residual coding syntax. Transformation coefficients can be derived based on the residual information (or information about the transformation coefficients), and scaled transformation coefficients can be derived via an inverse transformation (scaling) of the transformation coefficients. Residual samples can be derived based on an inverse transformation (transformation) of the scaled transformation coefficients. This can be applied / expressed similarly in other parts of this document.
[0463] In the embodiments described above, the method is explained based on a flowchart in a series of steps or blocks, but this document is not limited to the order of the steps, and some steps may occur with other steps, in a different order, or simultaneously. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of this document.
[0464] The method described in this document above can be implemented in software form, and the encoding and / or decoding devices described in this document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, and display devices.
[0465] In this document, when embodiments are embodied in software, the methods described above can be embodied in modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by a variety of well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described in this document can be embodied and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be embodied and executed on a computer, processor, microprocessor, controller, or chip.
[0466] Furthermore, decoding and encoding devices to which this document applies may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction equipment, real-time communication equipment such as video communications, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, image-phone video equipment, and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) video equipment may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0467] Furthermore, the processing methods to which this document applies can be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Multimedia data having the data structure described in this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data readable by a computer. The computer-readable recording medium can include, for example, Blu-ray discs (BDs), general-purpose serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored on a computer-readable recording medium or transmitted over a wireless network. The embodiments of this document can also be embodied in computer program products using program code, and the program code can be executed on a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.
[0468] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined and embodied in an apparatus, and the technical features of the apparatus claims herein can be combined and embodied in a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined and embodied in an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined and embodied in a method.
Claims
1. In a video decoding method performed by a decoding device, A step of receiving residual information from a bitstream, The steps include: deriving a conversion coefficient for the current block based on the residual information; The steps include: deriving a residual sample by applying at least one of LFNST (low frequency non-separable transform) or an inverse linear transform to the aforementioned transformation coefficients; The steps include generating a restored picture based on the said residual sample, The context increment for the first bin of the LFNST index is derived based on whether the tree type of the current block is a single tree, and the context increment for the second bin of the LFNST index is derived as a default value. Whether or not to parse the MTS index associated with the MTS (multiple transform selection) kernel for the inverse linear transformation is determined based on a first variable. The first variable is set to an initial value equal to 1 in the coding unit syntax, Based on the fact that at least one non-zero conversion coefficient exists in a region other than the upper-left 16x16 region within the current block, and the color index of the current block is a luma component, the first variable is changed to 0 in the residual coding syntax. A video decoding method in which the MTS index is parsed based on the fact that the first variable is equal to 1.
2. The video decoding method according to claim 1, further comprising the step of decoding the bin string of the LFNST index based on the context increments of the first bin and the second bin of the LFNST index to derive a value of the LFNST index.
3. Based on the fact that the tree type of the current block is the single tree, the context increment for the first bin is derived as the first value. The video decoding method according to claim 2, wherein, based on the fact that the tree type of the current block is not the single tree, the context increment for the first bin is derived as a second value.
4. The video decoding method according to claim 3, wherein the context increment for the second bin is derived as a third value different from the first and second values.
5. The LFNST conversion set includes two LFNST kernels, The video decoding method according to claim 2, wherein the value of the LFNST index includes one of the following: 0, which is associated with the case where the LFNST is not applied to the current block; 1, which is associated with the first of the two LFNST kernels; and 2, which is associated with the second of the two LFNST kernels.
6. The value of the LFNST index is binary-encoded as a truncated unalicode. The video decoding method according to claim 5, wherein the LFNST index having a value equal to 0 is binary as "0", the LFNST index having a value equal to 1 is binary as "10", and the LFNST index having a value equal to 2 is binary as "11".
7. In a video encoding method performed by an encoding device, The current step is to derive a predicted sample for the block, The steps include: deriving a residual sample for the current block based on the predicted sample; The steps include: deriving the conversion coefficient for the current block by applying at least one of LFNST (low frequency non-separable transform) or a linear transform to the residual sample; The process includes the step of encoding an LFNST index related to the LFNST and an MTS (multiple transform selection) index related to the primary transformation, The context increment for the first bin of the LFNST index is derived based on whether the tree type of the current block is a single tree, and the context increment for the second bin of the LFNST index is derived as a default value. Whether or not to encode the aforementioned MTS index is determined based on the first variable, The first variable is set to an initial value equal to 1 in the coding unit syntax, Based on the fact that at least one non-zero conversion coefficient exists in a region other than the upper-left 16x16 region within the current block, and the color index of the current block is a luma component, the first variable is changed to 0 in the residual coding syntax. A video encoding method in which the MTS index is encoded based on the fact that the first variable is equal to 1.
8. The step of encoding the LFNST index is: The steps include: deriving the value of the LFNST index, The steps include: deriving the context increment for the first and second bins of the LFNST index; The video encoding method according to claim 7, comprising the step of encoding the bin string of the LFNST index based on the context increment.
9. Based on the fact that the tree type of the current block is the single tree, the context increment for the first bin is derived as the first value. The video encoding method according to claim 8, wherein, based on the fact that the tree type of the current block is not the single tree, the context increment for the first bin is derived as a second value.
10. The video encoding method according to claim 9, wherein the context increment for the second bin is derived as a third value different from the first and second values.
11. The LFNST conversion set includes two LFNST kernels. The video encoding method according to claim 8, wherein the value of the LFNST index includes one of the following: 0, which is associated with the case where the LFNST is not applied to the current block; 1, which is associated with the first of the two LFNST kernels; and 2, which is associated with the second of the two LFNST kernels.
12. The value of the LFNST index is binary-encoded as a truncated unalicode. The video encoding method according to claim 11, wherein the LFNST index having a value equal to 0 is binary as "0", the LFNST index having a value equal to 1 is binary as "10", and the LFNST index having a value equal to 2 is binary as "11".
13. In a method for transmitting data related to video information, A step of obtaining a bitstream of video information including an LFNST (low frequency non-separable transform) index and residual information, wherein the bitstream is The current step is to derive a predicted sample for the block, The steps include: deriving a residual sample for the current block based on the predicted sample; The steps include: deriving a conversion coefficient for the current block by applying at least one of LFNST or a linear transformation to the residual sample; A step of encoding the LFNST index related to the LFNST and the MTS (multiple transform selection) index related to the primary transformation, and a step of generating the result, The step of transmitting the data, which includes the bitstream of the video information, The context increment for the first bin of the LFNST index is derived based on whether the tree type of the current block is a single tree, and the context increment for the second bin of the LFNST index is derived as a default value. Whether or not to encode the aforementioned MTS index is determined based on the first variable, The first variable is set to an initial value equal to 1 in the coding unit syntax, Based on the fact that at least one non-zero conversion coefficient exists in a region other than the upper-left 16x16 region within the current block, and the color index of the current block is a luma component, the first variable is changed to 0 in the residual coding syntax. A method in which the MTS index is encoded based on the fact that the first variable is equal to 1.